Electronic apparatus, method therefor, and computer-readable recording medium
Patent Information
- Application Number
- PCT/KR2026/001685
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-03
- Filing Date
- 2026-01-28
- Publication Date
- 2026-08-27
Smart Images

Figure KR2026001685_27082026_PF_FP_ABST
Abstract
Description
Electronic device, method, and computer-readable recording medium
[0001] The present disclosure relates to an electronic device, a method, and a computer-readable recording medium for image processing.
[0002] Image processing technologies can include techniques such as color correction, filter application, and resolution adjustment. Recently, with the advancement of machine learning and AI-based image processing technologies, more sophisticated image quality enhancement, automatic correction, and content recognition have become possible.
[0003] Meanwhile, object recognition and segmentation technologies, which identify specific objects contained in an image and separate their regions, can be utilized in image processing. After performing object recognition and segmentation, image processing regarding specific objects can be carried out. For example, image editing such as object deletion, background modification, and object style transformation is possible. With the advancement of artificial intelligence models (e.g., DNNs (deep neural networks) or CNNs (convolutional neural networks)), object recognition, object segmentation, and image processing technologies are becoming more sophisticated.
[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0005] According to one embodiment of the present disclosure, an electronic device may be provided. The electronic device may include a memory comprising at least one storage medium in which at least one instruction is stored, and at least one processor capable of executing said at least one instruction. The at least one processor may control a method of operation of the electronic device by executing said at least one instruction.
[0006] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: receive a plurality of image processing requests corresponding to each of a first plurality of images; identify at least one target object and at least one image processing type corresponding to the plurality of image processing requests; estimate a user intent corresponding to the plurality of image processing requests based on the fact that the plurality of image processing requests satisfy a specified condition; determine a second plurality of images according to the estimated user intent; provide a query for verifying the estimated user intent; and, based on a user response to the query, perform image processing corresponding to the at least one target object and the at least one image processing type in at least one image included in the second plurality of images.
[0007] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: perform a first type of image processing on a first object included in at least one first image in at least one first image based on a first request, count a first number of times the image processing based on the first request was performed, and, based on identifying that the first number is greater than or equal to a first threshold, perform the first type of image processing on a first object included in at least one first target image in at least one first target image, thereby obtaining at least one first edited image corresponding to each of the at least one first target image.
[0008] According to one embodiment of the present disclosure, a method of an electronic device may be provided. The method may include at least some of at least one operation. The at least one operation may include: an operation of performing a first type of image processing on a first object included in at least one first image in at least one first image based on a first request; an operation of counting a first number of times the image processing based on the first request has been performed; and an operation of performing the first type of image processing on a first object included in at least one first target image in at least one first target image based on identifying that the first number is greater than or equal to a first threshold, thereby obtaining at least one first edited image corresponding to each of the at least one first target image.
[0009] According to one embodiment of the present disclosure, a non-transient computer-readable recording medium may be provided for storing instructions that, when executed by at least one processor, cause said at least one processor to perform set operations. The set operations may include: an operation of performing a first type of image processing on a first object included in said at least one first image in at least one first image based on a first request; an operation of counting a first number of times the image processing based on said first request has been performed; and / or an operation of performing the first type of image processing on a first object included in said at least one first target image in at least one first target image based on identifying that said first number is greater than or equal to a first threshold, thereby obtaining at least one first edited image corresponding to each of said at least one first target image.
[0010] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0011] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments of the present disclosure.
[0012] FIG. 2 is a block diagram of a generative artificial intelligence (AI) system according to one embodiment.
[0013] FIG. 3 is a block diagram of an AI framework according to one embodiment.
[0014] FIG. 4 is a drawing for explaining an image processing operation according to one embodiment of the present disclosure.
[0015] FIGS. 5A, FIGS. 5B, and FIGS. 6 illustrate screens displayed by an electronic device according to an embodiment of the present disclosure.
[0016] FIG. 7 illustrates screens that change according to the operation of an electronic device according to one embodiment of the present disclosure.
[0017] FIG. 8 is a drawing for explaining the operation of an electronic device according to one embodiment of the present disclosure.
[0018] FIG. 9 is a drawing for explaining the operation of an electronic device according to one embodiment of the present disclosure.
[0019] FIG. 10 is a drawing for explaining relationship information according to one embodiment of the present disclosure.
[0020] FIG. 11 is a drawing for explaining the operation of an electronic device according to one embodiment of the present disclosure.
[0021] FIG. 12 is a drawing for explaining image processing according to one embodiment of the present disclosure.
[0022] FIG. 13 illustrates a screen displayed by an electronic device according to one embodiment of the present disclosure.
[0023] FIG. 14 is a drawing for explaining the operation of an electronic device according to one embodiment of the present disclosure.
[0024] FIGS. 15 and 16 are drawings for illustrating image processing and the context of an image according to one embodiment of the present disclosure.
[0025] FIG. 17 is a drawing for explaining an image and information related to the image according to one embodiment of the present disclosure.
[0026] FIG. 18 is a diagram illustrating the operation of updating an image on a cloud server according to image processing in an electronic device according to one embodiment of the present disclosure.
[0027] FIG. 19 is a drawing for explaining the operation of an electronic device according to one embodiment of the present disclosure.
[0028] FIG. 20 is a block diagram of an electronic device according to one embodiment of the present disclosure.
[0029] FIG. 21 is a drawing for explaining the operation of an electronic device according to one embodiment of the present disclosure.
[0030] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and should be understood to include various modifications, equivalents, or substitutions of the embodiments described herein, rather than being limited to the embodiments described herein. The present disclosure is capable of various modifications by those skilled in the art without departing from the gist of the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.
[0031] The purposes and effects of the present disclosure are not limited to those mentioned in the drawings and the following related description, and various modifications may be made within the technical scope of the present disclosure. The effects according to the embodiments of the present disclosure mentioned below are merely illustrative and are not limited thereto; depending on various modifications, different or additional effects may be realized.
[0032] In the following drawings and related descriptions, functions, configurations, technical terms, and technical details well known in the art to which this disclosure pertains may be omitted. This is intended to convey the essentials of this disclosure more clearly and concisely by minimizing unnecessary detailed descriptions.
[0033] In the drawings, each block of the flowcharts and combinations of the flowcharts may be performed by at least one instruction. The instruction may be loaded into a processor of a computer or other programmable data processing equipment to generate means for performing the functions described in the drawings. The instruction may also provide steps for performing the functions described in the drawings by being executed on a computer or other programmable data processing equipment.
[0034] Meanwhile, various elements and areas in the drawings are depicted schematically, and the technical concept of the present disclosure is not limited by the relative sizes, spacing, or arrangements depicted in the attached drawings. The electronic device (101) of the present disclosure is not limited to the configuration and / or operation shown in the drawings and may include all other configurations capable of performing the same or similar functions.
[0035] The individual components depicted in the drawings are not required to be implemented in a physically separate form, but are shown separately to aid in the description and understanding of the present disclosure. The present disclosure may be implemented in a form in which the individual components shown in the drawings are merged, modified, or have some components deleted and / or added. Each component may perform functions in conjunction with one another while existing in physically separated locations via a network or communication link.
[0036] Likewise, the operations depicted in the drawings are illustrative to aid in the description and understanding of the present disclosure, and the present disclosure may be modified by merging, changing the order of, or deleting and / or adding parts of the operations shown in the drawings. For example, two or more operations shown consecutively in the drawings may be performed simultaneously, in reverse order as necessary, repeatedly, or omitted depending on the actual situation.
[0037] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0038] The processor (120) can control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., a program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), a secure processing unit (SPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0039] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence is performed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network (DQN), or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0040] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0041] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146). One or more related applications may form a service configured to handle a series of user requests.
[0042] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0043] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0044] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0045] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0046] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0047] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0048] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0049] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0050] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0051] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) may be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0052] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0053] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., local area network (LAN) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0054] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. The NR access technology can support enhanced mobile broadband (Embb), massive machine type communications (mMTC), or ultra-reliable and low-latency communications (URLLC). The wireless communication module (192) can support high-frequency bands (e.g., mmWave bands) to achieve high data transmission rates, for example. The wireless communication module (192) can support various technologies for securing performance in high-frequency bands, for example, beamforming, multiple-input and multiple-output (massive MIMO), full-dimensional MIMO (FD-MIMO), array antennas, analog beamforming, or large-scale antennas. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., electronic device (104)), or a network system (e.g., a second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0055] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0056] According to one embodiment, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0057] At least some of the above components can be connected to each other and exchange signals (e.g., commands or data) through a communication method between peripheral devices (e.g., bus, general purpose input and output (GPIO), serial peripheral interface (SPI), or mobile industry processor interface (MIPI).
[0058] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least a part of the requested function or service, or additional functions or services related to the request, and transmit the result of the execution to the electronic device (101).
[0059] The electronic device (101) may process the above results, either as is or additionally, and provide them as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services, for example, by using distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199). The electronic device (101) may be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0060] FIG. 2 is a generative artificial intelligence (AI) system (200) according to one embodiment. Referring to FIG. 2, the generative AI system (200) may include a user interface (210), an AI framework (220), a generative AI model (230), a knowledge repository (240), and an application / service module (250). These components may be operated on one or more of an electronic device (101), an external electronic device (102 or 104), or a server (108). For example, the user interface (210) and the AI framework (220) may be operated on the electronic device (101), and the knowledge repository (240) and the generative AI model (230) may be operated on the server (108).
[0061] According to one embodiment, a user interface (210) may receive user input (e.g., user query). User input may be received in the form of text, images, voice (e.g., natural language), video, menu selection, or a combination thereof. The user interface (210) may include various context information (e.g., running application or user location) related to the generative artificial intelligence system (200) at the time the user input is received, in addition to or instead of the user input. The user interface (210) may provide the user input or the context information to the AI framework (220) and provide the result of processing therefrom to the user, for example, through the AI framework (220). According to one embodiment, in addition to user input, the electronic device may provide context information obtained using information included on the screen to the AI framework (220). The result may be provided in the form of text, images, voice, video, an action requested by the user (e.g., execution of a specified function or app), or a combination thereof.
[0062] According to one embodiment, the AI framework (220) can identify (e.g., estimate) a user intent based on at least some user input or context information received from a user interface (210), control each of the relevant modules (e.g., 221, 223, or 225) to perform a function or action corresponding to the identified user intent, and coordinate collaboration between two or more modules. The AI framework (220) may include a prompt design module (221), an API / plug-in management module (223), and an output management module (225), as illustrated in FIG. 2.
[0063] According to one embodiment, the prompt design module (221) can generate a prompt to be input to a generative AI model (230) based at least partially on user input or context information received from a user interface (210). For example, the prompt design module (221) can generate a prompt using user preferences, a prompt library, or prompt examples stored in a knowledge repository (240) based at least partially on user input or context information.
[0064] According to one embodiment, the API / plugin management module (223) may communicate, for example, via an API, with various resources (e.g., a knowledge repository (240)) that provide said additional information when there is a request for said additional information in relation to user input. Additionally or alternatively, when a specified action (e.g., a function, app, or service) is performed in response to said user input, the API / plugin management module (223) may request the application / service module (250) to perform said specified action via a corresponding API. The API / plugin management module (223) may provide information obtained from the knowledge repository (240), the application / service module 250, or another external resource to the prompt design module (221). The obtained information may be used by the prompt design module (221) to generate a prompt together with the user input or provided to a generative AI model (230).
[0065] According to one embodiment, the output processing module (225) can fine-tune the results obtained through the generative AI model (230) as at least part of the response to user input (e.g., user query). For example, the output processing module (225) can determine whether the content of the response obtained through the generative AI model (230) is appropriate as a response to a request made by the user input. For example, the output processing module (225) can determine the degree of relevance, degree of bias (e.g., political or social bias), or degree of harmfulness (e.g., sexual or profanity) of the difference between the response obtained through the generative AI model (230) and the user input. Additionally or generally, the output processing module (225) can request that additional AI processing be performed on the obtained response, or provide the user with a hint to avoid unwanted output. For example, the response can be obtained again through the generative AI model (230) by generating an additional prompt through the prompt design module.
[0066] According to one embodiment, the generative AI model (230) may form at least part of an artificial intelligence neural network and may include a model that generates images or a model that generates language. The image generation model may include, for example, a generative adversarial network (GAN), a variational autoencoder (VAE), or a Diffusion-based model using a VAE and a Transformer. The language generation model may include, for example, a large language model (LLM), a large multimodal model (LMM), a large vision model (LVM), or a large action model (LAM). The LAM may automatically generate actions for an environment (e.g., a robot, a car, an electronic device (101), or a program (140)). Additionally, for at least some AI models (e.g., LLM), there may be a low-rank adaptation (LoRA) adapter that is fine-tuned for, for example, a specific task or a specific situation.
[0067] FIG. 3 illustrates an AI framework (220) having on-device AI processing capabilities according to one embodiment. In this case, the AI framework (220) may generate and learn a response to the user input using resources within the device, instead of sending the user input received through a user interface (210) operating on the same device (e.g., electronic device (101)) to a generative AI model (230) operating on an external device (e.g., server (108)), or additionally. Referring to FIG. 3, the AI framework (220) may include a cross-application action module (310), a personal data management module (330), an on-device AI model (350), and an orchestration module (370).
[0068] According to one embodiment, the cross-application action module (310) determines one or more additional applications required for the operation of an executed application (e.g., an assistant app) and may link or suggest operations between the app and at least one additional application, or between a plurality of additional applications. For example, the cross-application action module (310) may execute one or more additional applications to be used to respond to a user request through the assistant app sequentially or at least partially, simultaneously. Additionally, the cross-application action module (310) may communicate with the additional applications so that the result of the execution of one additional application (e.g., content) can be shared with other additional applications.
[0069] According to one embodiment, the personal data management module (330) may provide personal information (e.g., schedule, contact, or message information) about a user of the application (e.g., assistant app) or the additional application running on the device (e.g., electronic device 101) or other related individuals (e.g., family or friends) to another module of the AI framework (220) or a related module (e.g., generative AI model 230) running on another device.
[0070] According to one embodiment, the on-device AI model (350) may include at least one model among one or more AI models (e.g., GAN, VAE, LLM, LMM, LVM, or LAM) operated on an external device (e.g., server (108)) or a corresponding lightweight AI model. Additionally, for said model or said lightweight model, there may be, for example, a LoRA adapter.
[0071] According to one embodiment, the orchestration module (370) may select one or more AI models to be used to obtain a response to user input (e.g., user query). For example, the orchestration module (370) may select one or more AI models from an on-device AI model (350), an AI model operating on an external device (e.g., server (108)) (e.g., generative AI model (230)), or a third AI model (not shown) operating on another external device. When multiple AI models are selected, the orchestration module (370) may communicate with the selected models or devices so that the operation between the selected AI models and the processing of the results thereof can be coordinated between the relevant models or devices.
[0072] According to one embodiment, two or more modules of a generative AI system (200) (e.g., a cross-application action module (310) and an orchestration module (370)) may be implemented as a single module to maintain the same functionality. Various variations are possible.
[0073] FIG. 4 is a drawing for explaining an image processing operation according to one embodiment of the present disclosure.
[0074] The 'image' described below in this disclosure may include static and / or dynamic images, i.e., still images and / or videos. For example, the 'image processing' described below may be understood to include the editing of at least one static image and / or video.
[0075] Referring to FIG. 4, according to one embodiment, an electronic device (e.g., the electronic device (101) of FIG. 1) can display a first image (410) through a display (e.g., the display module (160) of FIG. 1). The first image (410) may include at least one object, including a first object (411).
[0076] According to one embodiment, the electronic device (101) can identify at least one object included in the first image (410). For example, the electronic device (101) can identify the first object (411) using an artificial intelligence system (e.g., the generative artificial intelligence system (200) of FIG. 2).
[0077] According to one embodiment, the electronic device (101) may perform object segmentation to identify the region and / or boundary of an object included in an image. As a result of the object segmentation, the electronic device (101) may obtain mask data for the object. The electronic device (101) may, for example, apply image processing (e.g., object deletion) only to a specific region within an image (e.g., first image (410)) corresponding to the mask data for a specific object (e.g., first object (411)).
[0078] According to one embodiment, an electronic device (101) may perform image processing related to the first object (411) in a first image (410) containing the first object (411). Image processing related to the object may include analysis of the object (e.g., identification, classification) and / or visual adjustment (editing, transformation). Image processing related to the first object (411) may include at least one of various types of image processing, such as, for example, deletion of all or part of the first object, deletion of all or part of other parts (e.g., background or other objects) while leaving the first object, application of visual effects (e.g., brightness adjustment, sharpness adjustment, color adjustment, blurring, image compositing), resizing, shape modification, adding objects, and cutting out the first object.
[0079] According to one embodiment, image processing may include image processing using a generative artificial intelligence model. For example, image processing may include inpainting, which is a process of filling empty or damaged areas within a defined area of an image, and / or outpainting, which is a process of extending the image beyond an existing area.
[0080] According to one embodiment, inpainting may include a technique for filling in areas removed and / or damaged areas in relation to image processing based on surrounding pixel or object information. Inpainting may be utilized for processing such as, for example, restoring obscured (or damaged) parts, supplementing areas after removing unnecessary objects, noise removal, and image quality improvement.
[0081] According to one embodiment, outpainting may include a technique for predicting and / or expanding the remaining area of an object or background that is only partially represented within an image. Outpainting may be utilized for processing such as background expansion, field of view expansion, and the creation of additional graphic content.
[0082] According to one embodiment, if no other object other than the first object (411) exists in the first image (410), the electronic device (101) may delete the image file instead of performing image processing to delete the first object (411). To this end, the electronic device (101) may provide a query to the user regarding whether to perform image processing to delete the first object (411) within the image (410) or to delete the file itself.
[0083] According to one embodiment, if the first object (411) is a specific person, the electronic device (101) can identify the name of the object and / or the relationship between the user and the first object (411). For example, the electronic device (101) can identify what the name of the first object (411) is and what the relationship between the first object (411) and the user is based on user personalization information (e.g., contact information).
[0084] According to one embodiment, the electronic device (101) may perform image processing on the first object (411) and may also perform processing that is applied entirely to parts other than the first object (411). For example, the electronic device (101) may change visual characteristics such as brightness, luminance, saturation, mood, size, and resolution of the image. For example, the electronic device (101) may perform editing to delete or add audio to the video. For example, the electronic device (101) may delete the image or move the storage location.
[0085] According to one embodiment, when the first object (411) is a person object (including the face of a specific person), image processing for the first object (411) may include processing such as skin correction, eye, nose, and mouth correction, and expression change (e.g., correcting closed eyes to open eyes, correcting a blank expression to a smiling face).
[0086] According to one embodiment, the electronic device (101) can obtain and display a first edited image (420) by performing a first type of image processing on a first object (411). In FIG. 4, the first type of image processing is exemplified as object deletion, but is not limited thereto.
[0087] According to one embodiment, the electronic device (101) may determine that the user's intention is to perform the same type of image processing on the same object based on the user repeatedly performing the first type of image processing on the first object. Based on identifying the user's intention to perform the first type of image processing on the first object, the electronic device (101) may perform the first type of image processing on the first object on at least one other image.
[0088] According to one embodiment, the user's intention may be determined based on the number of times the same type of image processing is performed on the same object. The same object may include, for example, an object identified as the same person based on face recognition technology. For example, an electronic device (101) may count the number of times a first type of image processing is performed on a first object and determine the user's intention to perform the first type of image processing on the first object based on the number of times the number exceeds a threshold value.
[0089] According to one embodiment, the electronic device (101) may determine (or estimate) user intent by comparing only the number of times image processing is performed within a specified period with a threshold. For example, the electronic device (101) may compare only the number of times image processing of the first type for the first object has been performed during the last month with a threshold, and may not consider image processing of the first type for the first object performed earlier. For example, if the specified period elapses without satisfying a specific condition (e.g., exceeding a threshold), the electronic device (101) may reset (e.g., initialize to 0) the number of times image processing of the first type for the first object has been performed.
[0090] According to one embodiment, the electronic device (101) may also count the number of image processings performed by an external electronic device. For example, if a user performs a first image processing of a first object n times on the electronic device (101), and the user logged into the electronic device (101) performs the first image processing of the first object m times using another external electronic device (e.g., PC) that the user logged into, the electronic device (101) may count the number of first image processings of the first object as (n+m) times.
[0091] According to one embodiment, the electronic device (101) may consider user data when determining a user's intention. User data is data related to a specific user using the electronic device (101) and may include various data such as, for example, relationship information, context data, preference data, and image processing performance pattern data. For example, if the user performs image processing on a specific person object (e.g., object deletion), the electronic device (101) may determine the user's intention as batch image processing on the person object (e.g., deletion of the person object from all images containing the person object) based on relationship information (e.g., information that the user recently deleted a contact corresponding to the person).
[0092] According to one embodiment, the operation of the electronic device (101) described above determining the user's intention can be performed using an artificial intelligence model (e.g., the generative AI model (230) of FIG. 2). For example, the electronic device (101) can input information including the user's repetitive image processing history into an artificial intelligence model stored in the memory (130) of the electronic device (101) or on an external server, and obtain the user's intention determined and output by the artificial intelligence model.
[0093] According to one embodiment, the electronic device (101) may display a user interface (UI) through a display that is related to checking whether image processing corresponding to a user intent determined by the electronic device (101) is performed. For example, the electronic device (101) may provide information to the user through a first popup (431) such as, “Saving is complete. Repetitive erasing editing (first type of image processing) has been performed. Would you like to erase A (first object) from the selected photo (first type of image processing)?” The “selected photo” in the above example may refer to, for example, an image to which image processing according to the user intent determined by the electronic device (101) is to be applied, or an image selected based on user input from among the images determined by the electronic device (100) (target image, described later in FIG. 7).
[0094] According to one embodiment, the electronic device (101) may perform a first type of image processing on a first object in at least one image based on user input (e.g., touch input) for a “Yes” button (431a) of a first popup (431). According to one embodiment, the electronic device (101) may not perform image processing based on user input for a “No” button (431b) of a first popup (431).
[0095] FIGS. 5a, FIGS. 5b and FIGS. 6 illustrate a user interface displayed by an electronic device (e.g., the electronic device (101) of FIG. 1) according to embodiments of the present disclosure.
[0096] According to one embodiment, the electronic device (101) may determine that the user's intention includes a plurality of intentions corresponding to the performance of a plurality of different image processing (image processing for different objects or image processing of different types), or at least some of the plurality of intentions. For example, the electronic device (101) may determine the user's intention as a first intention to perform a first type of image processing on a first object and / or a second intention to perform a second type of image processing on a second object. For example, the electronic device (101) may determine the user's intention as either a first intention to perform a first type of image processing on a first object or a second intention to perform a second type of image processing on a second object. The first intention and the second intention may have the same or different priority, for example.
[0097] According to one embodiment, priority may be determined based on user data (e.g., relationship information) and / or the number of times image processing corresponding to each intention is performed. For example, the electronic device (101) may determine user intentions and priority between user intentions based on the number of object-specific image processings for images including person objects and / or user data for each object (e.g., relationship information with the user). For example, if (i) object deletion for the third object and the fourth object is performed in a first image including the first object, the third object, and the fourth object, (ii) object deletion for the third object and the fifth object is performed in a second image including the second object, the third object, and the fifth object, and (iii) object deletion for the fourth object and the sixth object is performed in a third image including the first object, the second object, the fourth object, and the sixth object, the electronic device (101) can determine the user's intention as follows: first priority “leaving (not deleting) the first object and the second object”, second priority “deleting the third object” and “deleting the fourth object”, third priority “deleting the fifth object” and “deleting the sixth object”.
[0098] According to one embodiment, in the example described above, the electronic device (101) can determine the user's intention by using information on the user's relationship with objects related to the user's image processing history (e.g., a first object and a second object that were not deleted). For example, if the electronic device (101) identifies that the first object and the second object are related to the user as family, the electronic device can determine the user's first priority intention as "leaving (not deleting) the person object that is related as family." For example, the electronic device (101) may not delete the seventh object in an image where a new seventh object identified as family is identified.
[0099] According to one embodiment, the electronic device (101) may determine the user's intention as a plurality of intentions based on the user repeatedly performing different image processing. For example, if the number of times a first type of image processing for a first object is greater than or equal to a first threshold and the number of times a second type of image processing for a second object is greater than or equal to a second threshold, the electronic device (101) may determine the user's intention as a first intention to perform first type of image processing for the first object and a second intention to perform second type of image processing for the second object.
[0100] According to one embodiment, if a user repeatedly performs the same image processing but the same image processing can be interpreted as multiple intentions, the electronic device (101) may determine the user intention as some of the multiple intentions (e.g., any one of the multiple intentions). For example, if a user repeatedly deletes B and C from images containing three objects (A, B, C), the electronic device (101) may identify the user's intention as either "delete objects B and C" or "leave A and delete the remaining objects." In such a case, the electronic device (101) may output a user interface for user confirmation and determine the user intention based on user input through the user interface.
[0101] Referring to FIGS. 5a through 6, according to one embodiment, an electronic device (101) can identify that image processing to delete A is repeatedly performed in images containing objects A and B. For example, it can identify that the number of times image processing to delete A is performed in images containing A and B is greater than a threshold. In this case, the electronic device (101) can determine that the user's intention is either "delete object A" or "leave B and delete the remaining objects."
[0102] According to one embodiment, the threshold value compared with the number of image processing steps, which serves as the basis for determining user intent, may vary depending on several factors. For example, the threshold value may be determined based on at least one of a predetermined setting, the type of image processing performed, characteristics of the object to be processed, user data (e.g., relationship information between the user and the person object to be processed), or user input. The threshold value may be set to a first threshold value (e.g., 3 times) for a first type of image processing (e.g., object deletion) and a second threshold value (e.g., 5 times) for a second type of image processing (e.g., object color correction).
[0103] FIGS. 5a to 6 describes an example in which the electronic device (101) determines the user's intention as either "delete object A" or "delete the remaining objects while leaving B," but this is merely an example, and if the electronic device (101) determines the user's intention as part of a single or multiple types of image processing for a single or multiple objects, the description in FIGS. 5a to 6 described below may be applied in the same or similar way.
[0104] According to one embodiment, if the electronic device (101) determines that the user intent is either “delete object A” or “delete remaining objects while leaving B,” it may display at least one user interface to confirm the user intent (i.e., to confirm which of the two image processing methods to perform) between “delete object A” and “delete remaining objects while leaving B.” For example, the electronic device (101) may display the second popup (510) of FIG. 5a and / or the third popup (520) of FIG. 5b.
[0105] Referring to FIG. 5a, according to one embodiment, an electronic device (101) may display a second popup (510) through a display (e.g., the display module (160) of FIG. 1) to confirm a user intent. The second popup (510) may include a user interface related to confirming whether to perform object deletion for A in at least one image, for example. The electronic device (101) may perform object deletion for A in at least one image based on identifying user input to the “Yes” button (510a) of the second popup (510).
[0106] Referring to FIG. 5b, the electronic device (101) may display a third pop-up (520) through a display to confirm user intent. The third pop-up (520) may include a user interface related to confirming whether to perform object deletion for the remaining objects while leaving B in at least one image, for example. The electronic device may perform object deletion for the remaining objects excluding B in at least one image.
[0107] According to one embodiment, the electronic device (101) may first display a second popup (510) and, based on user input for the “No” button (510b) of the second popup (510), display a third popup (520). According to one embodiment, the electronic device (101) may first display a third popup (520) and, based on user input for the “No” button (520b) of the third popup (520), display a second popup (510).
[0108] According to one embodiment, which of the second popup (510) and the third popup (520) is displayed first may be determined, for example, by priority. For example, if the electronic device (101) determines that the priority of the “delete object A” intention is higher than the priority of the “delete remaining objects while leaving B” intention, the electronic device (101) may display the second popup (510) first and display the third popup (520) based on user input for the “No” button (510b) of the second popup (510).
[0109] According to one embodiment, the electronic device (101) may delete information that serves as the basis for determining the user intention based on user input that the determined user intention is incorrect (e.g., input via the “No” button (510b, 520b) of FIG. 5a and FIG. 5b). For example, if the electronic device (101) determines the user intention as the first intention based on the fact that the number of first type image processing counts for the first object exceeds a threshold value, but receives user input that the first intention is incorrect, the electronic device (101) may initialize the number of first type image processing counts for the first object that was being counted to 0.
[0110] Referring to FIG. 6, according to one embodiment, an electronic device (101) may display a fourth popup (610) to confirm a user's intention. The fourth popup (610) may provide information that allows the user to make a selection by collectively including information corresponding to a plurality of user intentions determined to have the same or different priority. For example, the text “Shall we ‘delete’ all other people except B?” in the fourth popup (610) may correspond to a user intention determined to have the first priority, and the text “Shall we ‘delete’ A?” may correspond to a user intention determined to have the same or different priority as the first priority.
[0111] According to one embodiment, the fourth popup (610) may include information that allows the user to select either “delete object A” or “delete the remaining objects while leaving B.” For example, as illustrated, the fourth popup (610) may include a “Yes” button (610a) under the text “Shall we ‘delete’ all other people except B?” For example, as illustrated, the fourth popup (610) may include a user interface area (610b) with the text “Shall we ‘delete’ A?”
[0112] According to one embodiment, the electronic device (101) can perform deletion of the remaining objects excluding B in at least one image based on identifying user input through the “Yes” button (610a) of the fourth popup (610). According to one embodiment, the electronic device (101) can perform deletion of the object A in at least one image based on identifying user input through the user interface area (610b) of the fourth popup (610) where the text “Would you like to ‘delete’ A?” is written.
[0113] FIG. 7 illustrates a screen displayed on a display (e.g., a display module (160) of FIG. 1) of an electronic device (e.g., an electronic device (101) of FIG. 1) according to one embodiment of the present disclosure.
[0114] Referring to FIG. 7, according to one embodiment, an electronic device (101) may display a first screen (710) through a display before performing image processing corresponding to the determined user intent after determining (or estimating) the user intent (e.g., after determining the user intent based on the number of times the first type of image processing for the first object exceeds a threshold). The first screen (710) may be a screen for the electronic device (101) to determine at least one target image (or candidate image) to which a specific type of image processing for a specific object will be applied. For example, it may be a screen for the electronic device (101) to determine a target image to which the first type of image processing for the first object will be applied. In this case, the target image may be determined, for example, from at least some of the images that include the first object and have not performed the first type of image processing for the first object.
[0115] According to one embodiment, the target image may be determined from among the remaining images associated with the image on which the user has previously performed image processing. For example, if the user has performed a first type of image processing on some of a plurality of images shared or received through a messenger application, the electronic device (101) may determine the remainder of the plurality of images as the target image and provide a user interface for determining whether to perform or will perform the first type of image processing on the target image. If a user interface for determining whether to perform the first type of image processing on the target image is provided, the electronic device (101) may perform the first type of image processing on the target image based on user input through the provided user interface.
[0116] According to one embodiment, the target image may be determined based on user input. For example, the target image may be determined as images corresponding to the user's selection. The user's selection may include, in addition to the selection of individual images, the collective selection of images corresponding to specific characteristics (e.g., specific metadata). For example, the target image may be determined based on user input that collectively selects images of specific characteristics, such as “images saved on a specific date,” “images saved in a specific folder of a gallery,” “all images in a gallery,” or “images containing a specific person.” According to one embodiment, the electronic device (101) may provide a query or selection option to induce user input.
[0117] According to one embodiment, the first screen (710) may include first information (e.g., text information) (711) for guiding the user to interact, at least one image (712) that may be a target image, and / or a “confirm” button (713). To induce appropriate interaction with the user, the electronic device (101) may provide the first information (711) by including information about the image processing to be performed (e.g., an object to be edited, a type of image processing to be performed).
[0118] According to one embodiment, at least one image (712) that may be a target image may be determined based on a determined user intent (e.g., a user intent determined based on the number of times the same type of image processing is performed on the same object exceeding a threshold). For example, the electronic device (101) may estimate a user intent to perform the same image processing based on identifying that the user has repeated the first type of image processing on the first object more than a threshold, and may determine 'all images containing the first object for which the user has not previously performed the first type of image processing' as at least one image (712) that may be a target image.
[0119] According to one embodiment, the first information (711) may include information that guides the determination of a target image among at least one image (712) that can be a target image. For example, the first information (711) may include text information such as “Please uncheck unwanted photos” or “Please check desired photos”.
[0120] According to one embodiment, the first information (711) may include information for verifying a user intent determined by the electronic device (101). For example, the first information (711) may include information about the target object and / or type of image processing to be performed based on the user intent determined by the electronic device (101), and may be provided with a user interface that allows the user to approve or reject the information.
[0121] According to one embodiment, the electronic device (101) can perform image processing (e.g., image processing of a first type for a first object) on a selected image based on identifying user input through a “confirm” button (713) while at least some of at least one image (712) that can be a target image is selected as the target image.
[0122] According to one embodiment, the electronic device (101) may display a second screen (720) including an image obtained by performing image processing on a target image after performing image processing. The second screen (720) may be a screen for determining an image to be stored among the images obtained by performing image processing on a target image.
[0123] According to one embodiment, the second screen (720) may include second information (721) and a “confirm” button (722) to guide the user to interact. To induce appropriate interaction from the user, the electronic device (101) may provide the second information (721) by including information about the image processing performed (e.g., an object to which editing has been applied, the type of image processing performed).
[0124] According to one embodiment, the electronic device (101) can store the selected image in memory (e.g., memory (130) of FIG. 1) and / or an external server based on identifying user input through a “confirm” button (722) while at least a portion of the images obtained by performing image processing on a target image is selected as an image to be stored.
[0125] According to one embodiment, in addition to user input, the target image may be determined based on various factors, such as the characteristics of the object to be processed, the type of image processing to be performed, and / or the characteristics of an image on which the processing has previously been performed. For example, in the case of object deletion for a specific person object (e.g., an object corresponding to a person deleted from recent contacts), all stored images may be determined as the target image. For example, the target image may be determined differently for object deletion and object face correction for a specific person object. For example, in the case of face shape adjustment for a specific person object (e.g., jawline retouching), an image containing the object taken on a specific date may be determined as the target image. For example, if a user repeatedly performs the same type of image processing on the same object in images taken on a specific date and at a specific location, the photos taken on that date and / or location may be determined as the target image. For example, if a user repeatedly performs the same type of image processing on the same object in images of a graduation ceremony, the graduation ceremony image may be determined as the target image. For example, if the same object and / or the same type of image processing is repeatedly performed on images assigned specific tag information, the remaining images assigned the same specific tag may be selected as target images.
[0126] According to one embodiment, the electronic device (101) may provide information regarding the determination criteria for the determined target image. For example, the electronic device (101) may display information such as “Batch editing target: Image taken on OO / OO / OO” through a display.
[0127] FIG. 8 is a drawing for explaining the operation of an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment of the present disclosure.
[0128] In operation 810, the electronic device (101) may perform a first type of image processing on a first object included in at least one first image, based on at least one first request. The first object may include various objects, such as, for example, people, animals, and objects. The first request may be, for example, a request via user input (e.g., an image editing request through an image editing application) to perform the first type of image processing on the first object.
[0129] In operation 820, the electronic device (101) can identify whether a first condition related to a first request has been satisfied. The first condition may be, for example, a condition for determining a user intent related to a first type of image processing for a first object. The first condition may be, for example, a condition that the number of image processings performed based on the first request is greater than or equal to a threshold. For example, the first condition may include a condition that the number of first type of image processings for a first object included in a plurality of images exceeds a threshold.
[0130] In operation 830, the electronic device (101) may perform a first type of image processing for a first object in at least one first target image based on identifying that a first condition is satisfied, thereby obtaining at least one first edited image corresponding to each of the first target images. The first target image may be, for example, at least some of the images containing the first object that are not included in the first image in operation 810. The first target image may be, for example, at least some of the images containing the first object in which the first type of image processing for the first object has not been performed by user input.
[0131] According to one embodiment, by having the electronic device (101) perform a series of operations of FIG. 8, the electronic device (101) can identify the user's intention and perform repetitive image processing tasks in batches, thereby providing a fast and convenient user experience.
[0132] According to one embodiment, the electronic device (101) can verify user intent through a user query between operation 820 and operation 830. For example, after operation 820, the electronic device (101) can provide first information related to at least one of the first object, the first type, or the first target image (e.g., provided through a user interface) and perform operation 830 based on identifying user verification related to the first information.
[0133] According to one embodiment, when the electronic device (101) performs image processing in operation 830, it may apply different degrees of image processing depending on the characteristics of the first object (e.g., size, position, positional relationship with other objects) and / or the characteristics of the first target image (e.g., lighting, shadows). For example, the electronic device (101) may determine the extent to which image processing of the 'jawline reduction correction' type is performed (e.g., how much the jawline is reduced) based on the size of the person object and / or the lighting of the target image.
[0134] According to one embodiment, in the case of image processing to delete a specific person object, if there are no other objects other than the person object in the target image, the electronic device (101) may delete the target image instead of performing object deletion image processing.
[0135] According to one embodiment, when saving the first edited image, the electronic device (101) may retain the first target image corresponding to each original of the first edited image without deleting it. The electronic device (101) may, for example, display the first edited image through a display after operation 830, and then display the first target image based on user input (e.g., user input via the “View Original” user interface button, long press on the first edited image).
[0136] According to one embodiment, when a first edited image is obtained while the first target image is maintained, the electronic device (101) may provide information (e.g., an icon, text) indicating that the image was automatically edited in batches when displaying the first edited image. The information may be set to be provided only, for example, until a predetermined number of viewings or for a predetermined period. The information may include, for example, an element (e.g., a “Restore Original” user interface button) that allows the first target image to be restored through user input. The first target image and similar images restored based on user input may be excluded from the target image when the electronic device (101) performs the same batch image processing next.
[0137] According to one embodiment, when storing the first edited image, the electronic device (101) may delete the first target image corresponding to each original of the first edited image at all times or when a predetermined deletion criterion is met. The deletion criterion may be a criterion determined based on the type of image processing and / or the characteristics of the object being processed. For example, the electronic device (101) may delete the first target image if the image processing performed in operation 830 is of the 'object deletion' type for a 'person object that is disconnected from the user (e.g., recently deleted contact).
[0138] FIG. 9 is a drawing for explaining the operation of an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment of the present disclosure.
[0139] Operations 910 and 940 of FIG. 9 may be the same as operations 810 and 830 of FIG. 8. Therefore, a detailed description of operations 910 and 940 in FIG. 9 is omitted.
[0140] In operation 920, the electronic device (101) may count a first number of times that image processing based on a first request is performed. The first request may be, for example, the first request described in operation 810 of FIG. 8. Image processing based on the first request may include, for example, a specific type of image processing (e.g., object deletion) for a specific object (e.g., a specific person) included in a plurality of images.
[0141] In operation 930, the electronic device (101) can identify whether the first number is greater than or equal to the first threshold by comparing the first number with the first threshold. The first threshold is a natural number value, and may be a fixed value or a variable value.
[0142] According to one embodiment, the first threshold value may be determined based on the characteristics of the first object and / or the first type. For example, if the first object is a specific person object and relationship information between the first object and the user exists, the first threshold value may be determined based on the relationship information between the first object and the user. For example, if the first object is identified as a person recently deleted from contacts and the first type is image processing for deleting the object, the electronic device (101) may set the first threshold value to a small value. For example, the electronic device (100) may determine the user's intention for batch image processing even if the user performs image processing for deleting the person object with which the relationship has been severed only a few times.
[0143] According to one embodiment, the electronic device (101) can perform operation 940 based on identifying that the first number is greater than or equal to the first threshold.
[0144] Type 1 Type 2 Type 3 Type 4 Type 1 Object 0020 Type 2 Object 0020 Type 3 Object 1000 Type 4 Object 2000
[0145] Table 1 is a table regarding the number of image processing cycles for multiple types of multiple objects.
[0146] Referring to Table 1, according to one embodiment, an electronic device (101) can count the number of times each type of image processing is performed for each object from images including a first object, a second object, a third object and / or a fourth object. The electronic device (101) can compare the counted number with a threshold. The first type, second type, third type and fourth type are different image processing types for a specific object, for example, the first type may be object deletion, the second type may be object mosaic, the third type may be object color correction, and the fourth type may be object preservation (deleting other objects and not deleting the object itself).
[0147] According to one embodiment, based on identifying that the number of image processing operations is greater than or equal to a threshold, as in operations 930 and 830 of FIG. 9, the electronic device (101) can perform the same type of image processing on the same object in at least one target image. The threshold may vary, for example, depending on the type of image processing, but in Table 1, it is assumed that the threshold for all image processing is 2.
[0148] According to one embodiment, the electronic device (101) may perform a third type of image processing for a first object on at least one target image containing the first object, based on identifying that the number of third type of image processing 2 for the first object is greater than or equal to a threshold of 2. This may be the same for third type of image processing for a second object and first type of image processing for a fourth object.
[0149] According to one embodiment, the electronic device (101) may not perform the first type of image processing for the third object based on identifying that the number of first type of image processing operations for the third object is less than a threshold of 2.
[0150] According to one embodiment, the electronic device (101) can determine the priority of each type of image processing for each object based on the number of times it is performed. For example, the electronic device (101) may determine the third type of image processing for the first object, the third type of image processing for the second object, and the first type of image processing for the fourth object as the first priority, and the first type of image processing for the third object as the second priority.
[0151] According to one embodiment, similar to FIGS. 5a and 5b, the electronic device (101) may first display a popup for confirming a user intent regarding a first-priority image processing. Based on identifying that the first-priority image processing is not a user intent (e.g., identifying user input via the “No” button (510) in FIG. 5a), the electronic device (101) may display a popup for confirming a user intent regarding a second-priority image processing.
[0152] According to one embodiment, the electronic device (101) may display a popup to check user intents regarding first-priority and second-priority image processing at once. For example, the electronic device (101) may display a popup including a query regarding first-priority image processing and a query regarding second-priority image processing, similar to FIG. 7.
[0153] FIG. 10 is a drawing for explaining relationship information according to one embodiment of the present disclosure.
[0154] According to one embodiment, an electronic device (e.g., the electronic device (101) of FIG. 1) may obtain relationship information, which is information about the relationship between a user and a person. The relationship information may include information about the relationship with people and / or relationship patterns, obtained based on information stored through user input (e.g., contact group information) and / or other user data (e.g., user interaction data, SNS usage data, calendar data). The relationship information may include classification information of person groups (e.g., family, work, friends) obtained based on text stored in contacts, for example. The relationship information may include, for example, the frequency, type, and / or timing of interactions (e.g., messages, phone calls) between a specific person and the user. The relationship information may include, for example, the duration of the relationship between a specific person and the user (e.g., contact storage period, SNS follow period) and whether the relationship has been severed (e.g., whether the contact has been deleted, whether the SNS has been unfollowed).
[0155] According to one embodiment, when a person object is included in an image, the electronic device (101) can identify the person object. For example, the electronic device (101) can identify that the person object corresponds to a specific person stored in contacts.
[0156] Referring to FIG. 10, according to one embodiment, an electronic device (101) can identify at least one person object included in the images. For example, the electronic device (101) can identify person objects A, B, C, D, and E, each corresponding to a different person.
[0157] According to one embodiment, the electronic device (101) can classify person objects into groups based on relationship information for each person. For example, in FIG. 10, the electronic device (101) can classify A and B into a family group (1010), C and D into a friends group (1020), and E into a workplace group (1030).
[0158] Delete Object Object Mosaic Object Color Correction Object Preservation A0010B0010C1000D2000E1000Family0020
[0159] Table 2 is a table showing the number of different types of image processing for each person object.
[0160] Referring to Table 2, the number of object color corrections for A may be 1, the number of object color corrections for B may be 1, the number of object deletions for C may be 1, the number of object deletions for D may be 2, and the number of object deletions for E may be 1. As explained in Table 1, if counting is done individually without using relationship information for person objects, for example, the deletion of the object for D may be determined to be the first priority user intent.
[0161] According to one embodiment, the priority of user intent determined by the electronic device (101) may change when using relationship information. For example, if individual person objects are grouped and counted by person group as shown in FIG. 10, the number of object color corrections for the family group (1010) may be 2 times, as shown in the bottom row of Table 2. Although not listed in Table 2, the number of object deletions for the friends group (1020) may be 3 times, and the number of object deletions for the workplace group may be 1 time. In this case, if user intent is determined by person group classified based on relationship information, object deletion for the friends group (1020) may be determined as 1st priority, object color correction for the family group (1010) as 2nd priority, and object deletion for the workplace group (1030) as 3rd priority. That is, when counting the number of image processings per simple object without using relationship information, the first priority user intention is determined to be “deletion of object for D,” and when using relationship information, the first priority user intention can be determined to be “deletion of object for friend group (1020).”
[0162] FIG. 11 is a drawing for explaining the operation of an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment of the present disclosure.
[0163] In operation 1110, according to one embodiment, an electronic device (101) can identify person objects included in at least one image. The operation of identifying person objects may include a process of detecting and distinguishing each person object, and at least some of the identified person objects may correspond to a specific person. For example, at least some of the person objects identified and classified by the electronic device (101) may be mapped to correspond to a specific person based on pre-stored specific person data (e.g., contact data containing a person image) and / or user input (e.g., user tagging).
[0164] In operation 1120, according to one embodiment, the electronic device (101) can obtain relationship information related to person objects. The relationship information related to person objects may include, for example, relationship information between a specific person and a user corresponding to each person object. The relationship information may include, for example, the relationship information described in FIG. 10.
[0165] In operation 1130, according to one embodiment, the electronic device (101) can classify person objects into groups based on relationship information. For example, as shown in FIG. 10, the electronic device (101) can classify person objects into a “family” group, a “friends” group, and / or a “work” group.
[0166] FIG. 12 is a drawing for explaining image processing according to one embodiment of the present disclosure.
[0167] FIG. 13 illustrates a screen displayed by an electronic device according to one embodiment of the present disclosure.
[0168] According to one embodiment, an electronic device (101) performs image processing according to an image processing request for a plurality of objects (e.g., a first type of image processing request for a first object and a second type of image processing request for a second object), and when the image processing request satisfies a condition (e.g., a condition in which the number of first type of image processing for a first object is greater than or equal to a first threshold and the number of second type of image processing for a second object is greater than or equal to a second threshold), image processing for each of the plurality of objects can be performed on a single target image.
[0169] Referring to FIG. 12, according to one embodiment, an electronic device (101) can perform image processing N times to add a scarf to the neck position of a first object (1211) included in at least one first image (1210). The electronic device (101) can perform image processing to add a scarf to the neck position of a first object (1211) in at least one first target image based on identifying that the number of image processings to add a scarf to the neck position of the first object (1211) is greater than or equal to a first threshold.
[0170] According to one embodiment, an electronic device (101) may perform image processing M times to add a hat to the head position of a second object (1222) included in the second image (1220) in at least one second image (1220). The electronic device (101) may perform image processing to add a hat to the head position of a second object (1222) in at least one second target image based on identifying that the number of image processings to add a hat to the head position of the second object (1222) is greater than or equal to a second threshold.
[0171] According to one embodiment, the first threshold value may be N or less and the second threshold value may be M or less. For example, the electronic device (101) may automatically perform corresponding image processing on the first target image and the second target image, wherein there may be a common target image (1230) that is included in both the first target image and the second target image. For example, the common target image (1230) may include a first object (1231) and a second object (1232).
[0172] According to one embodiment, the electronic device (101) can perform image processing to add a scarf to the neck position of a first object (1211) included in a first image (1210) and image processing to add a hat to the head position of a second object (1222) included in a second image (1220), based on identifying that the number of image processings to add a scarf to the neck position of a first object (1231) included in a first image (1210) is greater than or equal to a first threshold, and the number of image processings to add a hat to the head position of a second object (1232) included in a second image (1220).
[0173] Referring to FIG. 13, according to one embodiment, an electronic device (101) may display a first screen (1310) containing an image including an object A (1311). The electronic device (101) may perform image processing on the object A (1311) based on user input (e.g., touch input) through a “create” user interface button on the first screen (1310), for example. For example, the electronic device (101) may repeatedly perform image processing to delete the remaining person objects excluding the object A (1311).
[0174] According to one embodiment, the electronic device (101) may display a second screen (1320) containing an image including a B object (1322). The electronic device (101) may perform image processing on the B object (1322) based on user input through a “create” user interface button on the second screen (1320), for example. For example, the electronic device (101) may repeatedly perform image processing to delete the remaining person objects excluding the B object (1322).
[0175] According to one embodiment, the electronic device (101) can perform image processing to delete the remaining person objects while leaving only the A object (1311) and the B object (1322) in a common target image containing the A object (1311) and the B object (1322), based on identifying that the number of image processings to delete the remaining person objects excluding the A object (1311) and the B object (1322) are each greater than or equal to a threshold value. Prior to this, the electronic device (101) can display a third screen (1330) related to user verification.
[0176] According to one embodiment, an electronic device (101) may display a common target image including object A (1311) and object B (1322), first information (1333) for receiving confirmation from the user whether to delete the remaining person objects while leaving only object A (1311) and object B (1322), and second information (1334) for receiving confirmation from the user whether there are any additional objects among the remaining person objects that the user wants to keep. The electronic device (101) may perform image processing to delete the remaining person objects while leaving only object A (1311) and object B (1322) in the common target image, for example, based on the identification of user input through the “Yes” button of the first information (1333) in the absence of user input through the second information (1334).
[0177] FIG. 14 is a drawing for explaining the operation of an electronic device according to one embodiment of the present disclosure.
[0178] FIGS. 15 and 16 are drawings for illustrating image processing and the context of an image according to one embodiment of the present disclosure.
[0179] Referring to FIG. 14, in operation 1410, the electronic device (101) applies image processing to a first target image (e.g., the first target image of FIG. 8) to obtain a first edited image (e.g., the first edited image of FIG. 8) and can identify whether the edited image maintains contextual consistency. The context of the image may refer to the overall visual and logical context of the image, including elements such as, for example, the type of scene the image represents, the relationships of objects included in the image, and their size and / or location. That the edited image maintains contextual consistency may mean, for example, that the image has not lost visual and / or logical consistency before and after editing.
[0180] According to one embodiment, visual consistency may mean that visual elements such as the arrangement, color, lighting, shadows, and / or perspective of elements within an image are maintained naturally. For example, visual consistency may be maintained when the lighting direction and color temperature match between the background and the object. For example, visual consistency may not be maintained if shadows remain after an object is edited and removed, or if the boundary lines of an object separated from the background appear unnatural.
[0181] According to one embodiment, logical consistency may mean that the relationships and arrangements between objects within an image are realistically valid and that the meaning conveyed by the image is maintained. For example, in an image of two people with their arms around each other's shoulders, logical consistency may be maintained if both people are retained. For example, in an image of two people with their arms around each other's shoulders, logical consistency may not be maintained if only one person is deleted and the other person's arm is left floating in the air.
[0182] According to one embodiment, whether an edited image maintains contextual consistency can be determined using an artificial intelligence model (e.g., the generative AI model (230) of FIG. 2). For example, an electronic device (101) can evaluate the visual consistency and / or logical consistency of an edited image using a machine learning-based image analysis model.
[0183] In operation 1420, the electronic device (101) can perform additional image processing or undo image processing that has already been performed based on identifying that the first edited image does not maintain context consistency.
[0184] Referring to FIG. 15, the electronic device (101) can perform image processing to delete one of the objects from a target image (1510) in which two person objects are linking arms. As a result of the image processing, for example, the electronic device (101) can obtain a first edited image (1521) in which the object is deleted while the remaining object is linking arms, or a second edited image (1522) in which the arm of the remaining object is deleted along with it.
[0185] According to one embodiment, the electronic device (101) may identify that the first edited image (1521) and / or the second edited image (1522) do not maintain contextual consistency. For example, the electronic device (101) may determine that the first edited image (1521) does not maintain logical consistency (e.g., when a person is in a shoulder-to-shoulder pose but there is no other person next to them) or determine that the second edited image (1522) does not maintain visual consistency (e.g., when a part of the person's body (e.g., arm) is unnaturally cut off).
[0186] According to one embodiment, the electronic device (101) may perform additional image processing based on identifying that the edited image does not maintain context consistency. For example, the electronic device (101) may perform additional image processing to naturally generate the awkward arm of the remaining person object, thereby obtaining a final edited image (1530).
[0187] Referring to FIG. 16, the electronic device (101) can obtain a third edited image (1610) by performing image processing to add a scarf (1612) to a target image containing a specific person object (1611). At this time, if the background of the third edited image (1610) represents hot weather (e.g., a summer beach background, a desert background), the electronic device (101) can identify that the third edited image (1610) does not maintain context consistency.
[0188] According to one embodiment, the electronic device (101) can obtain a final edited image (1620) by performing additional image processing to delete the scarf (1612) based on identifying that the third edited image (1610) does not maintain context consistency, or restore the original target image by reversing the image processing performed (image processing to add the scarf).
[0189] In operation 1430, the electronic device (101) may store the result of additional image processing or image processing reversal. For example, the electronic device (101) may store the final edited image resulting from performing additional image processing on the first edited image, or may reverse the image processing to delete the first edited image and restore the first target image to store it.
[0190] In operation 1440, the electronic device (101) may store the first edited image based on identifying that the first edited image maintains context consistency.
[0191] According to one embodiment, the electronic device (101) performs image processing on a target image to obtain an edited image, and then provides information (e.g., a query and / or a popup) to the user to perform additional image processing or to revert image processing if the edited image is awkward, and can apply additional image processing to the edited image or revert image processing based on user input.
[0192] According to one embodiment, the electronic device (101) may provide information to the user if it performs additional image processing or reverses image processing performed based on identifying that the edited image does not maintain context consistency. For example, in the example of FIG. 16, the electronic device (101) may provide information through a display such as “I deleted the scarf that does not match the summer background” or “I canceled the addition of the scarf because the added scarf does not match the summer background.”
[0193] FIG. 17 is a drawing for explaining an edited image and information related to the edited image according to one embodiment of the present disclosure.
[0194] The edited image (1710) of FIG. 17 may include, for example, a first edited image obtained as an electronic device (e.g., the electronic device (101) of FIG. 1) performs the operation of FIG. 8 or FIG. 9.
[0195] Referring to FIG. 17, according to one embodiment, an electronic device (101) may display information related to the edited image (1710) along with the edited image (1710) through a display (e.g., the display module (160) of FIG. 1). Information related to the edited image may include, for example, a date and time of shooting (or a date and time of saving), a filename, a storage path, a file size, a resolution, a number of pixels, hashtags, and / or usage history (1720). Usage history may include, for example, a history of transmitting the image or uploading it to a specific application (e.g., a social media application) and / or a cloud server.
[0196] According to one embodiment, some of the information related to the edited image (1710) may display the information of the original image corresponding to the edited image (1710) as is (e.g., in FIG. 8, the original image corresponding to the first edited image is the first target image). For example, the shooting date and time, hashtags previously entered by the user, and the history of use of the original image (1720) may remain unchanged before and after image processing and may be maintained identically in the original image and the edited image (1710).
[0197] According to one embodiment, the usage history (1720) displayed by the electronic device (101) together with the edited image (1710) may include a history of using the original image corresponding to the edited image (1720). For example, even if the edited image (1720) is obtained on February 1st by performing batch image processing, the usage history (1720) may include text information such as the previous usage history, such as “I uploaded the image on January 29th via Instagram,” or “I uploaded the image on January 30th to a cloud server.”
[0198] According to one embodiment, by providing a usage history (1720) to the user, the electronic device (101) can induce the user to decide whether to update the original image to the edited image (1710). For example, the user can view the usage history (1720) and update the original image stored on the cloud server to the edited image (1710), or delete the original image uploaded to SNS and upload the newly edited image (1710).
[0199] According to one embodiment, the electronic device (101) displays information (e.g., a user query popup) to confirm whether to update the original image to the edited image (1710) and may or may not update the original image to the edited image (1710) based on user input.
[0200] According to one embodiment, the electronic device (101) can automatically update the original image to the edited image (1710) without necessarily displaying the usage history (1720) or through user input. For example, the electronic device (101) can automatically delete the original image stored on a cloud server and upload the edited image (1710) to the storage location.
[0201] FIG. 18 is a diagram illustrating the operation of updating an image on a cloud server according to image processing in an electronic device according to one embodiment of the present disclosure.
[0202] The electronic device (1850) of FIG. 18 may be, for example, the electronic device (101) of FIG. 1.
[0203] Referring to FIG. 18, according to one embodiment, an electronic device (1850) may store a plurality of images, such as P1 and P2 (1851), in memory. The electronic device (1850) may obtain P2a by performing image processing on P2. For example, P2 (1851) may be a first target image of FIG. 8, and P2a may be a first edited image of FIG. 8 obtained by performing a first type of image processing on a first object on P2 (1851).
[0204] According to one embodiment, the electronic device (1850) can synchronize an image stored in memory to a cloud server. Synchronization may include the operation of automatically uploading the image to the cloud server based on a set synchronization condition (e.g., conditions related to network connection status, battery level, etc.) and / or the operation of manually uploading the image to the cloud server (e.g., based on user input). For example, an existing cloud server (1810) may store images such as P1 and P2 (1811) by synchronizing the image stored in the electronic device (1850).
[0205] According to one embodiment, after the electronic device (1850) obtains P 2a by performing image processing, it may transmit a modification notification to an existing cloud server (1810) that P 2 (1851) has been modified to P 2a and image data of the obtained P 2a. At this time, the electronic device (101) may also transmit a command to the existing cloud server (1810) to update P 2 (1811) to the modified image.
[0206] According to one embodiment, the existing cloud server (1810) may be configured to automatically update the image upon receiving only the image modification notification and / or the modified image data.
[0207] According to one embodiment, an update to P 2 (1811) is performed on the existing cloud server (1810), so that the updated cloud server (1820) may include P 2a (1821) instead of P 2. The existing cloud server (1810) and the updated cloud server (1820) are the same cloud server, but may represent the cloud server at a time before and after the update from P 2 to P 2a, respectively.
[0208] FIG. 19 is a drawing for explaining the operation of an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment of the present disclosure.
[0209] In operation 1910, the electronic device (101) may receive a prompt for image processing. The prompt may refer to input data that induces an artificial intelligence model for image processing (e.g., the generative artificial intelligence model (230) of FIG. 2) to generate and output a desired response. The prompt may be data in various formats, such as natural language text, code, or images, for example.
[0210] According to one embodiment, a prompt for image processing may include a prompt that performs a specific type of image processing on a specific object. The prompt may include a natural language prompt, for example, “Delete everyone else except me,” and may be entered by a user.
[0211] In operation 1920, the electronic device (101) can interpret the input prompt. The electronic device (101) can analyze the keywords and / or context contained in the prompt using, for example, a natural language processing (NLP) and / or command analysis model. For example, if the prompt “Leave only me and delete everyone else,” the electronic device (101) can recognize “the user’s intention to keep a specific object (an object identified as the same person as the user) and remove other objects.” Additionally, based on the structure and context of the prompt, the type of image processing (e.g., object removal, background change, style application, etc.) can be classified.
[0212] In operation 1930, the electronic device (101) can determine a target image to which image processing is to be applied. The electronic device (101) can determine an image explicitly selected by the user or an image automatically selected based on the content of a prompt as the target image. For example, if the user selects a specific image from the gallery and enters a prompt, that image may be set as the target image. Additionally, if a request is made to perform image processing on a specific person or a specific situation (e.g., an image taken during a specific period, an image taken at a specific location, etc.), the electronic device (101) can automatically select an appropriate target image based on metadata and image analysis.
[0213] In operation 1940, the electronic device (101) can perform image processing on a target image based on the result of interpreting a prompt. The electronic device (101) can perform image processing based on the command interpreted from the prompt using a generative artificial intelligence model. For example, the electronic device (101) can perform image processing to remove the background from the target image while leaving only a specific person.
[0214] FIG. 20 is a block diagram of an electronic device according to one embodiment of the present disclosure.
[0215] The electronic device (2000) of FIG. 20 may be, for example, an electronic device that illustrates the electronic device (101) of FIG. 1 from a different perspective. An operation performed using the configuration of the electronic device (2000) may be performed using the configuration of the electronic device (101) of FIG. 1. For example, an operation performed using the configuration of the electronic device (2000) may be implemented by executing at least one instruction stored in the memory (130) of the electronic device (101) of FIG. 1 in the processor (120).
[0216] According to one embodiment, the electronic device (2000) may include an image processing module (2010), a user intent identification module (2020), an AI module (2030), a relationship information management module (2040), and / or a user intent verification module (2050).
[0217] According to one embodiment, the image processing module (2010) can perform image processing as intended by the user based on user input. For example, the electronic device (2000) can perform object segmentation based on user input (e.g., long press) regarding a specific object included in the image, or perform image editing based on user input through a specific application (e.g., Photoshop).
[0218] According to one embodiment, the image processing module (2010) can estimate the user's intent and perform batch image processing on at least one target image according to the estimated user intent. For example, if the user intent is estimated to be “image processing of a first type for a first object,” the image processing module (2010) can perform image processing of a first type for a first object on at least one image containing the first object.
[0219] According to one embodiment, the image processing module (2010) may include an object detection module (2011) and / or a batch processing module (2012).
[0220] According to one embodiment, the object detection module (2011) receives data of an image and an object and can detect the object in the image. The object detection module (2011) can obtain corresponding mask data by dividing the detected object.
[0221] According to one embodiment, the batch processing module (2012) may receive an image, mask data for an object included in the image obtained from the object detection module (2011), and / or a specific type of image processing method. The batch processing module (2012) may apply the received type of image processing to the portion of the received image that corresponds to the received mask data.
[0222] According to one embodiment, the user intent identification module (2020) can estimate the user's intent to perform a specific edit, that is, a specific type of image processing for a specific object, on at least one image in batches. The user intent identification module (2020) may, for example, utilize relationship information received from the relationship information management module (2040) (described later) to identify the user's intent. The user intent identification unit may also estimate the user's intent based on information input through an AI model.
[0223] According to one embodiment, there may be multiple user intentions estimated by the user intention identification module (2020), and each of the multiple intentions may have the same or different priority.
[0224] According to one embodiment, the user intent identification module (2020) may include a count recording module (2021), a threshold setting module (2022), and / or an image processing type and target object recognition module (2023).
[0225] According to one embodiment, the count recording module (2021) can count and / or record the number of times a specific type of image processing is performed on a specific object in at least one image. For example, the count recording module (2021) can separately count the number of times a first type of image processing is performed on a first object and the number of times a second type of image processing is performed on a second object. The counted counts can be used to compare with a predetermined value (e.g., a threshold value that serves as a condition for triggering user intent estimation). When counting image processing for a person object, multiple person objects can be grouped together and counted using relationship information regarding the person object.
[0226] According to one embodiment, the threshold setting module (2022) can determine a threshold that serves as a condition for triggering user intent estimation. If the number of counts recorded by the count recording module (2021) is greater than or equal to the threshold set by the threshold setting module (2022), the user intent identification module (2020) can estimate that it is the user's intention to perform the corresponding image processing in batches. In determining the threshold, for example, relationship information may be considered. For example, in the case of a person whose relationship has recently been severed, the threshold for object deletion for that person object may be set lower than before.
[0227] According to one embodiment, the image processing type and target object recognition module (2023) can analyze manually performed image processing to identify the target object in the image processing and what type of image processing was performed. Whenever the number of times a specific type of image processing is performed on a specific object is counted by the count recording module (2021), the image processing type and target object recognition module (2023) can rank the corresponding image processing according to the number of times.
[0228] According to one embodiment, the AI module (2030) may include an artificial intelligence system, such as the generative artificial intelligence system (200) of FIG. 2. The AI module (2030) may be implemented to run in an on-device environment within an electronic device (2000), for example, or may be implemented in a cloud-based manner on an external server connected via a network.
[0229] According to one embodiment, the relationship information management module (2040) may be a module that generates relationship information based on contact information (2045) stored in an electronic device (2000). The relationship information management module (2040) may, for example, use the contact information (2045) to define social relationships between a user and each person (a person corresponding to an identified person object) and transmit this to the user intent identification module (2020).
[0230] According to one embodiment, the user intent verification module (2050) can verify whether the user intent estimated by the user intent identification module (2020) is correct. For example, the user intent verification module (2050) can provide information (e.g., popup, query) for verifying the estimated user intent to the electronic device (2000) and identify whether the estimated user intent is correct based on the user response. If there are multiple estimated user intents, the user intent verification module (2050) can provide information for selecting at least some of the multiple intents.
[0231] FIG. 21 is a diagram illustrating the operation of an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment of the present disclosure. At least some of the operations included in FIG. 21 may be performed in a different order. At least some of the operations included in FIG. 21 may be performed simultaneously (e.g., in parallel). At least some of the operations included in FIG. 21 (e.g., operation 2150) may be omitted. Additional operations may be performed between at least some of the operations included in FIG. 21. At least some of the operations included in FIG. 21 may be performed on another electronic device (e.g., a server).
[0232] Referring to FIG. 21, in operation 2110, the electronic device (101) may receive a plurality of image processing requests corresponding respectively to a first plurality of images. The plurality of image processing requests may be, for example, image processing requests via user input. The plurality of image processing requests may be, for example, requests for a specific type of image processing (e.g., object deletion) for a specific object (e.g., a specific person).
[0233] According to one embodiment, a plurality of image processing requests may include a sequence of repetitive commands based on user input that are received continuously. The electronic device (101) may, for example, use an artificial intelligence model to generate a first plurality of edited images corresponding to each of the first plurality of images.
[0234] In operation 2120, the electronic device (101) can identify at least one target object and / or at least one image processing type corresponding to a plurality of image processing requests. For example, the electronic device (101) can identify the target objects and image processing types and count the number of image processings for each request by object and by image processing type, as in Table 1 or Table 2.
[0235] In operation 2130, the electronic device (101) can identify whether a plurality of image processing requests satisfy a specified condition. The specified condition may include, for example, a condition that the number of times image processing corresponding to at least one target object and at least one image processing type is performed based on the plurality of image processing requests is greater than or equal to a specified threshold. The specified condition (e.g., threshold) may be determined, for example, in correspondence with the characteristics of the target object and / or the image processing type. For example, the electronic device (101) may set the specified condition differently for a first image processing type and a second image processing type.
[0236] In operation 2140, the electronic device (101) can estimate a user intent corresponding to a plurality of image processing requests based on the plurality of image processing requests satisfying specified conditions. The operation of estimating a user intent can correspond, for example, to the operation of determining a user intent described in FIG. 4.
[0237] According to one embodiment, the electronic device (101) may estimate a user intent to perform a specific image processing on a specific target object based on the condition that the number of times image processing corresponding to a specific target object and a specific image processing type is performed is greater than or equal to a predetermined threshold. Meanwhile, if at least one other object to which image processing of a specific image processing type is not applied is included in at least some of the first plurality of images, the electronic device (101) may estimate a user intent to not perform image processing of a specific image processing type on the at least one other object.
[0238] In operation 2150, the electronic device (101) may determine a second plurality of images according to an estimated user intent. The second plurality of images may include, for example, target images (or candidate images) for performing image processing according to the user intent.
[0239] According to one embodiment, the second plurality of images may be determined based on user input among the plurality of images determined according to the user intention estimated by the electronic device (101). For example, the operation of determining the second plurality of images among the plurality of images may be the same as the operation of determining the target image described in FIG. 7. For example, the electronic device (101) may determine all or part of the plurality of images as the second plurality of images based on user input.
[0240] In operation 2160, the electronic device (101) may provide a query to verify the estimated user intent. For example, if the electronic device (101) estimates the user intent to be “delete person object A,” it may provide a query such as “Would you like to delete object A in bulk from the displayed images (second plurality of images)?” through a display. The query may be displayed, for example, through a user interface (e.g., a popup) containing information about the estimated user intent and / or the second plurality of images.
[0241] In operation 2170, the electronic device (101) may perform image processing corresponding to at least one target object and / or at least one image processing type in at least one image (the second at least one image) included in the second plurality of images based on a user response to a query. For example, the electronic device (101) may perform A object deletion processing in the second at least one image based on user input that (i) selects only some of the second plurality of images (the second at least one image) and (ii) approves the batch A object deletion image processing in response to the query “Would you like to delete A object in batches from the selected images (the second plurality of images)?”. The user response may include, for example, user input confirming that the second at least one image corresponds to the user intent, which is entered through a user interface corresponding to the user query (e.g., touch input via a button included in the query popup).
[0242] According to one embodiment, the electronic device (101) may perform image processing corresponding to at least one image processing type to different degrees depending on the scene displayed by the second at least one image. For example, the electronic device (101) may or may not perform image processing based on the context of the second at least one image.
[0243] According to one embodiment, the electronic device (101) may perform additional image processing based on the positional relationship between at least one target object and at least one other object when at least one other object is included in addition to at least one target object in a second at least one image. For example, as in FIG. 15, when deleting a target person object, if there is another person object adjacent to the target person object, the electronic device (101) may perform additional image processing to make the unnaturally edited body of the other person object look natural after deleting the target person object.
[0244] According to one embodiment, the electronic device (101) may perform image processing corresponding to at least one target object and / or at least one image processing type in a second at least one image to generate at least one edited image corresponding to the second at least one image. The at least one edited image may be generated, for example, using an artificial intelligence model (e.g., a generative artificial intelligence model for image generation).
[0245] According to one embodiment, the electronic device (101) may display the generated image through a display. At this time, the electronic device (101) may display information indicating that at least one edited image was generated based on the user response and / or the second at least one original image corresponding to each of the at least one edited image. For example, the electronic device (101) may intuitively convey to the user which original image and which automatic image processing were used to generate the edited image, thereby making subsequent processing (e.g., saving the edited image, canceling the edit, deleting the image, etc.) convenient for the user.
[0246] According to one embodiment, a plurality of image processing requests in operation 2110 are received through a first application (e.g., a gallery application), and at least one edited image described above can be displayed through the first application. At this time, if there is an image that has already been uploaded through a second application (e.g., an SNS application, a messenger application) among the second at least one original image, the electronic device (101) can replace the uploaded image with a corresponding edited image and upload it through the second application.
[0247] According to one embodiment, after generating at least one edited image, the electronic device (101) may determine whether to delete or retain a second at least one image, which is a corresponding original image, based on user input and / or the type of image processing performed. For example, the electronic device (101) may provide a query to the user asking whether to delete the original image, and may determine whether to delete it based on user input through the query. For example, the electronic device (101) may delete the original image if the type of image processing performed is object deletion, and retain the original image if the type of image processing performed is color correction.
[0248] An electronic device according to one embodiment of the present disclosure may include a memory comprising at least one storage medium for storing instructions; and at least one processor comprising a processing circuit.
[0249] According to one embodiment, when the instructions are executed individually or collectively by at least one processor of an electronic device, the electronic device may be caused to: receive a plurality of image processing requests corresponding to each of a first plurality of images; identify at least one target object and at least one image processing type corresponding to the plurality of image processing requests; estimate a user intent corresponding to the plurality of image processing requests based on the fact that the plurality of image processing requests satisfy a specified condition; determine a second plurality of images according to the estimated user intent; provide a query for verifying the estimated user intent; and, based on a user response to the query, perform image processing corresponding to the at least one target object and the at least one image processing type in at least one image included in the second plurality of images.
[0250] According to one embodiment, a plurality of image processing requests include a sequence of repetitive commands based on user input that are received sequentially, and the commands, when executed individually or collectively by at least one processor, may cause an electronic device to: generate a first plurality of edited images corresponding to each of the first plurality of images using an artificial intelligence model based on each of the plurality of editing requests.
[0251] According to one embodiment, the specified condition may include a condition in which the number of times image processing corresponding to at least one target object and at least one image processing type is performed based on a plurality of image processing requests is greater than or equal to a threshold value determined in relation to the at least one image processing type.
[0252] According to one embodiment, at least some of the first plurality of images further include at least one other object that does not correspond to the plurality of image processing requests, and the estimated user intent may include an intent not to process the at least one other object according to at least one image processing type.
[0253] According to one embodiment, at least one target object includes a first target object included in a first image among a first plurality of images and a second target object included in a second image among the first plurality of images, and at least one image processing type includes a first image processing type for the first image and a second image processing type for the second image, and instructions, when executed individually or collectively by at least one processor, may cause an electronic device to: generate an edited image corresponding to the third image based on the first image processing type and the second image processing type, when the third image included in the at least one image included in the second plurality of images includes the first target object and the second target object.
[0254] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: set different specified conditions for a first image processing type and a second image processing type.
[0255] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: determine a second plurality of images differently according to each of at least one image processing type.
[0256] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: perform image processing corresponding to at least one image processing type to a different degree in the at least one image, depending on the scene displayed by the at least one image included in the second plurality of images.
[0257] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may perform additional image processing according to the positional relationship between the at least one target object and the at least one other object, where the at least one image included in the second plurality of images includes at least one target object and at least one other object.
[0258] According to one embodiment, a query is displayed through a user interface that includes information about a user intent estimated by an electronic device and a second plurality of images, and a user response may include a user input confirming that at least one image included in the second plurality of images corresponds to the user intent.
[0259] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: generate at least one edited image corresponding to said at least one image by performing image processing on at least one image included in said second plurality of images using an artificial intelligence model, and to display said at least one edited image together with a first element indicating that said at least one edited image was generated based on said user response and a second element for displaying said at least one image which is the original corresponding to said at least one edited image.
[0260] According to one embodiment, a plurality of image processing requests are received through a first application, at least one edited image is displayed through the first application, and commands, when executed individually or collectively by at least one processor, may cause an electronic device to: replace and upload an image that has been uploaded through the second application among at least one image included in the second plurality of images with a corresponding edited image through the second application.
[0261] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may determine, based on at least one image processing type, whether to delete or retain at least one image included in a second plurality of images after generating at least one edited image.
[0262] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to: perform a first type of image processing on a first object included in at least one first image in at least one first image based on a first request, count a first number of times the image processing based on the first request was performed, and, based on identifying that the first number is greater than or equal to a first threshold, perform the first type of image processing on a first object included in at least one first target image in at least one first target image to obtain at least one first edited image corresponding to each of the at least one first target image.
[0263] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: perform a second type of image processing for a second object different from a first object in at least one second image based on a second request, wherein the second type is an image processing type that is the same or different from the first type, count a second number which is the number of times the image processing based on the second request has been performed, and, based on identifying that the first number is greater than or equal to the first threshold and the second number is greater than or equal to the second threshold, perform the first type of image processing for the first object included in the at least one second target image and the second type of image processing for the second object included in the second target image in at least one second target image, thereby obtaining at least one second edited image corresponding to each of the at least one second target image.
[0264] According to one embodiment, the threshold value may be determined based on at least one of the characteristics of the object or the type of image processing.
[0265] According to one embodiment, the object may include an object identified as the same person. According to one embodiment, the object may include an object identified as one of the people belonging to the same group.
[0266] According to one embodiment, the threshold value may be determined based on at least one of relationship information corresponding to an object or a type of image processing.
[0267] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: identify whether the first edited image maintains contextual consistency, and based on identifying that the first edited image does not maintain contextual consistency, cause additional image processing to be performed or cause the performed image processing to be undoed.
[0268] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: store the usage history of each of the first target images in the memory, wherein the usage history includes an upload history, and based on the usage history of each of the at least one first target image, update at least one uploaded image among the at least one first target image to an image corresponding to at least one uploaded image among the first edited images.
[0269] A method of an electronic device according to one embodiment of the present disclosure may include: an operation of performing a first type of image processing on a first object included in at least one first image based on a first request; an operation of counting a first number of times the image processing based on the first request has been performed; and an operation of obtaining at least one first edited image corresponding to each of the at least one first target image by performing the first type of image processing on a first object included in at least one first target image based on identifying that the first number is greater than or equal to a first threshold value.
[0270] In a non-transient computer-readable recording medium storing instructions according to one embodiment of the present disclosure, the instructions may cause the at least one processor to perform set operations when executed by at least one processor. The set operations may include: an operation of performing a first type of image processing on a first object included in at least one first image in at least one first image based on a first request; an operation of counting a first number of times the image processing based on the first request has been performed; and an operation of obtaining at least one first edited image corresponding to each of the at least one first target image by performing the first type of image processing on a first object included in at least one first target image in at least one first target image based on identifying that the first number is greater than or equal to a first threshold.
[0271] Meanwhile, the various embodiments described above may be implemented as software containing instructions stored on a device-readable storage medium, included in a computer program product in the form of a device-readable storage medium (e.g., flash memory, SSD, HDD, optical disc, magnetic tape, etc.), or distributed online through an application store, website, or cloud server. Additionally, they may be implemented within a recording medium readable by a computer or similar device using software, hardware, firmware, or a combination thereof.
[0272] Each component according to the various embodiments described above may be composed of a single or multiple entities, and some auxiliary components may be omitted or additionally included. Some components may be integrated into a single entity to perform the same or similar functions as those performed by each corresponding component prior to integration.
[0273] Each component according to the various embodiments described above may be composed of a single or multiple entities, and some auxiliary components may be omitted or additionally included. Some components may be integrated into a single entity to perform the same or similar functions as those performed by each corresponding component prior to integration.
[0274] The operations according to the various embodiments described above may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
Claims
1. In an electronic device, Memory comprising at least one storage medium for storing instructions; and at least one processor including a processing circuit; comprising, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Receive multiple image processing requests corresponding respectively to the first plurality of images; Identifying at least one target object and at least one image processing type corresponding to the above plurality of image processing requests; Based on the fact that the above plurality of image processing requests satisfy specified conditions, the user intent corresponding to the above plurality of image processing requests is estimated; Determining a second plurality of images according to the above-mentioned estimated user intent; Providing a query to verify the above-mentioned estimated user intent; and Based on the user response to the above query, causing image processing corresponding to the at least one target object and the at least one image processing type to be performed on at least one image included in the second plurality of images, Electronic device.
2. In Paragraph 1, The above plurality of image processing requests include a sequence of repetitive commands based on user input that are received continuously, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device, Based on each of the above multiple editing requests, causing to generate a first plurality of edited images corresponding to each of the first plurality of images using an artificial intelligence model, Electronic device.
3. In Paragraph 1 or 2, The above specified condition includes a condition in which, based on the plurality of image processing requests, the number of times image processing corresponding to the at least one target object and the at least one image processing type is performed is greater than or equal to a predetermined threshold value. Electronic device.
4. In any one of paragraphs 1 through 3, At least some of the first plurality of images further include at least one other object that does not correspond to the plurality of image processing requests, and The above-mentioned estimated user intent includes an intention not to process the at least one other object according to the at least one image processing type, Electronic device.
5. In any one of paragraphs 1 through 4, The above at least one target object includes a first target object included in a first image and a second target object included in a second image, and the first image and the second image are included in the first plurality of images, and The above at least one image processing type includes a first image processing type for the first image and a second image processing type for the second image, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device, When a third image included in at least one image included in the second plurality of images includes the first target object and the second target object, based on the first image processing type and the second image processing type, causing to generate an edited image corresponding to the third image. Electronic device.
6. In any one of paragraphs 1 through 5, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: By performing image processing on at least one image included in the second plurality of images using an artificial intelligence model, at least one edited image corresponding to each of the at least one image is generated; and Causing to display the at least one edited image, together with a first element indicating that the at least one edited image was generated based on the user response, and a second element for displaying the at least one original image corresponding to each of the at least one edited image. Electronic device.
7. In Paragraph 6, The above plurality of image processing requests are received through a first application, and the at least one edited image is displayed through the first application, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device, Causing to replace and upload an image uploaded through a second application among at least one image included in the second plurality of images with an edited image corresponding to the uploaded image through the second application. Electronic device.
8. In electronic devices, Memory comprising at least one storage medium for storing instructions; and at least one processor including a processing circuit; comprising, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on a first request, at least one first image, a first type of image processing is performed for a first object included in said at least one first image; Counting the first number, which is the number of times image processing based on the first request was performed; Based on identifying that the first number is greater than or equal to a first threshold, the first type of image processing is performed on the first object included in the at least one first target image in at least one first target image, thereby causing to obtain at least one first edited image corresponding to each of the at least one first target image. Electronic device.
9. In Paragraph 8, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on a second request, at least one second image performs a second type of image processing for a second object different from a first object, wherein the second type is an image processing type that is the same as or different from the first type; Counting the second number, which is the number of times image processing based on the second request was performed; Based on identifying that the first number is greater than or equal to the first threshold and the second number is greater than or equal to the second threshold, the first type of image processing is performed on the first object included in the at least one second target image and the second type of image processing is performed on the second object included in the at least one second target image, thereby causing to obtain at least one second edited image corresponding to each of the at least one second target image. Electronic device.
10. In Paragraph 8 or 9, The first threshold value is determined based on at least one of the characteristics of the first object or the first type, Electronic device.
11. In any one of paragraphs 8 through 10, The above-mentioned first object includes an object identified as the same first person, Electronic device.
12. In any one of paragraphs 8 through 10, The above first object includes an object identified as one of the persons belonging to the same first group, Electronic device.
13. In Paragraph 11 or 12, The first threshold value is determined based on at least one of relationship information corresponding to the first object or the first type. Electronic device.
14. In any one of paragraphs 8 through 13, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Identify whether the first edited image maintains contextual consistency; Based on identifying that the first edited image does not maintain context consistency, additional image processing is performed, or the first type of image processing on a first object included in the at least one first target image that was performed is caused to be undoed. Electronic device.
15. In any one of paragraphs 8 through 14, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: The usage history of each of the above at least one first target image is stored in the memory, and the usage history includes an upload history; Based on the usage history of each of the at least one first target image, causing at least one uploaded image among the at least one first target image to be updated to an image among the first edited images corresponding to at least one uploaded image. Electronic device.