Device, method, and recording medium for editing image
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-08-13
Smart Images

Figure KR2026001197_13082026_PF_FP_ABST
Abstract
Description
Device, method, and recording medium for image editing
[0001] The present disclosure relates to an apparatus, a method, and a recording medium that support image editing in an electronic device.
[0002] With the development of digital technology, various types of electronic devices such as mobile communication terminals, PDAs (personal digital assistants), electronic notebooks, smartphones, tablet PCs (personal computers), and wearable devices are widely used. These electronic devices serve as the foundation for creating an environment where still images or videos, such as photos, can be rapidly shared among users based on communication technology and social networking services.
[0003] With the recent advancement of video and content production technologies, a variety of personalized multimedia content is being produced and distributed and consumed through applications that support social networking services. However, as applications operate independently on electronic devices, it may not be easy to process content such as photos or videos by linking multiple applications.
[0004] The information described above may be provided as related art for the purpose of aiding understanding of this document. None of the foregoing is to be claimed as prior art related to this document, nor is it to be used to determine prior art.
[0005] As an example, the electronic device may include a display. The electronic device may include a memory comprising one or more storage media for storing instructions. The electronic device may include at least one processor comprising a processing circuit. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to perform at least one operation. The at least one operation may include an operation of determining a specific area on the execution screen of the first application. The at least one operation may include an operation of obtaining context information reflecting an editing intention by considering information obtained through the use of the first application. The at least one operation may include an operation of downloading an original image associated with the determined specific area. The at least one operation may include an operation of executing a second application to generate an edited image by editing the downloaded original image based on the obtained context information.
[0006] As an example, the operation method of an electronic device may include an operation of determining a specific area on the execution screen of a first application. The operation method may include an operation of obtaining context information reflecting an editing intention by considering information obtained through the use of the first application. The operation method may include an operation of downloading an original image related to the determined specific area. The operation method may include an operation of executing a second application to generate an edited image by editing the downloaded original image based on the obtained context information.
[0007] As an example, a recording medium may store instructions that can be read by a computer. When executed by at least a part of at least one processor included in the electronic device, said instructions may cause said electronic device to perform at least one operation. That at least one operation may include an operation of determining a specific area on the execution screen of a first application. That at least one operation may include an operation of obtaining context information reflecting an editing intention by considering information obtained through the use of said first application. That at least one operation may include an operation of downloading an original image associated with said determined specific area. That at least one operation may include an operation of executing a second application to generate an edited image by editing said downloaded original image based on said obtained context information.
[0008] As an example, the electronic device may include a display. The electronic device may include a memory comprising one or more storage media for storing instructions. The electronic device may include at least one processor comprising a processing circuit. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to perform at least one operation. The at least one operation may include downloading one or more images included in an area selected by a user in a dialog box of an interactive application. The at least one operation may include analyzing the dialogue content displayed in the dialog box of the interactive application. The at least one operation may include generating an edited image by editing some or all of the one or more downloaded images based on the analyzed dialogue content.
[0009] As an example, a method of operation of an electronic device may include the operation of downloading one or more images included in an area selected by a user in a dialog window of an interactive application. The method of operation may include the operation of analyzing dialogue content displayed in the dialog window of the interactive application. The method of operation may include the operation of generating an edited image by editing some or all of the one or more downloaded images based on the analyzed dialogue content.
[0010] As an example, a recording medium may store instructions that can be read by a computer. When executed by at least part of at least one processor included in the electronic device, said instructions may cause said electronic device to perform at least one operation. said at least one operation may include downloading one or more images included in an area selected by a user in a dialog box of an interactive application. said at least one operation may include analyzing the dialogue content displayed in said dialog box of the interactive application. said at least one operation may include generating an edited image by editing some or all of said one or more downloaded images based on said analyzed dialogue content.
[0011] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0012] FIG. 1 is an exemplary block diagram of an electronic device in a network environment according to various embodiments.
[0013] FIG. 2 is a block diagram illustrating a program according to various embodiments.
[0014] FIG. 3 is an exemplary block diagram for image editing in an electronic device according to one embodiment.
[0015] FIG. 4 is a control flowchart for image editing in an electronic device according to one embodiment.
[0016] FIGS. 5a to 5i are drawings for exemplarily illustrating image editing performed in an electronic device according to one embodiment.
[0017] FIG. 6 is a diagram illustrating an image editing operation in an electronic device according to one embodiment.
[0018] FIG. 7 is a block diagram of an exemplary AI system capable of performing the operations described in the present disclosure.
[0019] FIG. 8 is a diagram illustrating the exemplary operation of a multimodal AI model that performs image editing according to one embodiment.
[0020] FIG. 9a, FIG. 9b, or FIG. 9c is a drawing for illustrating an electronic device according to one embodiment selecting a specific area in a captured image.
[0021] FIGS. 10a to 10e are drawings for exemplarily illustrating an operation of editing an image included in a capture screen in an electronic device according to one embodiment.
[0022] FIGS. 11a to 11d are drawings for exemplarily illustrating an editing operation of an image to be edited in an electronic device according to one embodiment.
[0023] FIG. 12 is a diagram illustrating an example of editing an image included in a web browser in an electronic device according to one embodiment.
[0024] FIG. 13 illustrates the configuration of an artificial system inside or outside an electronic device according to one embodiment.
[0025] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.
[0026] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments.
[0027] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0028] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0029] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence is performed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0030] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, software (e.g., program (140)) and input data or output data for related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0031] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0032] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0033] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0034] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0035] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0036] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0037] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0038] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0039] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0040] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0041] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0042] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0043] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0044] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) may support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0045] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0046] According to one embodiment, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0047] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0048] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0049] FIG. 2 is a block diagram (200) illustrating a program according to various embodiments (e.g., the program (140) of FIG. 1).
[0050] According to one embodiment, the program (140) may include an operating system (e.g., the operating system (142) of FIG. 1) for controlling one or more resources of an electronic device (e.g., the electronic device (101) of FIG. 1), middleware (e.g., the middleware (144) of FIG. 1), or an application executable on the operating system (142) (e.g., the application (146) of FIG. 1). The operating system (142) is, for example, Android TM , iOS TM , Windows TM , Symbian TM , Tizen TM , or Bada TM It may include. At least some of the programs (140) may be preloaded into the electronic device (101) at manufacturing time, for example, or downloaded or updated from an external electronic device (e.g., the electronic device (102 or 104) of FIG. 1) or a server (e.g., the server (108) of FIG. 1) when used by a user.
[0051] The operating system (142) can control the management (e.g., allocation or reclamation) of one or more system resources (e.g., process, memory, or power) of the electronic device (101). The operating system (142) additionally or substantially includes other hardware devices of the electronic device (101), for example, an input module (e.g., input module (150) of FIG. 1), an acoustic output module (e.g., acoustic output module (155) of FIG. 1), a display module (e.g., display module (160) of FIG. 1), an audio module (e.g., audio module (170) of FIG. 1), a sensor module (e.g., sensor module (176) of FIG. 1), an interface (e.g., interface (177) of FIG. 1), a haptic module (e.g., haptic module (179) of FIG. 1), a camera module (e.g., camera module (180) of FIG. 1), a power management module (e.g., power management module (188) of FIG. 1), a battery (e.g., battery (189) of FIG. 1), a communication module (e.g., communication module (190) of FIG. 1), a subscriber identification module (e.g., subscriber identification module (196) of FIG. 1), or an antenna module (e.g. It may include one or more driver programs for driving the antenna module (197) of Fig. 1.
[0052] Middleware (144) may provide various functions to an application (146) so that functions or information provided from one or more resources of an electronic device (101) can be used by the application (146). Middleware (144) may include, for example, an application manager (201), a window manager (203), a multimedia manager (205), a resource manager (207), a power manager (209), a database manager (211), a package manager (213), a connectivity manager (215), a notification manager (217), a location manager (219), a graphics manager (221), a security manager (223), a call manager (225), or a voice recognition manager (227).
[0053] The application manager (201) can, for example, manage the life cycle of the application (146). The window manager (203) can, for example, manage one or more GUI resources used on the screen. The multimedia manager (205) can, for example, identify one or more formats required for the playback of media files and perform encoding or decoding of the corresponding media files using a codec that matches the selected corresponding format. The resource manager (207) can, for example, manage the source code of the application (146) or the memory space of the memory (130). The power manager (209) can, for example, manage the capacity, temperature, or power of the battery (189) and, using the relevant information, determine or provide relevant information required for the operation of the electronic device (101). According to one embodiment, the power manager (209) can interact with the BIOS (basic input / output system) (not shown) of the electronic device (101).
[0054] The database manager (211) can, for example, create, search, or modify a database to be used by the application (146). The package manager (213) can, for example, manage the installation or update of an application distributed in the form of a package file. The connectivity manager (215) can, for example, manage a wireless or direct connection between the electronic device (101) and an external electronic device. The notification manager (217) can, for example, provide a function to notify the user of the occurrence of a specified event (e.g., an incoming call, a message, or an alarm). The location manager (219) can, for example, manage location information of the electronic device (101). The graphics manager (221) can, for example, manage one or more graphic effects or related user interfaces to be provided to the user.
[0055] The security manager (223) may, for example, provide system security or user authentication. The telephony manager (225) may, for example, manage voice call functions or video call functions provided by the electronic device (101). The voice recognition manager (227) may, for example, transmit user voice data to the server (108) and receive from the server (1308) a command corresponding to a function to be performed on the electronic device (101) based on at least part of the voice data, or text data converted based on at least part of the voice data. According to one embodiment, the middleware (244) may dynamically delete some existing components or add new components. According to one embodiment, at least part of the middleware (144) may be included as part of the operating system (142) or implemented as separate software different from the operating system (142).
[0056] The application (146) may include, for example, a home (251), a dialer (253), an SMS / MMS (255), an IM (instant message) (257), a browser (259), a camera (261), an alarm (263), a contact (265), a voice recognition (267), an email (269), a calendar (271), a media player (273), an album (275), a watch (277), a health (279) (e.g., measuring biometric information such as exercise volume or blood sugar), or an environmental information (281) (e.g., measuring atmospheric pressure, humidity, or temperature information). According to one embodiment, the application (146) may further include an information exchange application (not shown) capable of supporting information exchange between the electronic device (101) and an external electronic device. The information exchange application may include, for example, a notification relay application configured to transmit information (e.g., a call, a message, or an alarm) designated to an external electronic device, or a device management application configured to manage the external electronic device. The notification relay application may transmit notification information corresponding to a designated event (e.g., receiving mail) generated in another application of the electronic device (101) (e.g., an email application (269)) to the external electronic device. Additionally or alternatively, the notification relay application may receive notification information from the external electronic device and provide it to the user of the electronic device (101).
[0057] A device management application can control the power (e.g., turn-on or turn-off) or function (e.g., brightness, resolution, or focus) of an external electronic device or a part of its components (e.g., a display module or camera module of the external electronic device) that communicates with the electronic device (101). The device management application can additionally or substantially support the installation, deletion, or updating of applications running on the external electronic device.
[0058] FIG. 3 is an exemplary block diagram (300) for image editing in an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment.
[0059] Referring to FIG. 3, the application (330) (e.g., the application (146) of FIG. 1 or FIG. 2) may include a plurality of applications to perform specific functions, such as an interactive application or an image editing application. The application (330) may further include an information exchange application (not shown) capable of supporting information exchange between the electronic device (101) and an external electronic device. For example, the application (330) may include one or more applications that can be executed or are currently running on the electronic device (101). To perform multimodal services, at least two applications (e.g., a first application (331) or a second application (333)) may be executed on the application (330). For example, the second application (333) may be executed in response to the occurrence of a specific event while the first application (331) is running in advance on the electronic device (101). Data obtained by the first application (331) (e.g., downloaded image data) may be provided to the second application (333). The second application (333) may use the data provided by the first application (331) to perform a multimodal service (e.g., image editing service) associated with the operation of the first application (331).
[0060] According to one example, the first application (331) may be an information exchange application and may be a conversational application capable of providing a chat service between two or more participants. A conversational application is also referred to as a messenger. The first application (331) may provide a chat service that allows conversation with a designated counterpart using emoticons, text, images, or audio. For example, the first application (331) running on an electronic device (101) may transmit emoticons, text, or images to one or more participants through a chat window, or receive emoticons, text, images, or audio from one or more participants. For example, the first application (331) may transmit or receive images and / or audio in file format through a chat window.
[0061] According to one example, the second application (333) may be an application capable of editing a video or image (hereinafter referred to as 'image'). In the disclosure described below, the term 'image' may be used to refer to either a still image (e.g., a photograph) or a video. An image editing application may be called an image editor. The second application (333) may provide a service capable of editing an image that is a still image or a video by corresponding to a predetermined command.
[0062] Middleware (320) (e.g., middleware 144 of FIG. 1 or FIG. 2)) may provide various functions to an application (330) so that functions or information provided from one or more resources of an electronic device (101) can be used by the application (330). Middleware (320) may include, for example, an operations manager (321) or an analysis manager (323).
[0063] According to one example, the operation manager (321) can capture the execution screen of the first application (331) and analyze the captured image to determine a function that can select an action to apply to the target image. For example, the operation manager (321) can identify that it is possible to perform an action that can process the image included in the captured image (e.g., sharing, editing, sharing after editing, setting as wallpaper), and can display function identifiers that can select the corresponding actions.
[0064] According to one example, the operation manager (321) can obtain an original image corresponding to a captured image. For example, when an image included in the execution screen of the first application (331) is selected as a target image for editing, the operation manager (321) may request the first application (331) to download the original image of the target image.
[0065] According to one example, the operations manager (321) may save or analyze the edits of the target image. For example, if the operations manager (321) identifies an editing intention to remove a specific person from the target image, it may save this as the edit content.
[0066] According to the example, the analysis manager (323) can analyze a captured image of the execution screen of the first application (331) or a selected portion image of the execution screen and generate a description regarding the screen. The analysis manager (323) can identify a shared image from the captured image by checking the context of the screen. For example, the analysis manager (323) can identify a person participating in a conversation by analyzing information visible on the screen, such as in a messenger. The analysis manager (323) can obtain information to be reflected in the edit by analyzing the conversation content included in the captured image, for example, inferring a situation where an edit is requested.
[0067] According to one example, the hardware (310) may include a processor (311) (e.g., the processor (120) of FIG. 1) or memory (313) (e.g., the memory (130) of FIG. 1).
[0068] The processor (311) can control at least one other electrically connected component (e.g., hardware or software component) by executing software (e.g., a program). The processor (311) can perform various data processing or operations. As at least part of the data processing or operations, the processor (311) can store instructions or data received from other components (e.g., an application (330) and / or middleware (320)) in memory (313) (e.g., volatile memory, but without limitation). As at least part of the data processing or operations, the processor (311) can process instructions or data stored in memory (313). As at least part of the data processing or operations, the processor (311) can store the resulting data of the processing instructions or data in memory (313).
[0069] Memory (313) may store various data used by at least one component (e.g., processor (311), middleware (320) and / or application (330)). The data may include, for example, input data or output data for software (e.g., program) and related instructions. Memory (313) may also store at least one AI model (e.g., sLLM (smaller large language model), LLM (Large language model), LVM (Large vision model), LMM (Large multimodal model)) for executing one or more instances.
[0070] Memory (313) can store at least one instruction. Processor (311) can execute at least one instruction stored in memory (313). When executed by processor (311), at least one instruction may cause an AI system (e.g., AI system (700) of FIG. 7) to perform at least one action. For example, as at least one instruction is executed by processor (311), at least one other component may be controlled, and / or various data processing or operations may be performed. An action performed by processor (311) may mean that the action is performed by (or controlled by) one entity included in processor (311), for example (e.g., a main processor, but not limited thereto). An action performed may mean that a specific action is performed by (or controlled by) multiple entities (e.g., multiple processors). The execution of multiple operations may mean that all of the multiple operations are performed by (or by control of) one entity, for example (e.g., the main processor (121) of FIG. 1, but is not limited to). The execution of multiple operations may mean that some of the multiple operations are performed by at least one entity, and some of the remaining operations are performed by at least one other entity. At least one instruction that causes the execution of one or more operations may be stored in, for example, one memory, or distributed among multiple memories.
[0071] FIG. 4 is a control flowchart for image editing in an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1).
[0072] The control flow for image editing in FIG. 4 assumes a situation in which an electronic device (101) edits an image displayed in a chat window using another application while chatting with at least one participant via an interactive application.
[0073] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0074] Referring to FIG. 4, an electronic device (101) can analyze an image and / or text included in a selected area and / or the entire area of a screen displayed on a display, infer an editing intent based on the analysis results, and perform editing on a target image for editing (hereinafter referred to as 'target image for editing') by reflecting the inferred editing intent. For example, the electronic device (101) can share the edited image with another electronic device through a specific application (e.g., an interactive application). For example, the electronic device (101) can automatically delete the target image for editing and / or the edited image immediately after sharing, or after a predetermined time has elapsed after sharing.
[0075] To explain this in detail, the electronic device (101) can select a specific area on the execution screen of a specific application (e.g., the first application (331) of FIG. 3) in operation 410. For example, the electronic device (101) can select an area to obtain information based on image analysis from the current screen displayed on the display by the execution of the specific application. The electronic device (101) can, for example, set the entire area of the current screen as the target area to be analyzed for obtaining information. The electronic device (101) can, for example, set a part of the entire area of the current screen as the target area to be analyzed for obtaining information.
[0076] The electronic device (101) can determine a specific area for obtaining an image to be analyzed from the execution screen of a specific application in response to a specific request event generated by a user's request.
[0077] According to one example, a specific request event may be triggered by a user selecting all or part of the current screen using one of various input methods, such as a touch using a tool like a finger or an electronic pencil. The electronic device (101) may use the screen of the selected area as an image to obtain information for determining a supported function. In this case, the electronic device (101) may obtain an image of the selected area, but is not required to capture and store the image.
[0078] According to one example, a specific request event may occur after the assistant is called by the user, and the user selects a specific area on the execution screen. The electronic device (101) may use the screen of the selected area as an image to obtain information in order to determine the supported functions. In this case, the electronic device (101) may obtain an image of the selected area, but is not required to capture and save the image.
[0079] According to one example, a specific request event may be triggered by a screen capture request. A screen capture request may be triggered by the pressing of one or a number of keys forming a specific combination provided in the electronic device (101), for example. A screen capture request may be triggered in response to the input of a speech command, such as “screen capture,” after an assistant (e.g., Bixby) is invoked in the electronic device (101), for example. A screen capture request may be triggered in response to the selection of a capture function identifier (e.g., an icon) from a displayed menu in response to a swipe gesture from one side of the screen (e.g., the left side of the screen) toward the center of the screen in the electronic device (101).
[0080] The electronic device (101) can determine a supported function by analyzing a target image in operation 420. The target image may be, for example, an entire image of an execution screen or an image that captures the entire screen. The target image may be, for example, an image obtained from a specific area selected in the execution screen or an image that captures a specific area. The electronic device (101) can determine a supported function by analyzing a target image corresponding to the entire execution screen of the first application (331). The electronic device (101) can determine a supported function by analyzing a target image corresponding to a specific area selected in the execution screen of the first application (331). The target image to be analyzed by the electronic device (101) may include an image corresponding to text. The image corresponding to text may include, for example, conversation content received from a counterpart in a chat window. The image corresponding to text may include, for example, conversation content entered by a user in a chat window. The image corresponding to text may include, for example, conversation content received from a counterpart in a chat window and conversation content entered by a user.
[0081] According to one example, the electronic device (101) can analyze a target screen. The electronic device (101) can analyze a target image obtained from the target screen. The electronic device (101) can obtain information regarding a valid object (e.g., image or text) by analyzing the target screen or target image (hereinafter referred to as 'target screen' for convenience of explanation). The electronic device (101) can, for example, identify an image and / or text included in the target screen. The electronic device (101) can determine an image to be edited on the target screen. The electronic device (101) can obtain context information such as a place, surrounding objects, people, weather, or date by analyzing the image identified on the target screen. The electronic device (101) can convert an image determined to be text on the target screen into text, and analyze the content of the converted text to obtain context information corresponding to content such as a conversation partner and / or commands for image editing, such as “Please delete the person behind this” or “Please delete everyone except my family.”
[0082] According to one example, the electronic device (101) may additionally collect relevant information that needs to be considered for image editing. The electronic device (101) may obtain additional context information for image editing by analyzing profile information, such as photos, names, classification groups, or workplaces of people registered in contacts for voice / video calls or chats, for example. For example, if the electronic device (101) receives a request to “remove everyone except my family” based on an analysis of the target screen, it may obtain personal data (e.g., profile pictures) registered as family in the contacts as context information. For example, if the electronic device (101) receives a request to “remove everyone except us” based on an analysis of the target screen, it may obtain profile pictures registered for participants, including itself, who are participating in the chat room in the contacts as context information.
[0083] As described above, the electronic device (101) may identify images and / or text included in a target screen, select one or more images to be edited based on the identified targets, and / or obtain context information for editing one or more images to be edited. As an example, the electronic device (101) may directly analyze conversation content collected by the execution of the first application (331) and / or images transmitted or received, rather than captured images, to select images to be edited, and / or obtain context information for editing images to be edited.
[0084] According to one example, when an image to be edited is selected, the electronic device (101) may determine one or more supported functions applicable to the selected image to be edited. The electronic device (101) may determine, for example, as supported functions that can be performed on the image to be edited, such as sharing, editing, sharing after editing, or setting the background image. For example, the electronic device (101) may output an identifier (e.g., an icon or an emoticon) corresponding to the determined supported function through a display so that it can be selected by a user.
[0085] The electronic device (101) can perform editing on an original image corresponding to one or more images to be edited based on a selection function in operation 430. According to one example, the electronic device (101) can request the download of an original image corresponding to one or more images to be edited through the first application (331). For example, when an event for editing occurs, the electronic device (101) may perform or not perform an automatic download of the original image by considering whether the automatic download function is on or off.
[0086] When the download of the original image is completed through the first application (331), the electronic device (101) can perform editing on the original image based on context information obtained by analyzing the target screen. For example, the electronic device (101) can execute an application that provides an image editing function (e.g., the second application (333) of FIG. 3) for editing the original image. The electronic device (101) can transmit the original image and context information to be referenced for editing to the executed second application (333). For example, if the first application (331) is an interactive application, the electronic device (101) can analyze the conversation content and perform editing on the original image so that inferences based on the analysis results can be reflected. In this case, the electronic device (101) can guide the user with information regarding the direction of editing and proceed with editing in response to the user's request.
[0087] According to one example, if an electronic device (101) intends to perform image editing by a multimodal AI model (e.g., the multimodal AI model (800) of FIG. 7), it may request the multimodal AI model to edit the original image based on the analysis results and / or context information of the image to be edited. In this case, the multimodal AI model can generate data that is difficult to obtain with only a single type of data (e.g., edited images) by processing or analyzing the original image by considering at least two different types of data (or modalities, e.g., original image and text information, e.g., context information and / or prompts) together. The multimodal AI model can identify complex information and / or patterns by combining different types of data (e.g., images, text, or audio). The multimodal AI model can improve the accuracy of the prediction model by integrating multiple types of data or compensate for a lack of data by reducing noise. The multimodal AI model can support decision-making in dynamic environments by analyzing the source of the data to understand the situation more accurately. Therefore, the multimodal AI model can generate the desired type of output by fusing various types of data sources. Thus, when utilizing the multimodal AI model, an image edited in a desired form can be generated by considering different types of modalities (e.g., images, text, or audio) obtained through the execution of the first application (331). As an example, an image by the multimodal AI model When editing is performed, the multimodal AI model may provide a prompt for image editing based on the analysis results provided by the middleware (320).
[0088] The electronic device (101) can perform an editing operation on one or more original images by means of a second application (333) or a multimodal AI model. For example, if an editing request such as “delete everyone except my family” is received, “delete” can be mapped to a deletion function, and the object to be deleted can be specified by “everyone except my family.” In this way, the electronic device (101) can identify images of others who are not family members from one or more original images. For example, if an editing request such as “delete everyone except us” is received, the electronic device (101) can identify images of others from one or more original images, excluding participants including itself who are participating in the chat. The electronic device (101) can perform an editing operation to delete the identified images of others from one or more original images. The editing described above is merely exemplary and can be applied in the same or similar manner to various other image editing techniques.
[0089] In operation 440, the electronic device (101) can share one or more edited images with another electronic device. According to one example, the electronic device (101) can upload one or more edited images by the first application (331). In this case, the uploaded edited images can be shared with participants attending the chat.
[0090] According to one example, the electronic device (101) may delete one or more original images and / or one or more edited images that were downloaded after sharing one or more edited images. The electronic device (101) may delete one or more original images and / or one or more edited images that were downloaded, for example, immediately after sharing one or more edited images. The electronic device (101) may delete one or more original images and / or one or more edited images that were downloaded, for example, after a certain amount of time has elapsed after sharing one or more edited images.
[0091] According to one example, the electronic device (101) deletes one or more original images and / or one or more edited images that were downloaded immediately after sharing one or more edited images, and can also completely delete them from a temporary storage area such as a trash bin after a certain period of time has elapsed. Here, the deletion of one or more original images and / or one or more edited images may mean complete deletion that cannot be verified in a second application (333) that performed image editing, such as a gallery, in addition to erasing from memory.
[0092] FIGS. 5a to 5i are drawings for exemplarily illustrating image editing performed in an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1).
[0093] Referring to FIG. 5a, the electronic device (101) provides a chat service with a conversation partner named 'Young-hee' in a conversation window (500), which is the execution screen of an interactive application (e.g., the first application (331) of FIG. 3). For example, according to the conversation content included in the conversation window (500), it can be confirmed that Young-hee receives a photo (502) and a message (503) saying "Please delete the person behind this" in response to the user's sent message "What do you need me to do?". The user sends a message (505) saying "Okay" accepting Young-hee's request through the conversation window (500). The electronic device (101) can analyze the conversation content in the conversation window (500) and infer that an edit to delete a specific object (e.g., a person) from the photo sent by Young-hee has been requested.
[0094] Referring to FIG. 5b, the electronic device (101) may have a specific area (512) designated by an area selection indicator (511) in a layer (510) that is displayed to cover a screen corresponding to an image (510) of a conversation window or a conversation window (500). For example, the user may designate a specific area (512) so that only the image sent by Young-hee is selected in the captured image (510) or the layer (510). For example, the electronic device (101) may automatically designate a part of the captured image (510) or the layer (510) that seems meaningful and requires analysis as a specific area (512).
[0095] Referring to FIG. 5c, the electronic device (101) may display a selection window (520) containing function identifiers (521, 522) corresponding to supported functions for processing an image of a specific region (512) when a specific region (512) is selected by a region selection indicator (511) in a captured image (510) or a layer (510). For example, supported functions may include functions for performing actions such as sharing or editing.
[0096] Referring to FIG. 5d, the electronic device (101) may have a specific area (532, 533) designated by an area selection indicator (531) in an image (510) that captures a conversation window or a layer (510) that is displayed to cover a screen corresponding to the conversation window. For example, the user may designate a specific area in the captured image (510) or layer (510) to include an image (532) sent by Young-hee and a message (e.g., "Please erase the person behind this") (533) that can be inferred as an editing intent. When a specific area is selected by the area selection indicator (531) in the captured image (510) or layer (510), the electronic device (101) may analyze the image (532) and text (533) contained in the specific area and display a selection window (540) containing a function identifier (541, 542) corresponding to a supported function based on the analysis result. For example, supported features may include functions that perform actions such as sharing or editing.
[0097] Referring to FIG. 5e, the electronic device (101) may have a specific area (532, 534) designated by an area selection indicator (531) in an image (510) that captures a conversation window or a layer (510) that is displayed to cover a screen corresponding to the conversation window. For example, the user may designate a specific area in the captured image (510) or layer (510) to include an image (532) sent by Young-hee and a message (e.g., "Delete everyone except my family") (534) that can be inferred as an editing intent. When a specific area is selected by the area selection indicator (531) in the captured image (510) or layer (510), the electronic device (101) may analyze the image (532) and text (534) contained in the specific area and display a selection window (540) containing a function identifier (541, 542) corresponding to a supported function based on the analysis results. For example, supported features may include functions that perform actions such as sharing or editing.
[0098] Referring to FIG. 5f, the electronic device (101) can download the image sent by Young-hee through a dialog box for editing. The electronic device (101) can display the downloaded image (551) on the execution screen (550) of an application capable of video editing (e.g., the second application (333) of FIG. 3). The electronic device (101) can identify the main object (552) to be retained and the object (553) to be deleted by reflecting the inferred editing intent (e.g., a request to delete a specific object in a photo) by analyzing the content of the message sent by Young-hee. The electronic device (101) can create an edited image by deleting the identified object (553) included in the image (551). The electronic device (101) may request confirmation of the editing intent from the user before editing the original image (551). The electronic device (101) may, for example, display an object (553) to be deleted from the screen to distinguish it and display an identifier (554) (e.g., an icon or button) that can request the creation of an edited image. When the identifier (554) is selected, the electronic device (101) may create an edited image with the identified object (553) deleted (see FIG. 5g).
[0099] Referring to FIG. 5h, the electronic device (101) may display a selection window (560) to allow the user to select whether to proceed with further editing of the edited image or to share the edited image after displaying the edited image on the screen. The selection window (560) may include a first identifier (561) that can request further editing of the edited image or a second identifier (562) that can request sharing of the edited image.
[0100] Referring to FIG. 5i, when the electronic device (101) requests the sharing of an edited image by a user, it can send an edited image (571) and a conversation message (e.g., I'm done) (573) in a conversation window, which is the execution screen of an interactive application (e.g., the first application (331) of FIG. 3).
[0101] FIG. 6 is a diagram illustrating an image editing operation in an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1).
[0102] The image editing operation in FIG. 6 assumes a situation in which an electronic device (101) edits an image displayed in a chat window using another application while chatting with at least one participant via an interactive application. For example, the first application of the electronic device (101) (e.g., the first application (331) of FIG. 3) may be an interactive application, and the second application of the electronic device (101) (e.g., the second application (333) of FIG. 3) may be an application capable of editing images.
[0103] In the following embodiments, each signal processing procedure may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each signal processing procedure may be changed, and at least two signal processing procedures may be performed in parallel.
[0104] Referring to FIG. 6, the electronic device (101) may generate at least one edited image by editing one or a number of specific original images based on an analysis of the target screen (operations 611 to 621). According to one example, the target screen may be an execution screen of the first application (331) or a selected portion of the execution screen. If the first application (331) is an interactive application, the target screen may be, for example, a screen corresponding to a chat window where chatting is being performed with at least one participant. If the first application (331) is an interactive application, the target screen may be, for example, a selected portion of the chat window where chatting is being performed with at least one participant. According to one example, the target screen may be an image captured from the execution screen of the first application (331) or an image captured from a specific area selected from the execution screen of the first application (331).
[0105] More specifically, the middleware of the electronic device (101) (e.g., the middleware (320) of FIG. 3) may, in operation 611, determine a target screen in response to the occurrence of a specific request event and perform an analysis on the determined target screen. An AI model may be used to perform an analysis of the target screen by the middleware (320). A specific request event may be generated by a screen capture request. For example, a screen capture request may be generated by the pressing of one or a number of keys forming a specific combination provided in the electronic device (101). For example, a screen capture request may be generated in response to the input of a speech command, such as “screen capture,” after an assistant (e.g., Bixby) is called on the electronic device (101). For example, a screen capture request may be generated in response to the selection of a capture function identifier (e.g., an icon) from a displayed menu in response to a swipe gesture from one side of the screen (e.g., the left side of the screen) toward the center of the screen on the electronic device (101).
[0106] When a predetermined request event occurs, the middleware (320) may determine an image captured from an execution screen provided by the execution of the first application (331) as a target screen for analysis. For example, the target screen may be an image that captures the entire execution screen. For example, the target screen may be an image that captures a part of the execution screen. The target screen may include an image for editing (hereinafter referred to as 'image to be edited'). The target screen may include an image to be edited and an image corresponding to text. The image corresponding to text may include, for example, conversation content received from a counterpart in a chat window. The image corresponding to text may include, for example, conversation content entered by a user in a chat window. The image corresponding to text may include, for example, conversation content received from a counterpart in a chat window and conversation content entered by a user.
[0107] Middleware (320) can analyze the target screen. Middleware (320) can obtain information regarding valid objects (e.g., images or text) by analyzing the target screen. According to one example, middleware (320) can identify images and / or text included in the target screen. Middleware (320) can determine the image to be edited on the target screen. Middleware (320) can obtain context information such as location, surrounding objects, people, weather, or date by analyzing the image identified on the target screen. Middleware (320) can convert an image determined to be text on the target screen into text, and analyze the content of the converted text to obtain context information corresponding to content such as conversation partners and / or commands for image editing, such as “Please delete the person behind this” or “Please delete everyone except my family.”
[0108] Middleware (320) may additionally collect relevant information that needs to be considered for image editing. For example, middleware (320) may obtain additional context information for image editing by analyzing profile information such as photos, names, classification groups, or workplaces of people registered in contacts for voice / video calls or chats. For example, if a request to “remove everyone except my family” is received based on an analysis of the target screen, middleware (320) may obtain the profile photos of family members registered in contacts as context information. For example, if a request to “remove everyone except us” is received based on an analysis of the target screen, middleware (320) may obtain the profile photos of participants, including oneself, who are participating in the chat room from contacts as context information.
[0109] As described above, the middleware (320) may identify images and / or text included in the target screen, select one or more images to be edited based on the identified targets, and / or obtain context information for editing one or more images to be edited. However, the middleware (320) may also directly analyze conversation content collected by the execution of the first application (331) and / or images transmitted or received, rather than captured images, to select images to be edited or / or obtain context information for editing images to be edited.
[0110] Middleware (320) may request the first application (3231) to download original images corresponding to one or more images to be edited in operation 613. For example, middleware (320) may transmit identification information to the first application (331) indicating the images to be edited corresponding to one or more original images to be downloaded.
[0111] The first application (331), in operation 615, recognizes one or more images to be edited for which the download of the original image has been requested from the middleware (320), and can perform the download of the original image corresponding to the image to be edited. The first application (331) can download one or more original images, for example, using a pre-configured API (application programming interface). When the download of the original image is completed, the first application (331) can notify the middleware (320) that the download of the original image is completed in operation 617.
[0112] When the middleware (320) receives a notification from the first application (331) that the original image download is complete, in operation 619, it may transmit an editing request and analysis results to the second application (333). For example, the middleware (320) may transmit one or more original images downloaded by the first application (331) to the second application (333). For example, the middleware (320) may transmit context information obtained by target screen analysis to the second application (333). For example, if based on a multimodal AI model (e.g., the multimodal AI model (800) of FIG. 7), the middleware (320) may generate a prompt requesting editing of the original image based on the analysis results and / or context information of one or more images to be edited, and provide this to the multimodal AI model.
[0113] In the description above, it is stated that the original image is transferred from the first application (331) to the middleware (320) or from the middleware (320) to the second application (333). However, considering that the first application (331), the middleware (320), or the second application (333) are all configurations that can be executed on the electronic device (101), the configuration may be described as providing a path for the original image to be transferred, rather than actually performing the operation of transferring the original image. In the following description, the operation of transferring the original image and / or the edited image between the first application (331), the middleware (320), or the second application (333) may be disclosed, but this may not necessarily be based on the premise that an operation to transfer or receive the image must be performed. That is, only a path may be provided so that the image can be shared between the first application (331), middleware (320), or second application (333), and the operation by the configuration may not actually need to be performed.
[0114] When the second application (333) receives an editing request from the middleware (320) along with one or more original images and / or context information, it can perform an editing operation on one or more original images in operation 621. For example, if an editing request such as “delete everyone except my family” is received from the middleware (320), the second application (333) can identify images of other people who are not family from one or more original images. For example, if an editing request such as “delete everyone except us” is received from the middleware (320), the second application (333) can identify images of other people excluding participants, including itself, who are participating in the chat from one or more original images. The second application (330) can perform an editing operation to delete the images of other people identified from one or more original images. The editing described above is merely exemplary and may be applied in the same or similar manner to various other image editing techniques.
[0115] In operation 623, the second application (333) may notify the middleware (320) that editing of one or more original images has been completed. For example, when notifying the completion of editing, the second application (333) may transmit one or more edited images (hereinafter referred to as 'edited images') to the middleware (320). For example, a path may be provided between the second application (333) and the middleware (320) to allow one or more edited images to be shared.
[0116] As previously described, it was assumed that image editing is performed by the second application (333), but it may be performed by a multimodal AI model other than the second application (333), or the second application (333) may perform it using a multimodal AI model.
[0117] A multimodal AI model can generate data that is difficult to obtain with only a single type of data (e.g., edited images) by processing or analyzing an original image by considering at least two different types of data (or modalities, e.g., original images and text information, e.g., context information and / or prompts) together. A multimodal AI model can identify complex information and / or patterns by combining different types of data (e.g., images, text, or audio). By integrating multiple types of data, a multimodal AI model can improve the accuracy of a prediction model or reduce noise to compensate for a lack of data. A multimodal AI model can support decision-making in dynamic environments by analyzing the source of the data to understand the situation more accurately. Therefore, a multimodal AI model can generate a desired type of output by fusing various types of data sources. Thus, when utilizing a multimodal AI model, an image edited in a desired form can be generated by considering different types of modalities (e.g., images, text, or audio) obtained through the execution of the first application (331). As an example, image editing is performed by a multimodal AI model. In this case, the multimodal AI model may provide a prompt for image editing based on the analysis results provided by the middleware (320).
[0118] Middleware (320) can transmit one or more edited images to the first application (331) so that one or more edited images can be used by the first application (331) in operation 625. For example, middleware (320) can request the first application (331) to upload one or more edited images.
[0119] The first application (331) can perform specific functions using one or more edited images received from the middleware (320). For example, if an upload is requested from the middleware (320), the first application (331) can send one or more edited images through a chat window. In this case, the edited images can be shared with participants attending the chat.
[0120] Middleware (320) can delete one or more original images and / or one or more edited images that were downloaded after delivering one or more edited images to the first application (331). For example, middleware (320) can delete one or more original images and / or one or more edited images that were downloaded immediately after delivering one or more edited images to the first application (331). For example, middleware (320) can delete one or more original images and / or one or more edited images that were downloaded after a certain amount of time has elapsed after delivering one or more edited images to the first application (331).
[0121] For example, the middleware (320) may delete one or more original images and / or one or more edited images that were downloaded immediately after delivering one or more edited images to the first application (331), and may also completely delete them from a temporary storage area such as a trash bin after a certain period of time has elapsed. Here, the deletion of one or more original images and / or one or more edited images may mean complete deletion that cannot be verified in the second application (333) that performed the image editing, such as a gallery, in addition to erasing them from memory.
[0122] FIG. 7 is a block diagram of an exemplary AI system (700) capable of performing the operations described in the present disclosure. The AI system (700) may be a generative AI system, but will be referred to as 'AI system (700)' below.
[0123] The AI system (700) in FIG. 7 may include an on-device AI system or a cloud AI system. For example, an on-device AI system may be an AI system capable of processing information on its own within an electronic device (e.g., the electronic device (101) in FIG. 1) without needing to be connected to a server or the cloud. For example, a cloud AI system may be an AI system formed by a combination of AI technology and a cloud computing architecture, capable of providing AI services and / or computing capabilities via the cloud.
[0124] For example, the present disclosure assumes an on-device AI system; however, the proposed examples are not to be limited to on-device AI systems and may also be implemented by a cloud AI system or by linking an on-device AI system with a cloud AI system. For example, an on-device AI system may be implemented in one of various types or forms of electronic devices, such as a laptop, smartphones with various form factors, a tablet, an AR device, a cellular phone, and other similar computing devices. Smartphones with various form factors may include, for example, bar-type smartphones, foldable-type smartphones, or sliderable (or rollable)-type smartphones. The components, their relationships, and their functions illustrated in FIG. 7 are merely illustrative and do not limit the implementations described or claimed in this document.
[0125] According to one example, the AI system (700) may share resources (e.g., data processing or computational power) corresponding to part or all of at least one processor included in the processor (e.g., processor (311) of FIG. 3) and / or resources (e.g., data recording area) corresponding to part or all of the memory (e.g., memory (313) of FIG. 3). The AI system (700) may be operated by at least one of a CPU, a GPU, or an NPU. The AI system (700) may, for example, be allocated at least a portion of the memory (313) and be performed solely by the CPU. The AI system (700) may, for example, be allocated at least a portion of the memory (313) and be performed solely by the GPU. The AI system (700) may, for example, be allocated at least a portion of the memory (313) and be performed solely by the NPU. The AI system (700) may, for example, be allocated at least a portion of the memory (313) and be performed by the CPU and the GPU in cooperation. The AI system (700) can be performed by the CPU and NPU cooperating, for example, by allocating at least a portion of the memory (313). The AI system (700) can be performed by the GPU and NPU cooperating, for example, by allocating at least a portion of the memory (313). The AI system (700) can be performed by the CPU, GPU, and NPU cooperating, for example, by allocating at least a portion of the memory (313). Various embodiments described below in this disclosure are not limited to combinations of components for performing the AI system (700) and may be implemented and / or applied based on any combination.
[0126] Referring to FIG. 7, the AI system (700) may include a user query / response interface (710) (hereinafter referred to as 'I / F (710)'), an AI framework (720) (e.g., middleware (320) of FIG. 3), a generative AI model (730), a database (740), or an application / service component (750) (e.g., application (330) of FIG. 3). Each of these modules or its sub-modules (e.g., Prompt Design Module (221)) may be trained using a machine learning algorithm or a neural network and thereby produce improved results.
[0127] The I / F (710) may receive data or user input (e.g., user query) acquired or generated by an electronic device (e.g., electronic device (101) of FIG. 1). Data acquired or generated by the electronic device (101) may include image or video data generated using a processor (e.g., processor (120) of FIG. 1), or values received through a sensor or sensor hub. User input may be received in the form of text, images, voice and / or video. According to one example, user input may be of a data type such as natural language (e.g., text and / or audio), touch coordinates or stylus coordinates acquired through a touch panel or digitizer included in a display (e.g., display module (160) of FIG. 1), and / or images (e.g., photos and / or videos), but is not limited thereto. The I / F (710) may receive context information in addition to or on behalf of the user input. Additionally, context information may be transmitted together with the transmission of user input. Context information may include various additional information related to the generative artificial intelligence system (e.g., mobile phone) (700) at the time user input is received. For example, additional information may include information about the application currently being used by the user or information about the user's location. Additionally, user input (e.g., a user query entered into an electronic device) may be received in the form of text, images, voice, context information, or a combination thereof. Furthermore, user input may also be in a non-natural form, such as selecting a menu.
[0128] I / F (710) can output results of the AI system (700) and / or results of analyzing inputs. For example, I / F (710) can provide the user with the results of processing by the generative artificial intelligence system (700) on user input. The results of processing by the generative artificial intelligence system (700) may be in the form of natural language, such as text, images, or voice, or in the form of specific content. The results of processing by the generative artificial intelligence system (700) may also be provided in the form of an action requested by the user. The results of processing by the generative artificial intelligence system (700) may also be provided in the form of a specific value designated by the user. I / F (710) can output the results of the generative AI system (700) to the user. The results may be provided in the form of text, images, voice, an action requested by the user, or a combination thereof.
[0129] The AI framework (720) can receive input from the user and coordinate and control each component necessary to perform the user's intent based on the user's query. For example, the AI framework (720) may include a prompt design component (721), an API / plug-in management component (723), or an output modification component (or refiner component) (725).
[0130] User input received from the I / F (710) can be transmitted to a prompt design component (721). The prompt design component (721) can be used to generate prompts suitable for inputting user input into a generative AI model (730) (e.g., LLM, LVM, or LMM). The prompt design component (721) may be an AI component that uses machine learning algorithms or neural networks to develop better prompts over time. Based on user input, the prompt design component (721) can generate prompts by accessing user preference data (743), a prompt library (741), and a knowledge component containing prompt examples, and can transmit the generated prompts to a generative AI model (730), such as an LLM or LMM.
[0131] The API / plugin management component (723) can perform the role of communicating with external information when there is a request for additional information when user input is passed as input to the generative AI model (730). The API / plugin management component (723) establishes a channel to communicate with the outside of the AI interface via the API, and enables access to various data sources (e.g., knowledge repositories (745)) through the established channel.
[0132] The API / plugin management component (723) can request the application / service component (723) via the API to perform the action that ultimately performs user input, rather than an intermediate result, when the application or service needs to perform such action. The information obtained from the outside can be used to generate a prompt in the prompt design component (721) along with the user input, or can be passed as input to the generative AI model (730).
[0133] The output adjustment component (725) (or refiner component) can fine-tune or reprocess the output from the generative AI model (730). The output adjustment component (725) can, for example, verify whether the content generated by the generative AI model (730) is irrelevant, contains biased content, or contains harmful content. The output adjustment component (725) can determine the extent to which the output matches the desired result and, if additional processing is required, proceed with that process. Additionally, the output adjustment component (725) can configure and provide hints to the user to avoid unwanted output.
[0134] A generative AI model (730) generally refers to an AI neural network that generates new forms of data based on user input information. A generative AI model (730) may include a model that generates images and / or a model that generates language. A model that generates images may include, for example, a generative adversarial network (GAN) or a variational auto encoder (VAE). A model that generates images may be a diffusion-based AI model that uses, for example, a VAE and a transformer structure. A model that generates language may be a model trained to output the statistically most appropriate output value based on input values. Representative examples include models such as CHAT-GPT 3 and CHAT-GPT 4. Additionally, LMM is an AI model (730) capable of recognizing various forms of data input, such as text, images, voice, and video, and generating new data corresponding to them.
[0135] According to one example, the electronic device (101) may include a natural language recognition module, an NLP module, or a planner module for performing a conversation with a user using natural language. The NLP module may be an AI model pre-trained to enable the machine, the electronic device (101), to perform a conversation with a user using natural language. The NLP module may be a software module implemented, for example, by executing instructions by at least one processor (e.g., the processor (120) of FIG. 1). Hereinafter, the actions performed by the NLP module may be referred to as actions that may cause the electronic device (101) to perform when the instructions are executed individually or collectively by at least one processor (120). The NLP module may, for example, obtain a character recognition result (e.g., data converted from one or more phonemes included in a sentence entered by a user). The NLP module may provide the processing result of the character recognition result to the planner module. The processing result based on character recognition may include an intent, a target device (e.g., information about the target device), a capsule, or a combination thereof.
[0136] According to one example, an NLP module can determine the user's intent by interpreting natural language recognition results (e.g., syntactic analysis and / or semantic analysis). A natural language understanding model can identify the meaning of words extracted from natural language recognition results using linguistic features (e.g., grammatical elements) of morphemes or phrases, and determine the user's intent based on the identified meaning of the words and / or other parameters (e.g., domains or categories associated with the words). Syntactic analysis may include the act of dividing user input (e.g., user text input) into grammatical units (e.g., words, phrases, and / or morphemes) and identifying the grammatical elements possessed by the divided units. Semantic analysis may be performed through semantic matching, rule matching, and / or formula matching. Here, data transformed from one or more phonemes may represent one or more words included in the sentence entered by the user, and / or tokens for each of one or more words. Intent may be data used by the natural language platform to generate a plan. Intent may include goals and / or parameters. Goals may be used in the planner module to specify the final objective of the plan. Parameters may be values input to one or more actions included in the plan in the planner module.
[0137] FIG. 8 is a diagram illustrating an exemplary operation of a multimodal AI model (800) that performs image editing according to one embodiment.
[0138] The multimodal AI model in this disclosure can generate data that is difficult to obtain using only a single type of data by considering and processing or analyzing at least two different types of data (or modalities, e.g., text and images) together. The multimodal AI model can identify complex information and / or patterns by combining different types of data (e.g., images, text, or audio). By integrating multiple types of data, the multimodal AI model can improve the accuracy of predictive models or compensate for data deficiencies by reducing noise. By analyzing the sources of data, the multimodal AI model can support decision-making in dynamic environments by more accurately understanding the context. Therefore, the multimodal AI model can generate desired types of results by fusing various types of data sources.
[0139] According to one example, a multimodal AI model for image editing can be implemented by utilizing natural language processing (NLP) technology and computer vision technology for image analysis (hereinafter referred to as 'image analysis technology'). For example, the multimodal AI model in the present disclosure may be an AI model capable of simultaneously understanding and processing data of different forms, such as text and images.
[0140] According to one example, NLP technology is a technology that enables electronic devices, such as computers or smartphones (e.g., the electronic device (101) of FIG. 1), to understand and generate natural language. For example, NLP technology can be utilized for the analysis, understanding, generation, or response of text data. In NLP technology, sentiment analysis, machine translation, text generation, or question answering systems may be used. Sentiment analysis can analyze the sentiments embedded in natural language, such as text, and classify them as positive, negative, or neutral (e.g., I am really happy: positive). Machine translation can translate text consisting of one or more languages into a specified language (e.g., apple -> apple). Text generation can generate natural language sentences corresponding to a given input. For example, if the text "Please correct this word. Appple" is input, it can generate a result such as "Yes, I will correct the word. The corrected word is Apple. I have corrected an error containing three p's." Q&A can, for example, generate appropriate answers to a given question or provide search results (e.g., What is the weather like today? -> Today's weather is sunny).
[0141] According to one example, image analysis technology is a technology in which an electronic device (101), such as a computer or a smartphone, analyzes an image, such as a photograph or video, to recognize and understand visual information. For example, image analysis technology can be utilized to recognize features such as objects, scenes, or activities contained in an image. Image analysis technology can be utilized for image classification, segmentation, or image generation. Image classification can, for example, classify a target image into a predetermined category (e.g., cat photo -> cat, dog photo -> dog). Segmentation can, for example, analyze a target image at the pixel level to distinguish the object to which each pixel belongs (e.g., obtain only pixels corresponding to a dog within a photo). Image generation can, for example, generate a new image based on predetermined information (e.g., context data) or edit and modify a target image (e.g., draw a dog that knows the sky -> generate a picture).
[0142] Referring to FIG. 8, a multimodal AI model (800) for image editing can generate a prompt (805) based on two different types of heterogeneous data (e.g., text and image).
[0143] According to one example, a prompt design component (810) (e.g., the prompt design component (721) of FIG. 7) can generate a prompt requesting editing of the original image based on an analysis result (801) and the original image (803). For example, the analysis result (801) may include analysis results for one or more images to be edited. For example, the analysis result (801) may include information obtained through the application, such as conversation content, information obtained by analyzing a selection screen, editing history requested by another person, personal profile information, or context information obtained by analyzing information about the person participating in the conversation. The prompt design component (810) can generate a prompt (805) based on the analysis result (801) and the original image (803), such as “Erase the person shown in the photo and modify the erased part to blend in with the surroundings,” and transmit it to an AI model (820). The prompt design component (810) can be transmitted to the AI model (820) by, for example, a prompt (805) generated based on the analysis result (801) and the original image (803), and by masking the object to be modified so that the target object to be modified can be identified.
[0144] An AI model (820) (e.g., a generative AI model (730) of FIG. 7) can edit an original image (803) according to a prompt (805) generated by a prompt design component (810). For example, the AI model (820) can erase a person requested to be deleted from the original image (803) and generate an edited image (807) in which the erased space is modified to blend in with the surroundings.
[0145] FIG. 9a, FIG. 9b, or FIG. 9c is a drawing for illustrating that an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment selects a specific area in a captured image.
[0146] Referring to FIG. 9a, FIG. 9b, or FIG. 9c, the electronic device (101) can output an image of a screen displayed on a display in response to a user's request, or switch to a state where a desired area on the screen can be selected.
[0147] For example, an electronic device (101) can capture a screen and output the captured image through a display in response to the pressing of a single key or a number of keys forming a specific combination. For example, the electronic device (101) can capture a screen and output the captured image through a display in response to the input of a speech command, such as “screen capture,” after an assistant (e.g., Bixby) is called. When the captured image is output in one of the two suggested ways, the electronic device (101) can display a function selection window (920) on the captured image (910). The function selection window (920) may include function identifiers (e.g., icons or emoticons) that allow selecting a function to process the captured image. For example, the function selection window (920) may include a region selection identifier (921) for selecting a desired region in the captured image (910) (see FIG. 9a). After selecting the region selection identifier (921), the user can select a desired region in the captured image (910).
[0148] For example, the electronic device (101) may display a menu window (930) in response to a swipe gesture toward the center of the screen from one side of the screen (910) (e.g., the left side of the screen) (see FIG. 9b). The menu window (930) may include function identifiers (e.g., icons or emoticons) that allow selecting a function to handle the corresponding screen (910). For example, the menu window (930) may include a region setting identifier (931) (e.g., icons or emoticons) for selecting a desired region on the screen (910) (see FIG. 9b). The user may select the region setting identifier (931). When the region setting identifier (931) is selected by the user, the electronic device (101) may display a layer (940) to cover the top of the displayed screen (910) of the display (see FIG. 9c). The user can select a desired portion of area (952) on the layer (940) by adjusting the area selection indicator (951) on the layer (940). When the desired portion of area (952) is selected by the user, the electronic device (101) can display a function identifier for processing the image of the selected area (952).
[0149] FIGS. 10a to 10e are drawings for exemplarily illustrating an operation of editing an image included in a capture screen in an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1).
[0150] The exemplary operation for image editing in FIGS. 10a to 10e assumes a situation in which an electronic device (101) is chatting with multiple participants via an interactive application.
[0151] Referring to FIG. 10a, an electronic device (101) may provide a chat service consisting of text messages or images with multiple participants in a conversation window (1000a) provided by an interactive application. For example, one of the participants, participant D (1010), may send three photos (1013, 1015, 1017) to be edited along with a message (1011) saying “Please delete everyone except us” and display them in the conversation window (1000a). The three photos (1013, 1015, 1017) displayed in the conversation window (1000a) by participant D (1010) may include at least one of the participants in the conversation. For example, one of the participants, Participant B (1020), may send a photo (1023) along with a message (1021) saying “I’m sending too” and display it in the conversation window (1000a). The photo (1023) displayed in the conversation window (1000a) by Participant B (1020) may include images of Participants A, C, and D among the conversation participants. For example, one of the participants, Participant C (1030), may send an image (1033) along with a message (1031) saying “My photo too” and display it in the conversation window (1000a). The photo (1033) displayed in the conversation window (1000a) by Participant C (1030) may include images of Participants A, B, D, and E among the conversation participants.
[0152] As described above, the user who received messages (1011, 1021, 1031) and photos (1013, 1015, 1017, 1023, 1033) from participant D (1010), participant B (1020), and participant C (1030) sent a message (1041) saying “Okay” through the chat window (1000a).
[0153] Referring to FIG. 10b, the electronic device (101) can acquire a captured image (1000b) of a conversation window (1000b). The electronic device (101) can select an area in the captured image (1000b) that includes photos (1013, 1015, 1017, 1023, 1033) sent by Participant D (1010), Participant B (1020), and Participant C (1030). The electronic device (101) can display a selection window that includes a function identifier (e.g., share, edit) that allows selecting an action to process the photos (1013, 1015, 1017, 1023, 1033) included in the selected area.
[0154] Referring to FIG. 10c, the electronic device (101) can download original images (1013, 1015, 1017, 1023, 1033) corresponding to photos included in the selection area when the user selects an identifier for editing from among the identifiers included in the selection window. For example, the electronic device (101) can download and store first to third original images (1013, 1015, 1017) corresponding to three photos sent by participant D (1010), a fourth original image (1023) corresponding to one photo sent by participant B (1020), and a fifth original image (1033) corresponding to one photo sent by participant C (1030).
[0155] Referring to FIG. 10d or FIG. 10e, the electronic device (101) can obtain images of the faces of the participants (Participants A, B, C, D, E) based on profile data corresponding to the participants (Participants A, B, C, D, E) participating in the conversation in the first to fifth photos (1013, 1015, 1017, 1023, 1033) downloaded and stored based on the message “Please delete everyone other than us” sent by Participant D (1010). The electronic device (101) can use the obtained images of the faces of the participants (Participants A, B, C, D, E) to determine the images (1060) of the person excluding the participants (Participants A, B, C, D, E) in the first to fifth photos (1013, 1015, 1017, 1023, 1033). The electronic device (101) can generate an edited image (1013') by deleting the images (1060) of the remaining people, excluding the images (1050a) of Participant A and the images (1050b) of Participant B, from a first photo (1013) displayed on a screen for editing, for example. The electronic device (101) can generate an edited image (1015') by deleting the images (1060) of the remaining people, excluding the images (1050a) of Participant A, the images (1050b) of Participant B, and the images (1050c) of Participant C, from a second photo (1015) that is not displayed on a screen but is downloaded and stored, for example. The electronic device (101) can generate an edited image (1017') in which the images (1060) of the remaining people are deleted, excluding the image (1050b) of participant B and the image (1050c) of participant C from the third photo (1017) that is downloaded and stored but not displayed on the screen.The electronic device (101) can generate an edited image (1023') by deleting the images (1060) of the remaining people, excluding the images of Participant A (1050a), Participant C (1050c), and Participant D (1050d), from a fourth photo (1023) that is downloaded and stored but not displayed on the screen. The electronic device (101) can generate an edited image (1033') by deleting the images (1060) of the remaining people, excluding the images of Participant A (1050a), Participant B (1050b), Participant D (1050d), and Participant E (1050e), from a fifth photo (1033) that is downloaded and stored but not displayed on the screen.
[0156] As described above, the electronic device (101) can generate edited images (1015', 1017', 1023', 1033') in which the modifications requested by one specific participant (e.g., participant D (1010)) among the conversation participants (e.g., participants A, B, C, D, E) for the modification of the photo are applied in common to the photos of other participants (e.g., participants A, B, C, E).
[0157] According to one example, when a request for modification of a photo is received by each of the conversation participants (e.g., participants A, B, C, D, E), the electronic device (101) can generate edited images that modify the photo provided by each participant to reflect the request of that participant. In this case, the electronic device (101) may be able to provide edited images that reflect unique requests for each participant.
[0158] According to one example, if a request for photo modification is received from at least two of the conversation participants (e.g., participants A, B, C, D, E), the electronic device (101) can generate edited images that apply the requests of at least two participants to all participants' photos in common.
[0159] FIGS. 11a to 11d are drawings for exemplarily illustrating an editing operation of an image to be edited in an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1).
[0160] Referring to FIGS. 11a through 11d, the electronic device (101) can download original photos (1110a, 1120a, 1130a, 1140a, 1150a) for editing. For example, the five photos (1110a, 1120a, 1130a, 1140a, 1150a) downloaded by the electronic device (101) may be family photos (e.g., people A, B, C, D, E). For example, among the five photos (1110a, 1120a, 1130a, 1140a, 1150a), the first photo (1110a) may include A, B, and C among the family members. For example, among the five photos (1110a, 1120a, 1130a, 1140a, 1150a), the second photo (1120a) may include A, C, and D among the family members. For example, among the five photos (1110a, 1120a, 1130a, 1140a, 1150a), the third photo (1130a) may include A and B among the family members. For example, among the five photos (1110a, 1120a, 1130a, 1140a, 1150a), the fourth photo (1140a) may include B and C among the family members. For example, among the five photos (1110a, 1120a, 1130a, 1140a, 1150a), the fifth photo (1150a) may include A, B, D, and E among the family members.
[0161] According to one example, if editing (e.g., skin tone correction) of A and B among family members is requested, the electronic device (101) can generate primary edited photos (1110b, 1120b, 1130b, 1140b, 1150b) with skin tone correction for A and / or B included in each of the first to fifth photos (1110a, 1120a, 1130a, 1140a, 1150a) including A and B (see FIG. 11b).
[0162] According to one example, if an electronic device (101) is requested to edit C among the family members (e.g., red-eye correction), it can generate secondary edited photos (1110c, 1120c, 1130c, 1140c, 1150c) in which red-eye is corrected in the image of C included in each of the edited photos (1120b, 1140b) that include C among the primary edited photos (1110b, 1120b, 1130b, 1140b, 1150b) (see FIG. 11c).
[0163] According to one example, if a request is made to disable editing (e.g., skin tone correction) on B among the family members, the electronic device (101) can generate third edited photos (1110d, 1120d, 1130d, 1140d, 1150d) that restore the image of B included in each of the edited photos (1110c, 1130c, 1140c, 1150c) that include B among the second edited photos (1110c, 1120c, 1130c, 1140c, 1150c) to a state before applying the skin tone correction that was applied to the image of B (see FIG. 11d).
[0164] According to one example, the electronic device (101) can identify a specific object (e.g., family) for editing in an image to be edited based on predetermined information. For example, the electronic device (101) can extract information such as photos from registration information in which a group is classified as family in contacts registered in a specific application, such as a phone application, and can identify the image of the family included in the image to be edited based on feature information obtained by analyzing the extracted information. For example, the image to be edited may be original photos (1110a, 1120a, 1130a, 1140a, 1150a), first edited photos (1110b, 1120b, 1130b, 1140b, 1150b), or second edited photos (1110c, 1120c, 1130c, 1140c, 1150c) described with reference to FIGS. 11a to 11d.
[0165] According to one example, the electronic device (101) can identify a specific object (e.g., family) for editing in an image to be edited based on predetermined information. For example, the electronic device (101) can extract information such as a profile picture of a friend whose group is classified as family from a list of friends registered in a specific application (e.g., a text application, a social media application) that provides an online service for building a network of relationships between people, and can identify an image of the family included in the image to be edited based on feature information obtained by analyzing the extracted information. For example, the images to be edited may be the original photographs (1110a, 1120a, 1130a, 1140a, 1150a), first edited photographs (1110b, 1120b, 1130b, 1140b, 1150b), or second edited photographs (1110c, 1120c, 1130c, 1140c, 1150c) described with reference to FIGS. 11a to 11d.
[0166] FIG. 12 is a drawing for illustrating an example of editing an image included in a web browser in an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment.
[0167] Referring to FIG. 12, the electronic device (101) can specify a specific area (1213) on the screen (1210) of a web browser by means of an area selection indicator (1211). For example, the user can adjust the area selection indicator (1211) to specify a specific area (1213) to include an image of Dabotap on the screen (1210) of the web browser. The electronic device (101) can download the image of Dabotap included in the specific area (1213).
[0168] The electronic device (101) may display a selection window (1215) for processing an image of the Dabotap pagoda in a specific area (1213) when a specific area (1213) is selected by an area selection indicator (1211). For example, the selection window (1215) displayed by the electronic device (101) may include function identifiers (shared identifier (1217), edit identifier (1219)) corresponding to supported functions.
[0169] When one of the function identifiers (sharing identifier (1217), editing identifier (1219)) included in the selection window (1215) is selected, the electronic device (101) can share or edit the downloaded Dabotap image.
[0170] FIG. 13 illustrates the configuration of an artificial system inside or outside an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment. The device (1300) may include the following configuration.
[0171] According to one embodiment, applications and services (1310) may refer to applications and services running on a device. An application (e.g., the application (146) of FIG. 2) is a program for performing a specific task and can be operated by a user running the application. A service is one of the programs for performing a specific task and may refer to running applications necessary for the system regardless of whether the user runs them. For example, applications and services (1310) may include an assistant (e.g., Bixby).
[0172] According to one embodiment, a framework (1320) (e.g., middleware (144) of FIG. 2) provides a set of structured code necessary for application operation and may include commonly available modules, reusable modules, classes, interfaces, etc.
[0173] According to one embodiment, the Cross Application Actions Platform (1322) may provide the ability to determine the application required for the Assistant to perform an action, and to connect or suggest actions between the Assistant and the application or between different applications. For example, it may operate an application to respond to a user request through the Assistant (e.g., Bixby), or operate multiple applications sequentially or simultaneously to respond to a user request. Content required for this process may be shared between applications. Additionally, the Cross Application Actions Platform (1322) may provide the ability to share and integrate the actions and content of the Assistant and the application with systems across the platform, such as the Assistant, widgets, and controls.
[0174] According to one embodiment, the personal data engine (1324) can enable the assistant to recognize the personal situation and operate.
[0175] For example, after a user launches an application using an assistant (e.g., Bixby), the personal information engine (1324) may request the assistant to perform actions on behalf of the assistant in the application or system. Additionally, the personal information engine (1324) may transmit personal data held by the application to the assistant to create a semantic index and enable the assistant to search for information held by the application. For example, the personal information engine (1324) may request the assistant to retrieve personal schedule information from a calendar application and share it with others using a mail application or a messaging application. For example, the personal information engine (1324) may request the assistant to retrieve information that can predict the relationship between the user and the subject (e.g., information such as family relationships) from an address book application and share it so that it can be used in an image editing application.
[0176] According to one embodiment, a semantic index is an important concept in text analysis and can be used to extract and understand the meaning of words or phrases within a text. Through the semantic index, the electronic device (101) can identify the topic or context of the text. For example, the electronic device can utilize the semantic index for various natural language processing tasks such as text classification, search, and sentiment analysis. Based on the semantic index, the electronic device (101) can understand stored data and find information more quickly. The electronic device (101) can define and vectorize relationships regarding the content of an application, personalized information, and other information. Vectors can be arranged or mapped as numbers placed close to each other to indicate similarity. Vectors can be stored in a multidimensional space where semantically similar data points are clustered together in the vector space, and the electronic device (101) can process extensive search queries using the vectors.
[0177] According to one embodiment, the orchestration (1326) can coordinate and manage an AI model for operation (e.g., the AI model (800) of FIG. 7). For example, the orchestration (1326) can determine whether to use an AI model within the device or an AI model from an external device to perform an AI operation requested by an application and service (1310). Additionally, the orchestration (1326) can combine and use an AI model within the device and an AI model from an external device depending on the operation. According to one embodiment, the orchestration (1326) can operate in a single module with a cross-application action platform (1322). For example, the orchestration (1326) can make a decision to have the editing of the image performed in response to an editing request either on an on-device AI model (1328) or on an AI model provided on an external device (1340).
[0178] According to one embodiment, the on-device AI model (1328) may be at least one AI model implemented within the device. For example, the AI model may include a large language model (LLM), a large vision model (LVM), or a large action model (LAM). The AI model may be used for natural language understanding, image analysis, behavior prediction, etc. The AI model may include at least one adapter. For example, the adapter may be implemented as a low-rank adaptation (LoRA) as a module for efficiently fine-tuning the model. According to one embodiment, the electronic device (101) may use the AI model to adjust weights for specific information and may reflect personalized information when analyzing or generating results through the AI model. For example, the electronic device (101) may use LoRA to represent updates to a pre-trained weight matrix for adapting to a specific task as a low-rank matrix. For example, the electronic device (101) can learn by replacing four weight matrices in the self-attention module of the Transformer model and two weight matrices in the Multi-Layer Perceptron (MLP) module with low-rank matrices, and instead of updating the entire weight matrix, use a method of updating weights using two low-rank matrices that are multiplied by the input and then added by coordinate.
[0179] According to one embodiment, the Hardware Abstraction layer (HAL) (1330) is a layer that abstracts the interface between hardware and software, and can enable communication between the operating system and the hardware. Through this, the operating system can operate in a consistent manner without being dependent on the hardware. For example, among the hardware, the processor (processor (120) of FIG. 1) may include a main processor (e.g., main processor (121) of FIG. 1) or an auxiliary processor (e.g., auxiliary processor (123) of FIG. 1). An example of the auxiliary processor (123) is a processor with enhanced security features, which may further include a secure processor unit (SPU) that can be used to safely process sensitive data.
[0180] According to one embodiment, an external device (1340) may include a server (e.g., the server (108) of FIG. 1) or a third AI solution (1344). For example, the server (108) may include an AI model such as LLM or LVM, and may process an AI function requested by the device (1300) using the AI model and transmit the processing result to the device (1300).
[0181] According to one embodiment, the 3rd AI Solution (1344) may include an AI model or AI service operated by another operator in addition to the AI model implemented in the device (1300) and the external device (1340) by the manufacturer. For example, the orchestration (1326) may control the selection of one of the device (1300), the external device (1340), or the 3rd AI Solution (1344) when selecting the AI model required for operation.
[0182] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains.
[0183] As an example, the electronic device (101) may include a display. The electronic device (101) may include a memory (313) comprising one or more storage media for storing instructions. The electronic device (101) may include at least one processor (311) comprising a processing circuit. When the instructions are executed individually or collectively by the at least one processor (311), the electronic device (101) may be caused to perform at least one operation. The at least one operation may include an operation of determining a specific area on the execution screen of the first application (331). The at least one operation may include an operation of obtaining context information reflecting an editing intention by considering information obtained through the use of the first application (311). The at least one operation may include an operation of downloading an original image associated with the determined specific area. The above at least one operation may include the operation of executing the second application (333) to generate an edited image by editing the downloaded original image based on the acquired context information.
[0184] As an example, when the above instructions are executed individually or collectively by at least one processor (311), the electronic device (101) may be caused to perform an operation of transmitting the edited image generated by the second application (333) so that it can be used in the first application (331).
[0185] As an example, when the above instructions are executed individually or collectively by at least one processor (311), the electronic device (101) may be caused to perform an operation to delete the downloaded original image and / or the edited image.
[0186] For example, when the above instructions are executed individually or collectively by at least one processor (311), the electronic device (101) may be caused to perform an operation to obtain the context information reflecting the editing intention of the counterpart or user by analyzing the image and / or text included in the dialog box according to the use of the first application (331), if the first application (331) is an interactive application.
[0187] For example, when the above instructions are executed individually or collectively by at least one processor (311), the electronic device (101) may be caused to perform an operation to obtain the context information reflecting the editing intention of the counterpart or user by analyzing the conversation content included in the conversation window and / or the profile information of the counterpart participating in the chat by the conversation window, if the first application (331) is an interactive application.
[0188] As an example, when the above instructions are executed individually or collectively by at least one processor (311), the electronic device (300) may be caused to perform the operation of sharing the edited image through the dialog box.
[0189] As an example, when the above instructions are executed individually or collectively by at least one processor (311), the electronic device (300) may be caused to perform the operation of capturing the execution screen of the first application (331) and displaying the captured image.
[0190] As an example, when the above instructions are executed individually or collectively by at least one processor (311), the electronic device (300) may be caused to perform the operation of selecting the specific area in the captured image.
[0191] As an example, when the above instructions are executed individually or collectively by at least one processor (311), the electronic device (300) may be caused to perform an operation of displaying a layer so as to be superimposed on the execution screen of the first application (331).
[0192] As an example, when the above instructions are executed individually or collectively by at least one processor (311), the electronic device (300) may be caused to perform the operation of selecting the specific area included in the execution screen on the layer.
[0193] For example, when the above instructions are executed individually or collectively by at least one processor (311), the electronic device (300) may be caused to perform an operation to edit a specific original image included in the plurality of original images by reflecting an editing intention based on the acquired context information when downloading the plurality of original images.
[0194] As an example, when the above instructions are executed individually or collectively by at least one processor (311), the electronic device (300) may be caused to perform an operation of editing one or more of the remaining original images among the plurality of original images to reflect the editing intention of the specific original image.
[0195] For example, the operation method of the electronic device (101) may include an operation of determining a specific area on the execution screen of the first application (331). The operation method may include an operation of obtaining context information that reflects an editing intention by considering information obtained through the use of the first application (311). The operation method may include an operation of downloading an original image related to the determined specific area. The operation method may include an operation of creating an edited image by executing the second application (333) and editing the downloaded original image based on the obtained context information.
[0196] As an example, the above operation method may include an operation of transmitting the edited image generated by the second application (333) so that it can be used in the first application (331).
[0197] As an example, the above operation method may include an operation to delete the downloaded original image and / or the edited image.
[0198] As an example, if the first application (331) is a conversational application, the operation of obtaining the context information may include analyzing an image and / or text included in a conversation window resulting from the use of the first application (331) to obtain the context information that reflects the editing intention of the counterparty or user.
[0199] As an example, if the first application (331) is a conversational application, the operation of obtaining the context information may include analyzing the conversation content included in the conversation window resulting from the use of the first application (331) and / or the profile information of the counterpart participating in the chat via the conversation window to obtain the context information that reflects the counterpart or user's editing intention.
[0200] As an example, the above operation method may include the operation of sharing the edited image through the dialog box.
[0201] As an example, the operation of determining the specific area may include capturing the execution screen of the first application (331) and displaying the captured image, and selecting the specific area from the captured image.
[0202] As an example, the operation of determining the specific area may include the operation of displaying a layer so as to overlap the execution screen of the first application (331) and the operation of selecting the specific area included in the execution screen on the layer.
[0203] As an example, the operation of generating the edited image may include, when downloading a plurality of original images, an operation of editing a specific original image included in the plurality of original images by reflecting an editing intention based on the acquired context information.
[0204] As an example, the operation of generating the edited image may include the operation of editing one or more of the remaining original images among the plurality of original images to reflect the editing intention of the specific original image.
[0205] For example, a recording medium may store instructions that can be read by a computer. When these instructions are executed by at least part of at least one processor (311) included in the electronic device (300), they may cause the electronic device (300) to perform at least one operation. The at least one operation may include an operation of determining a specific area on the execution screen of the first application (331). The at least one operation may include an operation of obtaining context information that reflects an editing intention by considering information obtained through the use of the first application (311). The at least one operation may include an operation of downloading an original image related to the determined specific area. The at least one operation may include an operation of creating an edited image by executing the second application (333) and editing the downloaded original image based on the obtained context information.
[0206] The electronic devices according to the various embodiments disclosed in this document may be devices of various types or forms. The electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or consumer electronics devices. The electronic devices according to the embodiments of this document are not limited to the devices described above.
[0207] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as "coupled" or "connected" to another (e.g., 2nd) component, with or without the terms "functionally" or "communicationly," it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0208] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0209] Various embodiments of this document may be implemented as software (e.g., a program) comprising one or more instructions stored in a storage medium (e.g., memory (313)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (331)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term is used when data is stored semi-permanently in the storage medium and It does not distinguish cases where it is temporarily stored.
[0210] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0211] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device (101), display; Memory (313) comprising one or more storage media for storing instructions; and It includes at least one processor (311) including a processing circuit, and When the above instructions are executed individually or collectively by at least one processor (311), the electronic device (101) is caused to perform at least one operation, and The above at least one operation is, An action to determine a specific area on the execution screen of the first application (331); An operation to obtain context information that reflects an editing intention by considering information obtained through the use of the first application (311); The operation of downloading the original image related to the specific area determined above; and The operation of executing the second application (333) to generate an edited image by editing the downloaded original image based on the acquired context information. An electronic device (101) including 2. In Paragraph 1, When the above instructions are executed individually or collectively by at least one processor (311), the electronic device (101) is made to: The operation of transmitting the edited image generated by the second application (333) so that it can be used in the first application (331); and The action of deleting the original image downloaded above and / or the edited image above. An electronic device (101) that causes to perform.
3. In Paragraph 1 or 2, When the above instructions are executed individually or collectively by at least one processor (311), the electronic device (101) is made to: If the first application (331) is an interactive application, the operation of obtaining the context information that reflects the editing intention of the counterparty or user by analyzing the image and / or text included in the conversation window according to the use of the first application (331); and The action of sharing the above edited image through the above dialog box An electronic device (101) that causes to perform.
4. In Paragraph 1 or 2, When the above instructions are executed individually or collectively by at least one processor (311), the electronic device (101) is made to: If the first application (331) is a conversational application, the operation of obtaining the context information reflecting the editing intention of the counterpart or user by analyzing the conversation content included in the conversation window resulting from the use of the first application (331) and / or the profile information of the counterpart participating in the chat via the conversation window; and The action of sharing the above edited image through the above dialog box An electronic device (101) that causes to perform.
5. In any one of paragraphs 1 through 4, When the above instructions are executed individually or collectively by at least one processor (311), the electronic device (300) is: An operation to capture the execution screen of the first application (331) and display the captured image; and The action of selecting the specific area in the above-mentioned captured image An electronic device (101) that causes to perform.
6. In any one of paragraphs 1 through 4, When the above instructions are executed individually or collectively by at least one processor (311), the electronic device (300) is: An operation to display a layer so as to overlap the execution screen of the first application (331); and The operation of selecting the specific area included in the execution screen on the above layer. An electronic device (101) that causes to perform.
7. In any one of paragraphs 1 through 6, When the above instructions are executed individually or collectively by at least one processor (311), the electronic device (300) is: When downloading multiple original images, an operation to edit a specific original image included in the multiple original images by reflecting an editing intention based on the acquired context information; and The operation of editing one or more of the remaining original images among the plurality of original images to reflect the above editing intention. An electronic device (101) that causes to perform.
8. In the method of operating the electronic device (101), An action to determine a specific area on the execution screen of the first application (331); An operation to obtain context information that reflects an editing intention by considering information obtained through the use of the first application (311); The operation of downloading the original image related to the specific area determined above; and The operation of executing the second application (333) to generate an edited image by editing the downloaded original image based on the acquired context information. A method of operation including 9. In Paragraph 8, The operation of transmitting the edited image generated by the second application (333) so that it can be used in the first application (331); and The action of deleting the original image downloaded above and / or the edited image above. A method of operation including 10. In Paragraph 8 or 9, The operation of obtaining the above context information is, If the first application (331) is an interactive application, the operation of obtaining the context information that reflects the editing intention of the counterparty or user by analyzing the image and / or text included in the conversation window according to the use of the first application (331); and The action of sharing the above edited image through the above dialog box A method of operation including 11. In Paragraph 8 or 9, The operation of obtaining the above context information is, If the first application (331) is a conversational application, the operation of obtaining the context information reflecting the editing intention of the counterpart or user by analyzing the conversation content included in the conversation window resulting from the use of the first application (331) and / or the profile information of the counterpart participating in the chat via the conversation window; and The action of sharing the above edited image through the above dialog box A method of operation including 12. In any one of paragraphs 8 through 11, The operation of determining the aforementioned specific area is, An operation to capture the execution screen of the first application (331) and display the captured image; and The action of selecting the specific area in the above-mentioned captured image A method of operation including 13. In any one of paragraphs 8 through 11, The operation of determining the aforementioned specific area is, An operation to display a layer so as to overlap the execution screen of the first application (331); and The operation of selecting the specific area included in the execution screen on the above layer. A method of operation including 14. In any one of paragraphs 8 through 13, The operation of generating the above edited image is, When downloading multiple original images, an operation to edit a specific original image included in the multiple original images by reflecting an editing intention based on the acquired context information; and The operation of editing one or more of the remaining original images among the plurality of original images to reflect the editing intention of the specific original image mentioned above. A method of operation including 15. In a recording medium storing computer-readable instructions, When the above instructions are executed by at least part of at least one processor (311) included in the electronic device (300), the electronic device (300) causes at least one operation to be performed, and The above at least one operation is: An action to determine a specific area on the execution screen of the first application (331); An operation to obtain context information that reflects an editing intention by considering information obtained through the use of the first application (311); The operation of downloading the original image related to the specific area determined above; and The operation of executing the second application (333) to generate an edited image by editing the downloaded original image based on the acquired context information. A recording medium including