Electronic device and method for transferring and displaying data between applications by using artificial intelligence model in electronic device

WO2026205744A1PCT designated stage Publication Date: 2026-10-01SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/001748
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-06-10
Filing Date
2026-01-29
Publication Date
2026-10-01

Smart Images

  • Figure KR2026001748_01102026_PF_FP_ABST
    Figure KR2026001748_01102026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to an embodiment comprises: a display; a memory storing instructions; and at least one processor, wherein the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to: identify a first area of a first user interface of a first application on the basis of a first user input for the first user interface while the first user interface is displayed on the display; obtain first data corresponding to the first area, the first data including at least one of UI component unit data corresponding to the first area, UI space data corresponding to the first area, first metadata corresponding to the first area, first image data corresponding to the first area, or first text data corresponding to the first area; obtain first extraction data including first main image data, from the first data through an artificial intelligence model on the basis of first context information for the first data; display a second user interface of a second application on the display; obtain second data corresponding to the second application on the basis of a second user input for applying the first extraction data to the second user interface, the second data including second metadata of the second application and second image data and / or second text data corresponding to the second user interface; obtain second extraction data extracted from the second data, through the artificial intelligence model on the basis of second context information for the second data, the second extraction data including one or more UI components applicable to the second user interface and additional information for changing the first extraction data to be applicable to the second data; obtain at least one content item through the artificial intelligence model on the basis of the first extraction data and the second extraction data; and display the first main image data and the at least one content item on the second user interface through the display. Various other embodiments may be possible.
Need to check novelty before this filing date? Find Prior Art

Description

Method for data transfer and display between electronic devices and applications using artificial intelligence models in electronic devices

[0001] Various embodiments of the present disclosure relate to a method for data transfer and display between an electronic device and applications using an artificial intelligence model in the electronic device.

[0002] Thanks to the remarkable advancements in information and communication technology and semiconductor technology, the distribution and use of various electronic devices are increasing rapidly. Electronic devices are being developed to allow users to carry them around and communicate. The term "electronic device" may refer to a device that performs specific functions according to an installed program, such as mobile communication terminals, tablet PCs, smartphones, wearable electronic devices, video / audio devices, desktop / laptop computers, or vehicle navigation systems.

[0003] The electronic device can execute at least one program (e.g., at least one application). Depending on the execution of the application, the electronic device can display a user interface (UI) corresponding to the application on the display.

[0004] An electronic device can acquire data through the user interface of an application and transmit it to the user interface of another application. The electronic device can transmit data of a defined data type that is mutually usable between applications. For example, an electronic device can copy text data from the screen of a note application and transmit it to the screen of a schedule management application, and the schedule management application can utilize the text data. However, if the application's data is not of a defined data type for mutual use between applications, the electronic device may experience inconvenience because it cannot transmit the data to another application, or it must convert the data into a data type usable by the other application before transmitting it. For example, even if a user intends to transmit data to another application while the application's user interface is displayed on the electronic device, the data may not be transmitted if it is not of a defined data type for mutual use. When data from an application is transmitted from an electronic device to another application, if the transmitted data does not match the context of the data in the other application, the efficiency of utilizing the transmitted data may be low.

[0005] According to the present disclosure, when a portion of an image is selected in a user interface (e.g., a first user interface) of an application (e.g., a first application) in an electronic device, an image of the selected area and data associated with the selected area are obtained, and the image of the selected area and data associated with the selected area are converted into an image and data that can be displayed in a user interface (e.g., a second user interface) of another application (e.g., a second application) and transmitted to the second application, thereby providing convenience to the user.

[0006] According to the present disclosure, when a portion of an image shown in a first user interface of a first application in an electronic device is selected, an image and data according to a first context corresponding to the selected portion of the image and data may be obtained, and the obtained image and data may be converted into an image and at least one content item according to a second context of a second user interface of a second application and displayed in the second user interface of the second application.

[0007] An electronic device according to one embodiment of the present disclosure comprises a display, a memory for storing commands, and at least one processor. When the commands are executed individually or collectively by the at least one processor, the electronic device may identify a first area of ​​a first user interface based on a first user input for the first user interface while displaying a first user interface of a first application on the display. When the commands are executed individually or collectively by the at least one processor, the electronic device may acquire first data corresponding to the first area. The first data may include at least one of UI component unit data corresponding to the first area, UI spatial data corresponding to the first area, first metadata corresponding to the first area, first image data corresponding to the first area, or first text data corresponding to the first area. When the above commands are executed individually or collectively by the at least one processor, the electronic device may be enabled to obtain first extracted data including first key image data from the first data through an artificial intelligence model based on first context information regarding the first data. When the above commands are executed individually or collectively by the at least one processor, the electronic device may be enabled to display a second user interface of a second application on the display. When the above commands are executed individually or collectively by the at least one processor, the electronic device may be enabled to obtain second data corresponding to the second application based on a second user input for applying the first extracted data to the second user interface.The second data may include second metadata of the second application, second image data corresponding to the second user interface, and / or second text data. When the commands are executed individually or collectively by the at least one processor, the electronic device may obtain second extracted data extracted from the second data through the artificial intelligence model based on second context information regarding the second data. The second extracted data may include one or more UI components applicable to the second user interface and additional information for modifying the first extracted data to be applicable to the second data. When the commands are executed individually or collectively by the at least one processor, the electronic device may obtain at least one content item through the artificial intelligence model based on the first extracted data and the second extracted data. When the above commands are executed individually or collectively by the at least one processor, the electronic device may display the first main image data and the at least one content item on the second user interface through the display.

[0008] A method for data transfer and display between applications using an artificial intelligence model in an electronic device according to one embodiment of the present disclosure may include an operation of identifying a first area of ​​a first user interface based on a first user input for the first user interface while displaying a first user interface of a first application on a display of the electronic device. The method may include an operation of acquiring first data corresponding to the first area. The first data may include at least one of UI component unit data corresponding to the first area, UI spatial data corresponding to the first area, first metadata corresponding to the first area, first image data corresponding to the first area, or first text data corresponding to the first area. The method may include an operation of acquiring first extracted data including first key image data from the first data through the artificial intelligence model based on first context information for the first data. The method may include an operation of displaying a second user interface of a second application through the display. The above method may include an operation of obtaining second data corresponding to the second application based on a second user input for applying the first extracted data to the second user interface. The second data may include second metadata of the second application, second image data corresponding to the second user interface, and / or second text data. The above method may include an operation of obtaining second extracted data extracted from the second data through the artificial intelligence model based on second context information regarding the second data.The second extracted data may include one or more UI components applicable to the second user interface and additional information for modifying the first extracted data to be applicable to the second data. The method may include an operation of obtaining at least one content item through the artificial intelligence model based on the first extracted data and the second extracted data. The method may include an operation of displaying the first main image data and the at least one content item on the second user interface through the display.

[0009] In a non-transient storage medium storing commands according to one embodiment of the present disclosure, the commands are configured to cause the electronic device to perform at least one operation when executed by the electronic device, wherein the at least one operation may include an operation of identifying a first area of ​​the first user interface based on a first user input for the first user interface while displaying a first user interface of a first application on the display of the electronic device. The at least one operation may include an operation of acquiring first data corresponding to the first area. The first data may include at least one of UI component unit data corresponding to the first area, UI spatial data corresponding to the first area, first metadata corresponding to the first area, first image data corresponding to the first area, or first text data corresponding to the first area. The at least one operation may include an operation of acquiring first extracted data including first key image data from the first data through an artificial intelligence model based on first context information for the first data. The at least one operation may include an operation of displaying a second user interface of a second application through the display. The at least one operation may include an operation of obtaining second data corresponding to the second application based on a second user input for applying the first extracted data to the second user interface. The second data may include second metadata of the second application, second image data corresponding to the second user interface, and / or second text data. The at least one operation may include an operation of obtaining second extracted data extracted from the second data through the artificial intelligence model based on second context information regarding the second data.The second extracted data may include one or more UI components applicable to the second user interface and additional information for modifying the first extracted data to be applicable to the second data. The at least one operation may include an operation of obtaining at least one content item through the artificial intelligence model based on the first extracted data and the second extracted data. The at least one operation may include an operation of displaying the first main image data and the at least one content item on the second user interface through the display.

[0010] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment.

[0011] FIG. 2 is a block diagram of an electronic device according to one embodiment.

[0012] FIG. 3 is a flowchart illustrating data transfer and display operations between applications using an artificial intelligence model in an electronic device according to one embodiment.

[0013] FIG. 4 is a diagram illustrating operations for obtaining first extracted data from first data based on first user input for a first user interface of a first application according to one embodiment.

[0014] FIG. 5 is a diagram showing operations for applying first extracted data obtained through a first user interface of a first application according to one embodiment to a second user interface.

[0015] FIG. 6 is a diagram illustrating operations for arranging multiple content items generated through a third artificial intelligence model according to one embodiment based on priority and displaying them on a display.

[0016] FIG. 7a is a diagram illustrating an example of obtaining first extracted data including first image data from first data of a first area of ​​a map application execution screen based on user input to a map application execution screen of a map application according to one embodiment.

[0017] FIG. 7b is a diagram showing an example of applying first extracted data obtained through a map application screen according to one embodiment to a messenger application screen.

[0018] FIG. 8 is a diagram illustrating an example of obtaining image data that has been edited from the first image data of a first area according to one embodiment.

[0019] FIG. 9 is a diagram showing an example of applying first extracted data, including first main image data obtained through a movie application screen according to one embodiment, to a messenger application screen.

[0020] FIG. 10 is a diagram showing an example of data transfer and display between a calendar application, a map application, and a messenger application according to one embodiment.

[0021] FIG. 11 is a diagram illustrating an example in which a change in some of the tracking data occurs during data transmission and display between a calendar application, a map application, and a messenger application according to one embodiment.

[0022] FIG. 12 is a diagram illustrating an example of providing information of an application capable of applying first data, including first image data of a first area obtained through a user interface of an application according to one embodiment.

[0023] FIG. 13 is a diagram illustrating an example of providing information indicating the previous usage history of a first data including first image data corresponding to a first area obtained through a user interface of an application according to one embodiment.

[0024] FIG. 14 is a block diagram of a generative artificial intelligence (AI) system according to one embodiment.

[0025] FIG. 15 is a block diagram of an AI Framework according to one embodiment.

[0026] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to one embodiment. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through the server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).

[0027] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.

[0028] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0029] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).

[0030] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0031] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0032] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0033] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0034] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).

[0035] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0036] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0037] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0038] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0039] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0040] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0041] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0042] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).

[0043] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.

[0044] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).

[0045] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0046] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0047] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0048] In the following detailed description, reference numbers in the drawings may be assigned identically or omitted for configurations that can be easily understood through prior embodiments, and detailed descriptions thereof may also be omitted. An electronic device according to one embodiment disclosed in this document may be implemented by selectively combining configurations of different embodiments, and a configuration of one embodiment may be replaced by a configuration of another embodiment. For example, it should be noted that the present invention is not limited to specific drawings or embodiments.

[0049] FIG. 2 is a block diagram of an electronic device according to one embodiment.

[0050] Referring to FIG. 2, an electronic device (201) according to one embodiment (e.g., the electronic device (101) of FIG. 1) may include a processor (220), memory (230), a display (260), and a communication circuit (290). The electronic device (201) according to one embodiment is not limited thereto and may be configured to include various additional components or to exclude some of the components. The electronic device (201) according to one embodiment may further include all or part of the electronic device (101) shown in FIG. 1.

[0051] A processor (220) according to one embodiment may be composed of one or more processors. In this case, the one or more processors may include a general-purpose processor such as a CPU, AP, DSP (digital signal processor), a graphics-dedicated processor such as a GPU, VPU (vision processing unit), or an artificial intelligence-dedicated processor such as an NPU (neural processing unit) or AI accelerator. The one or more processors may be controlled to process input data according to a predefined operation rule or at least one artificial intelligence model stored in memory (230). Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0052] Predefined behavioral rules or artificial intelligence models can be created through learning. Here, being created through learning means that a basic artificial intelligence model is trained using multiple learning data by a learning algorithm, thereby creating a predefined behavioral rule or artificial intelligence model configured to perform a desired characteristic (or objective). Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.

[0053] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values ​​and performs neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), or deep Q-networks, but are not limited to the examples mentioned above.

[0054] A processor (220) according to one embodiment can perform overall control operations of an electronic device (201) and can perform operations (or methods) for data transfer and display between applications using the artificial intelligence model of the present disclosure.

[0055] A processor (220) according to one embodiment may execute an application and display a user interface (UI) (e.g., an application screen) corresponding to the application being executed on a display (260). An application according to one embodiment may refer to a map application, a message application, a schedule management application, a movie application, a web browser application, or other applications executable on an electronic device (201), and may not be limited to a specific application. A processor (220) according to one embodiment may execute multiple applications and display multiple user interfaces corresponding to the multiple applications on a display (260).

[0056] A processor (220) according to one embodiment may receive user input (e.g., first user input) for a first user interface while displaying a user interface (e.g., first user interface or first application screen) of an application (e.g., first application) on a display (260). The first user input according to one embodiment may be a gesture input. For example, the gesture input may include a touch input, a long press input, a drag input after touch, a drawing input, or other gesture inputs. For example, the first user input may be an input for selecting (or extracting or copying) a rectangular area, a round area, a lasso-shaped area, or an area within a closed curve drawn by the user, or for selecting (or extracting or copying) an area of ​​a designated area unit of the first user interface (e.g., a street area or an alley area in the case of a map application screen).

[0057] A processor (220) according to one embodiment may identify (or select) at least some area (e.g., a first area) corresponding to the user input in the first user interface based on the first user input for the first user interface. A processor (220) according to one embodiment may change the boundary of the first area to include additional information that follows (or is associated with) the information included in the first area (e.g., text followed by text included in the first area or an image followed by an image included in the first area) according to the information included in the first area.

[0058] A processor (220) according to one embodiment may include image data of a first area (e.g., first image data) (e.g., partial screen image data corresponding to the first area in an application screen image) based on the identification of a first area of ​​a first user interface, and may obtain data (e.g., first data) corresponding to the first area (or image data of the first area).

[0059] A processor (220) according to one embodiment may display information of at least one application among applications included in an electronic device (201) capable of applying (or utilizing) the first data including the first image data, based on the acquisition of first data corresponding to a first area including the first image data of a first area. When a user inputs a query regarding information of at least one application among applications included in an electronic device (201) capable of applying (or utilizing) the first data including the first image data of a first area when the processor (220) according to one embodiment acquires and provides (displays or outputs) information of at least one application among applications included in the electronic device (201) capable of applying (or utilizing) the first data including the first image data of a first area, upon the acquisition of first data corresponding to a first area including the first image data of a first area, the processor (220) may acquire and provide (display or output) information of at least one application among applications included in the electronic device (201) capable of applying (or utilizing) the first data including the first image data.

[0060] A processor (220) according to one embodiment may display information indicating the usage history (or history) of the first data including the first image data based on the acquisition of first data including the first image data of the first region.

[0061] According to one embodiment, the first data may include at least one of UI component unit data for one or more UI (user interface) components corresponding to the first area, UI spatial data corresponding to the first area, metadata of an app associated with the first area (e.g., first metadata), and the first image data and / or first text data. According to one embodiment, the processor (220) may extract and obtain UI component unit data corresponding to the first area selected by the user. According to one embodiment, the processor (220) may extract and obtain information regarding one or more independent UI components included in the UI, such as a button, a figure, or a menu, unique attributes of each of the one or more UI components, characteristics of each of the one or more UI components (e.g., whether they are active and / or visible), and / or subsequent actions, sub-items, and / or functions based on user interaction (e.g., actions) with the one or more UI components. For example, one or more UI components corresponding to the first area include UI components that map text or images, such as Android-based TextView or ImageView, or web page-based <figure>or <menu>It may include a UI component representing an inserted image or menu, or a UI component corresponding to an HTML Tag. One or more UI components corresponding to the first area according to one embodiment may include a menu component, a button component, a checkbox component, a slider component, and / or an image component.

[0062] A processor (220) according to one embodiment can extract and obtain UI spatial data corresponding to a first area selected by a user. A processor (220) according to one embodiment can extract and obtain spatial relationship information between UI components (e.g., hierarchy and / or depth between UI components) as UI spatial data (422). For example, the hierarchy and / or depth between UI components may include the Android view structure or the HTML DOM structure.

[0063] A processor (220) according to one embodiment may obtain first metadata corresponding to a first area from metadata used in a first application. A processor (220) according to one embodiment may extract and obtain metadata (e.g., first metadata) of an app (e.g., first app) corresponding to a first area selected by a user. A processor (220) according to one embodiment may extract and obtain first metadata of a first area that is not displayed on the screen (e.g., important data or designated data) according to the characteristics of the first app corresponding to the first area selected by a user. The first metadata according to one embodiment may include descriptions, location information, place information, and / or traffic information associated with the first area included in the first application. For example, the first metadata corresponding to the first area may include additional information or descriptions regarding the first image of the first area and / or the first data corresponding to the first area. In one embodiment, if the first user interface is a map application screen, the first area may include a part of a map image, and the first metadata corresponding to the first area may include location information (e.g., GPS information), facility information, and / or other additional information or descriptions regarding the part of the map image and / or part of the map data. In one embodiment, if the first user interface is a photo (or image) application screen, the first area may include a part of a photo, and the first metadata corresponding to the first area may include EXIF ​​(exchangeable image file format), photo analysis information (e.g., object, background, and / or scene information), and / or other additional information or descriptions regarding the part of the photo and photo data.In the case where the first user interface according to one embodiment is a calendar application screen, the first area may include a part of the calendar, and the first metadata corresponding to the first area may include information obtained from an electronic device (201) and / or information obtained from an external server (e.g., 108) and / or other additional information or descriptions regarding the part of the calendar and data (e.g., memo, schedule, custom time) of the part of the calendar (date corresponding to the part of the calendar).

[0064] A processor (220) according to one embodiment may extract and obtain image data (e.g., first image data) and / or first text data corresponding to a first area selected by a user. For example, the processor (220) may obtain first image data corresponding to the first area selected by the user, a main context (e.g., scene and / or object) corresponding to the first image data corresponding to the first area selected by the user, and text data existing within the first image data obtained by performing OCR on the first image data corresponding to the first area selected by the user. A processor (220) according to one embodiment may obtain text included in the first image data by performing OCR (optical character recognition) on the first image data. A processor (220) according to one embodiment may obtain the text of "Submit" by performing OCR if the first image data of the first area includes a button with the text "Submit" written on it. A processor (220) according to one embodiment may obtain image data that has been edited from the first image data.

[0065] A processor (220) according to one embodiment may obtain image data that has been edited from the first image data. A processor (220) according to one embodiment may identify at least one cell corresponding to the first image data among the cells according to the grid structure of the first user interface using cell information (e.g., area information) according to the grid structure of the first user interface, and may identify at least one first cell among the at least one cell corresponding to the first image data in which important data (data corresponding to the first image data and / or the first context corresponding to the first data (e.g., UI component, action connected to the UI component, sub-item, and / or function, or first metadata)) is mapped or more than a specified number of data items are mapped, and at least one second cell in which data is not mapped or unimportant data is mapped or less than a specified number of data items are mapped. A processor (220) according to one embodiment can use an artificial intelligence model to identify the importance and / or association according to the first image data and / or the first context corresponding to the first data for the data included (or associated or corresponding) in each of the cells corresponding to the first image data among the cells according to the grid structure of the first user interface. A processor (220) according to one embodiment can use an artificial intelligence model to identify a first cell with high importance and / or association and a second cell with low importance and / or association compared to the first cell among the cells corresponding to the first image data, based on the importance and / or association of the data included (or associated or corresponding) in each of the cells corresponding to the first image data.

[0066] A processor (220) according to one embodiment can obtain edited image data by editing the image data in such a way that the portion of the first image data corresponding to at least one first cell becomes a high-quality image and the portion of the first image data corresponding to at least one second cell becomes a low-quality image.

[0067] A processor (220) according to one embodiment may obtain first context information through context analysis of first data including first image data through a first artificial intelligence model (e.g., a multimodal artificial intelligence model). The first context information according to one embodiment may include an association between UI components, a mapping table for actions connected to the first metadata and UI components (e.g., an action to navigate to an HTML page), sub-items (e.g., a dropdown menu, a multi-choice menu), and / or functions (e.g., a JavaScript function), and text-type output information obtained (or output) from the first artificial intelligence model by providing the first artificial intelligence model with a prompt to check the boundary of an important area within the first image.

[0068] A processor (220) according to one embodiment may preprocess (e.g., generalize, normalize, or standardize) the first data into a form that can be used in at least one other application or another platform (e.g., iOS, Android OS, or Web) or another format. A processor (220) according to one embodiment may normalize the text contained in the first data to obtain normalized text data, and use the normalized text data to obtain normalized first data corresponding to the first data. For example, text normalization, which normalizes the text contained in the first data, may mean converting text representing UI elements into standard or general text so that UI elements between different platforms (Android, iOS, Web, etc.) can be compared or converted. For example, text normalization may include actions such as unifying uppercase and lowercase letters (e.g., CheckBox → checkbox), removing unnecessary symbols, unifying abbreviations or syllabic words (e.g., btn → button), translation or multilingual processing, and mapping names that differ by platform to a common expression. A processor (220) according to one embodiment may obtain text data (e.g., normalized text data) indicating which class of iOS the Android OS-based "Checkbox" corresponds to when the first data includes an Android OS-based "Checkbox". A processor (220) according to one embodiment may obtain normalized first data corresponding to the first data by replacing the text included in the first data with normalized text, or obtain normalized first data corresponding to the first data by including normalized text information mapped to the text included in the first data in the first data.

[0069] A processor (220) according to one embodiment can extract partial image data (important image data) (e.g., first main image data) corresponding to (matching) the first context information from the first image data using first context information obtained through a first artificial intelligence model, and obtain partial data (important data) (e.g., first extracted data) corresponding to (matching) the first context information from the first data (or normalized first data).

[0070] A processor (220) according to one embodiment can optimize the first image data and the first data by redefining the region of interest using first context information obtained through the first artificial intelligence model, extracting partial image data (important image data) (e.g., first main image data) corresponding to (matching) the first context information from the first image data, and obtaining partial data (important data) (e.g., first extracted data) corresponding to (matching) the first context information from the first data (or normalized first data). When using the first main image and the first extracted data through optimization, it is possible to prevent the processing time from taking longer when the artificial intelligence model understands and generates data compared to when the first image data and the first data are used as is. The first image data and the first data can be optimized by selecting the first main image and the first extracted data from the first image data and the first data by considering UI features using the first context information obtained through the first artificial intelligence model according to one embodiment. For example, in the case of a map application, the processor (220) can divide major facilities or traffic information based on distance or divide specific areas into a rectangular grid. Even if the user has set a large area, the area can be redefined with data (text, entire image) corresponding to such criteria to reduce the size of the data while removing data of low importance.

[0071] A processor (220) according to one embodiment can display a second user interface of a second application on a display (260).

[0072] A processor (220) according to one embodiment may receive a user input (e.g., a second user input) for applying the first main image data and / or the first extracted data to the second user interface. The second user input according to one embodiment may be a gesture input. For example, the gesture input may include a touch input, a long press input, a touch-and-drag input, a drawing input, or other gesture inputs. For example, the gesture input may include a touch input, a long press input, a touch-and-drag input, a drawing input, or other gesture inputs. For example, the second user input may be an input (e.g., paste) for applying the first main image and / or the first extracted data, which reflects the first context, to the second user interface.

[0073] A processor (220) according to one embodiment may acquire second data corresponding to a second application based on a second user input. A processor (220) according to one embodiment may acquire second data corresponding to a second application to identify the features of the second application in order to apply (e.g., paste) the first main image data and / or the first extracted data to the second user interface. The second data according to one embodiment may include second metadata of the second application, a second image corresponding to the second user interface, and / or second text (text obtained through OCR of the second image). The second data according to one embodiment may further include information about applications (or software) installed on the electronic device (201) (e.g., application list and / or application category). The second data according to one embodiment may further include a category of the second application (e.g., messenger category, map category, payment category, or other categories) or user data stored in association with the second application.

[0074] If the first data including the first image data of the first area according to one embodiment is extracted mainly from data that can be extracted from the first user interface of the first application, the second data corresponding to the second application may include data that is helpful or necessary when converting the first main image and / or the first extracted data to fit the second user interface of the second application. In the case where the second application is a messenger application, the processor (220) according to one embodiment may include the second data corresponding to the second application, such as an image of a conversation screen currently being executed, text obtained by performing OCR on the image of the conversation screen, conversation history through the conversation screen, and UI components for configuring the second user interface of the second application.

[0075] A processor (220) according to one embodiment may obtain second context information by performing context analysis on second data corresponding to a second application through a second artificial intelligence model (e.g., a multimodal artificial intelligence model). The second context information according to one embodiment may include information for data (e.g., second extracted data) that can be used (or extracted) from the second data or must be additionally (or additionally) obtained in order to apply the first extracted data to the second user interface (e.g., data types available in the second application and information of a third application among the applications included in the electronic device that can obtain additional information).

[0076] A processor (220) according to one embodiment may obtain second extracted data from second data through a second artificial intelligence model (e.g., a multimodal artificial intelligence model) based on second context information. The second extracted data according to one embodiment may include a data type that is usable (or applicable or available) in a second user interface (e.g., one or more UI components usable on a second application screen or a data type that can be displayed on a messenger application screen (e.g., text, image, video, or URL)) and additional information for changing a first main image and / or the first extracted data so that it can be applied to the second data.

[0077] A processor (220) according to one embodiment may extract (or acquire) text data, image data, and conversation history data as data types that can be utilized (or applied or used) to apply (e.g., to convert so that it can be displayed) the first main image and / or first extracted data from the second data corresponding to the messenger application when the second application is a messenger application. A processor (220) according to one embodiment may extract (or acquire) text data, location data, and GPS data as data types that can be utilized (or applied or used) to apply (e.g., to convert so that it can be displayed) the second data corresponding to the map application when the second application is a map application.

[0078] A processor (220) according to one embodiment identifies a third application to be used for obtaining additional information among the applications of the electronic device (201) and can obtain additional information through the third application. A processor (220) according to one embodiment can execute the third application in the background using information of the third application capable of obtaining additional information and can extract (or obtain) additional information to change the first main image data and / or the first extracted data so that it can be applied to the second data through the third application. A processor (220) according to one embodiment can obtain additional information regarding the arrival information of a bus or subway, etc., by searching for stop information in a traffic information application to change the stop information of the first extracted data so that it can be applied to the second data of the message application when additional information regarding arrival information of a bus or subway, etc., is needed in the conversation history of the message application when applying (e.g., copying) the first extracted data of the map application to the message application screen of the message application. According to one embodiment, the processor (220) may omit the operation of obtaining additional information.

[0079] A processor (220) according to one embodiment may generate (or acquire) at least one content item through a third artificial intelligence model (e.g., a generative artificial intelligence model) based on first extracted data and second extracted data. A generative artificial intelligence model according to one embodiment may include an LLM that generates text and / or an LVM that generates images. At least one content item according to one embodiment may include at least one content item in a form usable in a second application (e.g., at least one content item that is displayed and accessible in a second user interface of the second application). At least one content item according to one embodiment may include a plurality of content items. The plurality of content items according to one embodiment may have different data types.

[0080] A processor (220) according to one embodiment can identify the priority of multiple content items when at least one content item includes multiple content items. A processor (220) according to one embodiment can identify the priority of multiple content items according to how much more meaningful the data is from a user's perspective. A processor (220) according to one embodiment can identify the priority of multiple content items in order of high similarity between text summary information for each of the multiple content items and second context information of the second application. For example, the processor (220) can obtain multiple first vector values ​​for each of the multiple content items by embedding text summary information for each of the multiple content items into a vector through a text encoder, obtain a second vector value by embedding second context information of the second application into a vector, obtain cosine similarity values ​​between each of the first vector values ​​and the second vector value, and identify the priority of the multiple content items according to the order of high similarity values ​​by comparing the cosine similarity values. The processor (220) according to one embodiment can sort (or reorder) the display order of the multiple content items according to the priority of the multiple content items.

[0081] A processor (220) according to one embodiment may display a first extracted image and at least one content item on a second user interface through a display (260). When displaying the first main image data and a plurality of content items, the processor (220) according to one embodiment may display the plurality of content items arranged according to the priority of the plurality of content items. The plurality of content items may include text, images, icons or link information that can be connected to other applications. A processor (220) according to one embodiment may apply at least one content item to the second application (e.g., by transferring or inputting into an input window) based on user input selecting at least one content item. If at least one content item is not selected by the user, the processor (220) according to one embodiment may regenerate (or re-acquire) at least one content item through a third artificial intelligence model (e.g., a generative artificial intelligence model) based on the first extracted data and the second extracted data. A processor (220) according to one embodiment can ensure that when at least one content item is regenerated, at least one regenerated content item does not overlap with at least one previously created content item.

[0082] A processor (220) according to one embodiment may store first image data, first data, first context information, first main image, first extracted data, second data, second context information, and second extracted data obtained from the time a first area is identified based on a first user input for a first user interface of a first application until a first main image and at least one content item are displayed on a second user interface. A processor (220) according to one embodiment may obtain third main image data and third extracted data for a second area identified by a third user input based on a third user input for a second user interface in a manner similar to the method of obtaining first main image data and first extracted data for a first area. A processor (220) according to one embodiment may acquire third data corresponding to the third application and fourth extracted data from the third data in a manner similar to the method of acquiring second data corresponding to the second application based on a fourth user input for applying third main image data and third extracted data to the third user interface of the third application. When acquiring at least one content item to be applied to the third user interface, the processor (220) according to one embodiment may acquire at least one other content item to be applied to the third user interface by using the first extracted data and second extracted data and the third extracted data and fourth extracted data stored in relation to the first application and the second application.

[0083] A processor (220) according to one embodiment may track and store data obtained during the data transfer process between the first application, the second application, and the third application after data transfer and display operation from the first application to the second application, and data transfer and display from the second application to the third application. A processor (220) according to one embodiment may use (or apply) the data tracked and stored during the data transfer process between the first application, the second application, and the third application in the next application data transfer process. A processor (220) according to one embodiment may reflect the data tracked and stored during the data transfer process from the first application to the N-1st application when data is transferred from the N-1st application to the Nth application. A processor (220) according to one embodiment may reflect the content of the changed data in the user interface of the application associated with the changed data when data is sequentially transferred from the first application to the Nth application and a change in some data is detected on the tracking path while the tracked and stored data exists. A processor (220) according to one embodiment may maintain data consistency based on the traceability of the data obtained during the data transfer process between applications.

[0084] A memory (230) according to one embodiment (e.g., memory (130) of FIG. 1) can store various data used by at least one component of an electronic device (201) (e.g., processor (220), display (260) and / or communication circuit (290)). The data may include, for example, input data or output data for software (e.g., software module or program (140)) and related commands.

[0085] A memory (230) according to one embodiment may include at least one artificial intelligence model (e.g., a first artificial intelligence model (231), a second artificial intelligence model (232), and a third artificial intelligence model (233)). The first artificial intelligence model (231), the second artificial intelligence model (232), and the third artificial intelligence model (233) according to one embodiment may be implemented as a single artificial intelligence model or as an artificial intelligence model in which at least a portion is integrated. Although the first artificial intelligence model (231), the second artificial intelligence model (232), and the third artificial intelligence model (233) have been described separately in the description of the present disclosure, the first artificial intelligence model, the second artificial intelligence model, and the third artificial intelligence model may be implemented as a single artificial intelligence model or may include additional artificial intelligence models, and the number of artificial intelligence models or the method of implementation (e.g., integrated implementation or separate implementation) may be optional for those skilled in the art. According to one embodiment, the first artificial intelligence model (231), the second artificial intelligence model (232), and the third artificial intelligence model (233) may include an LLM, LVM, or multimodal artificial intelligence model. According to one embodiment, the memory (230) may include a plurality of applications (e.g., a first application (235), a second application (236) to N applications (237)). According to one embodiment, the first application (235), the second application (236) to N applications (237) may refer to a map application, a message application, a schedule management application, a movie application, a web browser application, or other applications executable on the electronic device (201), and may not be limited to a specific application.

[0086] A memory (230) according to one embodiment may store commands (or instructions) that enable a processor (220) to perform operations (or methods) for data transfer and display between applications using the artificial intelligence model of the present disclosure.

[0087] A display (260) according to one embodiment (e.g., the display (160) of FIG. 1) can display various information based on the control of a processor (220). A display (260) according to one embodiment can display user interfaces (e.g., application execution screens) associated with performing operations (or methods) for data transfer and display between applications using an artificial intelligence model. According to one embodiment, the display (260) can be implemented in the form of a touch screen. When the display (260) is implemented in the form of a touch screen together with an input module, it can display various information generated according to the user's touch operation.

[0088] A communication circuit (290) according to one embodiment (e.g., communication module (190) of FIG. 1) may include a wireless communication module (e.g., cellular module, Wi-Fi (wireless-fidelity) module, Bluetooth module, or NFC (near field communication) module). A communication circuit (290) according to one embodiment may communicate with an external electronic device (not shown) (e.g., server (108) of FIG. 1 or electronic device (102) of FIG. 1) through a third application. A communication circuit (290) according to one embodiment may receive additional information from an external electronic device (e.g., electronic device (102) of FIG. 1 or server (108) of FIG. 1) based on the control of a processor (220).

[0089] According to one embodiment, the electronic device (201) is not limited to the configuration shown in FIG. 2 and may be configured to include various additional components. According to one embodiment, the electronic device (201) may further include a sound output module (not shown) (e.g., the sound output module (150) of FIG. 1) and may output sound for at least one content item through the sound output module.

[0090] In one embodiment, the main components of the electronic device were described through the electronic device (201) of FIG. 2. However, in various embodiments, not all components illustrated in FIG. 2 are essential components, and the connection relationships of the main components of the electronic device (201) described above through FIG. 2 may be changed according to various embodiments.

[0091] An electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1 or the electronic device (201) of FIG. 2) may include a display (e.g., the display module (160) of FIG. 1 or the display (260) of FIG. 2), a memory for storing commands (e.g., the memory (130) of FIG. 1 or the memory (230) of FIG. 2), and at least one processor (e.g., the processor (130) of FIG. 1 or the processor (230) of FIG. 2). When the commands are executed individually or collectively by the at least one processor, the electronic device may identify a first area of ​​a first user interface based on a first user input for the first user interface while displaying a first user interface of a first application on the display. When the commands are executed individually or collectively by the at least one processor, the electronic device may obtain first data corresponding to the first area. The first data may include at least one of UI component unit data corresponding to the first area, UI spatial data corresponding to the first area, first metadata corresponding to the first area, first image data corresponding to the first area, or first text data corresponding to the first area. When the commands are executed individually or collectively by the at least one processor, the electronic device may be enabled to obtain first extracted data including first key image data from the first data through an artificial intelligence model based on first context information regarding the first data. When the commands are executed individually or collectively by the at least one processor, the electronic device may be enabled to display a second user interface of a second application on the display.When the above commands are executed individually or collectively by the at least one processor, the electronic device may obtain second data corresponding to the second application based on a second user input for applying the first extracted data to the second user interface. The second data may include second metadata of the second application, second image data corresponding to the second user interface, and / or second text data. When the above commands are executed individually or collectively by the at least one processor, the electronic device may obtain second extracted data extracted from the second data through the artificial intelligence model based on second context information regarding the second data. The second extracted data may include one or more UI components applicable to the second user interface and additional information for modifying the first extracted data to be applicable to the second data. When the above commands are executed individually or collectively by the at least one processor, the electronic device may obtain at least one content item through the artificial intelligence model based on the first extracted data and the second extracted data. When the above commands are executed individually or collectively by the at least one processor, the electronic device may display the first main image data and the at least one content item on the second user interface through the display. The at least one content item may include sub-functions and / or actions of the first user interface.

[0092] According to one embodiment, when the commands are executed individually or collectively by the at least one processor, the electronic device may obtain normalized text data by normalizing the text included in the first extracted data. When the commands are executed individually or collectively by the at least one processor, the electronic device may obtain normalized first extracted data corresponding to the first extracted data using the normalized text data. For example, when the commands are executed individually or collectively by the at least one processor, the electronic device may obtain normalized first extracted data in which the first extracted data has been preprocessed (e.g., generalized, normalized, or standardized) into a form that can be used in at least one other application or another platform (e.g., iOS, Android OS, or Web) or another format. For example, text normalization, which normalizes the text included in the first extracted data, may mean converting text representing UI elements into standard or general text so that UI elements can be compared or converted between different platforms (Android, iOS, Web, etc.). For example, text normalization may include operations such as unifying uppercase and lowercase letters, removing unnecessary symbols, unifying abbreviations or synonymous terms, translation or multilingual processing, and mapping platform-specific names to common expressions. For example, when the above commands are executed individually or collectively by the at least one processor, the electronic device may obtain normalized first extracted data corresponding to the first extracted data by replacing the text contained in the first extracted data with normalized text, or obtain normalized first extracted data corresponding to the first extracted data by including normalized text information mapped to the text contained in the first extracted data in the first extracted data.

[0093] The commands according to one embodiment, when executed individually or collectively by the at least one processor, may enable the electronic device to identify the priority of the plurality of content items when the at least one content item includes the plurality of content items. The commands, when executed individually or collectively by the at least one processor, may enable the electronic device to display the plurality of content items based on the priority.

[0094] The commands according to one embodiment, when executed individually or collectively by the at least one processor, may enable the electronic device to identify an application among the applications of the electronic device to be used for obtaining the additional information. The commands, when executed individually or collectively by the at least one processor, may enable the electronic device to obtain the additional information through the application.

[0095] When the commands according to one embodiment are executed individually or collectively by the at least one processor, the electronic device may obtain a third principal image and third extracted data for third data corresponding to a second area of ​​the second user interface based on a third user input for the second user interface. When the commands are executed individually or collectively by the at least one processor, the electronic device may display a third user interface of a third application on the display. When the commands are executed individually or collectively by the at least one processor, the electronic device may obtain fourth extracted data for fourth data corresponding to a third area of ​​the third user interface based on a fourth user input for applying the third extracted data to the third user interface. When the above commands are executed individually or collectively by the at least one processor, the electronic device may obtain at least one other content item through the artificial intelligence model based on the first extracted data, the second extracted data, the third extracted data, and the fourth extracted data. When the above commands are executed individually or collectively by the at least one processor, the electronic device may display the third main image data and the at least one other content item on the third user interface through the display.

[0096] When the commands according to one embodiment are executed individually or collectively by the at least one processor, the electronic device may be able to display information of at least one other application to which the first data can be applied through the display when a first area of ​​the first user interface is identified.

[0097] When the commands according to one embodiment are executed individually or collectively by the at least one processor, the electronic device may display information indicating the previous usage history of the first data when a first area of ​​the first user interface is identified.

[0098] When the commands according to one embodiment are executed individually or collectively by the at least one processor, the electronic device may perform optical character recognition (OCR) on the first image data corresponding to the first area to obtain text corresponding to the first image data corresponding to the first area.

[0099] According to one embodiment, the one or more UI components include a menu component, a button component, a checkbox component, a slider component, and / or an image component, and the UI component may include an action, a sub-item, and / or a function associated with the UI component.

[0100] According to one embodiment, the first metadata may include a description, location information, place information, and / or traffic information associated with the first area and included in the first application.

[0101] FIG. 3 is a flowchart illustrating data transfer and display operations between applications using an artificial intelligence model in an electronic device according to one embodiment.

[0102] Referring to FIG. 3, a processor (e.g., processor (120) of FIG. 1 or processor (220) of FIG. 2) of an electronic device according to one embodiment (e.g., electronic device (101) of FIG. 1 or electronic device (201) of FIG. 2) can perform at least one of 310 operations to 370 operations.

[0103] In operation 310, a processor (220) according to one embodiment may identify (or select) a first area corresponding to the first user input in the first user interface of the first application (235) displayed on the display (260) based on the first user input. The first user input according to one embodiment may be an input for selecting (or extracting or copying) a rectangular area, a round area, a lasso-shaped area, or an area within a closed curve drawn by the user, or for selecting (or extracting or copying) an area of ​​a designated area unit of the first user interface (e.g., a street area or an alley area in the case of a map application screen). The processor (220) according to one embodiment may change the boundary of the first area to include additional information that follows (or is associated with) the information included in the first area (e.g., text followed by text included in the first area or an image followed by an image included in the first area) according to the information included in the first area.

[0104] In operation 320, a processor (220) according to one embodiment may obtain first data corresponding to the first area (or image data of the first area) including first image data of the first area (e.g., partial screen image data corresponding to the first area in an application screen image) based on the identification of the first area of ​​the first user interface. The first data according to one embodiment may include at least one of UI component unit data for one or more UI (user interface) components corresponding to the first area, UI spatial data corresponding to the first area, metadata of an app associated with the first area (e.g., first metadata), and the first image data and / or first text data.

[0105] A processor (220) according to one embodiment may extract and obtain UI component unit data corresponding to a first area selected by a user. A processor (220) according to one embodiment may extract and obtain information regarding one or more independent UI components included in the UI, such as a button, a figure, or a menu, unique attributes of each of the one or more UI components, characteristics of each of the one or more UI components (e.g., whether they are active and / or visible), and / or subsequent actions, sub-items, and / or functions based on user interaction (e.g., actions) with one or more UI components. For example, one or more UI components corresponding to the first area may include UI components where text or images are mapped, such as an Android-based TextView or ImageView, or web page-based <figure>or <menu>It may include a UI component representing an inserted image or menu, or a UI component corresponding to an HTML Tag. One or more UI components corresponding to the first area according to one embodiment may include a menu component, a button component, a checkbox component, a slider component, and / or an image component.

[0106] A processor (220) according to one embodiment can extract and obtain UI spatial data corresponding to a first area selected by a user. A processor (220) according to one embodiment can extract and obtain spatial relationship information between UI components (e.g., hierarchy and / or depth between UI components) as UI spatial data (422). For example, the hierarchy and / or depth between UI components may include the Android view structure or the HTML DOM structure.

[0107] A processor (220) according to one embodiment may obtain first metadata corresponding to a first area from metadata used in a first application. A processor (220) according to one embodiment may extract and obtain metadata (e.g., first metadata) of an app (e.g., first app) corresponding to a first area selected by a user. A processor (220) according to one embodiment may extract and obtain first metadata of a first area that is not displayed on the screen (e.g., important data or designated data) according to the characteristics of the first app corresponding to the first area selected by a user. The first metadata according to one embodiment may include descriptions, location information, place information, and / or traffic information associated with the first area included in the first application. For example, the first metadata corresponding to the first area may include additional information or descriptions regarding the first image of the first area and / or the first data corresponding to the first area. In one embodiment, if the first user interface is a map application screen, the first area may include a part of a map image, and the first metadata corresponding to the first area may include location information (e.g., GPS information), facility information, and / or other additional information or descriptions regarding the part of the map image and / or part of the map data. In one embodiment, if the first user interface is a photo (or image) application screen, the first area may include a part of a photo, and the first metadata corresponding to the first area may include EXIF ​​(exchangeable image file format), photo analysis information (e.g., object, background, and / or scene information), and / or other additional information or descriptions regarding the part of the photo and photo data.In the case where the first user interface according to one embodiment is a calendar application screen, the first area may include a part of the calendar, and the first metadata corresponding to the first area may include information obtained from an electronic device (201) and / or information obtained from an external server (e.g., 108) and / or other additional information or descriptions regarding the part of the calendar and data (e.g., memo, schedule, custom time) of the part of the calendar (date corresponding to the part of the calendar).

[0108] A processor (220) according to one embodiment may extract and obtain image data (e.g., first image data) and / or first text data corresponding to a first area selected by a user. For example, the processor (220) may obtain first image data corresponding to the first area selected by the user, a main context (e.g., scene and / or object) corresponding to the first image data corresponding to the first area selected by the user, and text data existing within the first image data obtained by performing OCR on the first image data corresponding to the first area selected by the user. A processor (220) according to one embodiment may obtain text included in the first image data by performing OCR (optical character recognition) on the first image data. A processor (220) according to one embodiment may obtain the text of "Submit" by performing OCR if the first image data of the first area includes a button with the text "Submit" written on it. A processor (220) according to one embodiment may obtain image data that has been edited from the first image data. A processor (220) according to one embodiment can identify at least one cell corresponding to a first image data among the cells according to the grid structure of the first user interface using cell information (e.g., area information) according to the grid structure of the first user interface, and can identify at least one first cell among the at least one cell corresponding to the first image data in which important data (data corresponding to the first image data and / or the first context corresponding to the first data (e.g., UI component, action connected to the UI component, sub-item, and / or function, or first metadata)) is mapped or more than a specified number of data items are mapped, and at least one second cell in which data is not mapped or unimportant data is mapped or less than a specified number of data items are mapped.A processor (220) according to one embodiment can use an artificial intelligence model to identify the importance and / or association according to the first image data and / or the first context corresponding to the first data for the data included (or associated or corresponding) in each of the cells corresponding to the first image data among the cells according to the grid structure of the first user interface. A processor (220) according to one embodiment can use an artificial intelligence model to identify a first cell with high importance and / or association and a second cell with low importance and / or association compared to the first cell among the cells corresponding to the first image data, based on the importance and / or association of the data included (or associated or corresponding) in each of the cells corresponding to the first image data.

[0109] A processor (220) according to one embodiment can obtain edited image data by editing the portion of the first image data corresponding to at least one first cell so that the image becomes high-quality image data, and the portion of the first image data corresponding to at least one second cell becomes low-quality image data. A processor (220) according to one embodiment can preprocess (e.g., generalization, normalization, or standardization) the first data into a form that can be used in at least one other application or other format (e.g., iOS, Android OS, or Web). A processor (220) according to one embodiment can obtain normalized text data by normalizing the text included in the first data, and obtain normalized first data corresponding to the first data using the normalized text data. A processor (220) according to one embodiment can obtain text data (e.g., normalized text data) indicating which class of iOS the Android OS-based "Checkbox" corresponds to when the first data includes an Android OS-based "Checkbox".

[0110] In operation 330, a processor (220) according to one embodiment may obtain first key image data corresponding to (matching) the first context information from the first image data through a first artificial intelligence model based on first context information for first data including first image data, and obtain first extracted data corresponding to (matching) the first context information from the first data (or normalized first data). A processor (220) according to one embodiment may obtain first context information through context analysis of first data including first image data through a first artificial intelligence model (e.g., a multimodal artificial intelligence model). The first context information according to one embodiment may include an association between UI components, a mapping table for actions, sub-items, and / or functions connected to the first metadata and UI components, and text-form output information obtained (or output) from the first artificial intelligence model by providing the first artificial intelligence model with a prompt to check the boundary of an important area within the first image.

[0111] In operation 340, a processor (220) according to one embodiment may acquire second data corresponding to the second application based on a second user input for applying the first main image data and the first extracted data to the second user interface of the second application displayed on the display (260). The second user input according to one embodiment may be an input (e.g., paste) for applying the first main image data and the first extracted data, which reflect the first context, to the second user interface. The processor (220) according to one embodiment may acquire second data corresponding to the second application to identify the features of the second application in order to apply the first main image data and the first extracted data to the second user interface (e.g., paste). The second data according to one embodiment may include second metadata of the second application, a second image corresponding to the second user interface, and / or second text (text obtained through OCR of the second image). According to one embodiment, the second data may further include information about applications (or software) installed on the electronic device (201) (e.g., an application list and / or application categories). According to one embodiment, the second data may further include a category of the second application (e.g., a messenger category, a map category, a payment category, or other categories) or user data stored in association with the second application. According to one embodiment, if the first data corresponding to the first area includes first image data of the first area and is extracted mainly from data extractable from the first user interface of the first application, the second data corresponding to the second application may include data that is helpful or necessary when converting the first main image data and / or the first extracted data to fit the second user interface of the second application.In the case where the second application is a messenger application, the processor (220) according to one embodiment may include a second data corresponding to the second application, such as an image of a conversation screen currently being executed, text obtained by performing OCR on the image of the conversation screen, conversation history through the conversation screen, and UI components for configuring a second user interface of the second application.

[0112] In operation 350, a processor (220) according to one embodiment may obtain second extracted data from second data through a second artificial intelligence model (e.g., a multimodal artificial intelligence model) based on second context information for second data corresponding to a second application. A processor (220) according to one embodiment may obtain second context information by performing context analysis on second data corresponding to a second application through a second artificial intelligence model (e.g., a multimodal artificial intelligence model). The second context information according to one embodiment may include information for data (e.g., second extracted data) that can be used (or extracted) from the second data or additionally (or additionally) obtained in order to apply the first extracted data to a second user interface (e.g., data types available in the second application and information of a third application among the applications included in the electronic device that can obtain additional information). According to one embodiment, the second extracted data may include data types that are usable (or applicable or available) in the second user interface (e.g., one or more UI components usable on the second application screen or data types that can be displayed on the messenger application screen (e.g., text, image, video, or URL)) and additional information for changing the first main image and / or the first extracted data so that it can be applied to the second data. According to one embodiment, if the second application is a messenger application, the processor (220) may extract (or obtain) text data, image data, and conversation history data as data types that are usable (or applicable or available) in the second data corresponding to the messenger application to apply the first main image and / or the first extracted data to the messenger application screen (e.g., to convert it so that it can be displayed).A processor (220) according to one embodiment may extract (or acquire) text data, place data, and GPS data as data types that can be utilized (or applied or used) to apply (e.g., to convert for display) to the map application screen (e.g., to convert for display) when the second application is a map application. A processor (220) according to one embodiment may identify a third application among the applications of the electronic device (201) to be used for acquiring additional information, and may acquire additional information through the third application. A processor (220) according to one embodiment may execute the third application in the background using information of the third application capable of acquiring additional information, and may extract (or acquire) additional information through the third application to change the first extracted data so that it can be applied to the second data. According to one embodiment, when applying (e.g., copying) the first extracted data of a map application to the message application screen of a message application, if additional information regarding arrival information of a bus or train is needed in the conversation history of the message application, the processor (220) may search for stop information in a traffic information application to obtain additional information regarding arrival information of a bus or train, etc., in order to change the stop information of the first extracted data so that it can be applied to the second data of the message application. According to one embodiment, the processor (220) may omit the operation of obtaining additional information.

[0113] In operation 360, a processor (220) according to one embodiment may generate (or acquire) at least one content item through a third artificial intelligence model (e.g., a generative artificial intelligence model) based on first extracted data and second extracted data. A generative artificial intelligence model according to one embodiment may include an LLM that generates text and / or an LVM that generates images. At least one content item according to one embodiment may include at least one content item in a form usable in a second application (e.g., at least one content item that is displayable and accessible in a second user interface of the second application). At least one content item according to one embodiment may include a plurality of content items. The plurality of content items according to one embodiment may have different data types. A processor (220) according to one embodiment may identify the priority of the plurality of content items when at least one content item includes a plurality of content items. A processor (220) according to one embodiment can identify the priority of multiple content items based on how much more meaningful the data is from a user's perspective. A processor (220) according to one embodiment can identify the priority of multiple content items in order of high similarity between text summary information for each of the multiple content items and second context information of the second application. For example, the processor (220) can obtain multiple first vector values ​​for each of the multiple content items by embedding text summary information for each of the multiple content items into a vector through a text encoder, obtain second vector values ​​by embedding second context information of the second application into a vector, obtain cosine similarity values ​​between each of the first vector values ​​and the second vector values, and identify the priority of multiple content items in order of high similarity value by comparing the cosine similarity values.A processor (220) according to one embodiment can sort (or reorder) the display order of multiple content items according to the priority of multiple content items.

[0114] In operation 370, a processor (220) according to one embodiment may display a first main image data and at least one content item on a second user interface through a display (260). When displaying the first main image data and a plurality of content items, the processor (220) according to one embodiment may display the plurality of content items arranged according to the priority of the plurality of content items. The plurality of content items may include text, images, actions, icons or link information that can be connected to other applications. The processor (220) according to one embodiment may apply at least one content item to the second application (e.g., forwarding or inputting into an input window) based on user input selecting at least one content item. If at least one content item is not selected by the user, the processor (220) according to one embodiment may regenerate (or reacquire) at least one content item through a third artificial intelligence model (e.g., a generative artificial intelligence model) based on the first extracted data and the second extracted data. Although the first artificial intelligence model (231), the second artificial intelligence model (232), and the third artificial intelligence model (233) have been described separately in the description of the present disclosure, the first artificial intelligence model (231), the second artificial intelligence model (232), and the third artificial intelligence model (233) may be implemented as a single artificial intelligence model or may include additional artificial intelligence models, and the number of artificial intelligence models or the method of implementation (e.g., integrated implementation or separate implementation) may be optional by those skilled in the art. A processor (220) according to one embodiment may ensure that when at least one content item is regenerated, the regenerated at least one content item does not overlap with at least one previously created content item.

[0115] A method for data transfer and display between applications using an artificial intelligence model in an electronic device (101, 201) according to one embodiment of the present disclosure may include an operation of identifying a first area of ​​a first user interface based on a first user input for the first user interface while displaying a first user interface of a first application on a display (160, 260) of the electronic device. The method may include an operation of acquiring first data corresponding to the first area. The first data may include at least one of UI component unit data corresponding to the first area, UI spatial data corresponding to the first area, first metadata corresponding to the first area, first image data corresponding to the first area, or first text data corresponding to the first area. The method may include an operation of acquiring first extracted data including first key image data from the first data through the artificial intelligence model based on first context information for the first data. The method may include an operation of displaying a second user interface of a second application through the display. The above method may include an operation of obtaining second data corresponding to the second application based on a second user input for applying the first extracted data to the second user interface. The second data may include second metadata of the second application, second image data corresponding to the second user interface, and / or second text data. The above method may include an operation of obtaining second extracted data extracted from the second data through the artificial intelligence model based on second context information regarding the second data.The second extracted data may include one or more UI components applicable to the second user interface and additional information for modifying the first extracted data to be applicable to the second data. The method may include an operation of obtaining at least one content item through the artificial intelligence model based on the first extracted data and the second extracted data. The method may include an operation of displaying the first main image data and the at least one content item on the second user interface through the display.

[0116] The method according to one embodiment may include an operation of obtaining normalized text data by normalizing the text included in the first extracted data. The method may include an operation of obtaining normalized first extracted data corresponding to the first extracted data using the normalized text data.

[0117] The method according to one embodiment may include an operation of identifying the priority of the plurality of content items when the at least one content item includes a plurality of content items. The method may include an operation of displaying the plurality of content items based on the priority.

[0118] The method according to one embodiment may include an operation of identifying an application among the applications of the electronic device to be used for obtaining the additional information. The method may include an operation of obtaining the additional information through the application.

[0119] The method according to one embodiment may include an operation of acquiring a third main image and third extracted data for third data corresponding to a second area of ​​the second user interface based on a third user input for the second user interface. The method may include an operation of displaying a third user interface of a third application on the display. The method may include an operation of acquiring fourth extracted data for fourth data corresponding to a third area of ​​the third user interface based on a fourth user input for applying the third extracted data to the third user interface. The method may include an operation of acquiring at least one other content item through the artificial intelligence model based on the first extracted data, the second extracted data, the third extracted data, and the fourth extracted data. The method may include an operation of displaying the third main image data and the at least one other content item on the third user interface through the display.

[0120] The method according to one embodiment may include an operation of displaying information of at least one other application capable of applying the first data through the display when a first area of ​​the first user interface is identified.

[0121] The method according to one embodiment may include an operation of displaying information indicating the previous usage history of the first data when a first area of ​​the first user interface is identified.

[0122] The method according to one embodiment may include the operation of performing OCR on the first image data corresponding to the first region to obtain text included in the first image data corresponding to the first region.

[0123] In the method according to one embodiment, the one or more UI components include a menu component, a button component, a checkbox component, a slider component, and / or an image component, and the UI component may include an action, a sub-item, and / or a function associated with the UI component.

[0124] In the method according to one embodiment, the first metadata may include a description, location information, place information, and / or traffic information associated with the first area and included in the first application.

[0125] FIG. 4 is a diagram illustrating operations for obtaining first extracted data from first data based on first user input for a first user interface of a first application according to one embodiment.

[0126] Referring to FIG. 4, in operation 410, a processor (220) according to one embodiment can identify (or select) an area corresponding to the first user input (414) in the first user interface (412) (e.g., map application execution screen) of the first application (e.g., first app or map application) based on the first user input (414) (e.g., AI smart select input).

[0127] In operation 420, a processor (220) according to one embodiment may obtain data (e.g., first data) for an area selected by a user (e.g., first area). The first data according to one embodiment may include UI component unit data, UI spatial data, metadata of an app, image data, and / or text data.

[0128] A processor (220) according to one embodiment can extract and obtain UI component unit data corresponding to a first area selected by a user (421). A processor (220) according to one embodiment can extract and obtain information regarding one or more independent UI components included in the UI, such as a button, a figure, and a menu, unique attributes of each of the one or more UI components and characteristics of each of the one or more UI components (e.g., whether they are active and / or visible), and / or subsequent actions, sub-items, and / or functions based on user interaction (e.g., actions) with one or more UI components.

[0129] A processor (220) according to one embodiment can extract and obtain UI spatial data corresponding to a first area selected by a user. A processor (220) according to one embodiment can extract and obtain spatial relationship information between UI components (e.g., hierarchy and / or depth between UI components) as UI spatial data (422). For example, the hierarchy and / or depth between UI components may include the Android view structure or the HTML DOM structure.

[0130] A processor (220) according to one embodiment may extract and obtain metadata (e.g., first metadata) of an app (e.g., first app) corresponding to a first area selected by a user (423). A processor (220) according to one embodiment may extract and obtain first metadata of a first area that is not displayed on the screen (e.g., important data or designated data) depending on the characteristics of the first app corresponding to the first area selected by the user. For example, if the first app is a map application, the processor (220) may extract and obtain first metadata including GPS information and / or facility information that is not displayed on the screen.

[0131] A processor (220) according to one embodiment may extract and obtain image data (e.g., first image data) and / or first text data corresponding to a first area selected by a user (424). For example, the processor (220) may obtain first image data corresponding to a first area selected by a user, a key context (e.g., scene and / or object) corresponding to the first image data corresponding to the first area selected by a user, and text data existing within the first image data obtained by performing OCR on the first image data corresponding to the first area selected by a user.

[0132] In operation 430, a processor (220) according to one embodiment may perform context analysis by transmitting UI component unit data, UI spatial data, app metadata, image data, and / or text data for a user-selected area to an artificial intelligence model (multimodal AI model) (e.g., a first artificial intelligence model). A processor (220) according to one embodiment may perform context analysis on the first data through a text encoder (431), an image encoder (432), and / or a multimodal large language model (433). A processor (220) according to one embodiment may obtain first context information as a result of context analysis. The first context information according to one embodiment may include output information in the form of text obtained (or output) from the first artificial intelligence model by providing the first artificial intelligence model with a prompt to check the boundary of an important area within an image (e.g., the first image), metadata and a behavior mapping table (e.g., a mapping table for the first metadata and the behavior (e.g., action, sub-item, and / or function) connected to the UI components).

[0133] In operation 440, a processor (220) according to one embodiment may recognize the UI component entity name of the first data and normalize the first data. A processor (220) according to one embodiment may preprocess (e.g., generalize, normalize, or standardize) the first data into a form that can be used in at least one other application or other format (e.g., iOS, Android OS, or Web) through a text encoder (441) and a multi-modal large language model (LLM) (443). A processor (220) according to one embodiment may normalize the text contained in the first data to obtain normalized text data, and use the normalized text data to obtain normalized first data corresponding to the first data. For example, text normalization, which normalizes the text contained in the first data, may mean converting text representing UI elements into standard or general text so that UI elements can be compared or converted between different platforms (Android, iOS, Web, etc.). For example, text normalization may include standardizing uppercase and lowercase letters (e.g., CheckBox → checkbox), removing unnecessary symbols, standardizing abbreviations or syllabaries (e.g., btn → button), translation or multilingual processing, and mapping platform-specific names to common expressions.

[0134] In operation 450, a processor (220) according to one embodiment may perform first data optimization by redefining a primary region of interest. A processor (220) according to one embodiment may obtain optimized first extracted data by performing optimization of the first data by redefining a region of interest among the first regions using first context information obtained through a first artificial intelligence model. The first extracted data may include partial data (primary data or important data or representative data) corresponding to (matching) the first context information extracted from the first data (or normalized first data). The first extracted data may include first primary (or important or representative) image data for a first region selected by a user. According to one embodiment, the first primary image data may include partial image data for representing a partial image (primary image or important image or representative image) corresponding to (matching) the first context information from the image data.

[0135] In operation 460, a processor (220) according to one embodiment may store (or copy (e.g., AI smart copy)) first extracted data including first major (or important or representative) image data (e.g., boundary-edited image data) for a first area selected by a user through a first application.

[0136] FIG. 5 is a diagram showing operations for applying first extracted data obtained through a first user interface of a first application according to one embodiment to a second user interface.

[0137] Referring to FIG. 5, in operation 510, a processor (220) according to one embodiment may receive a second user input (514) (e.g., AI smart paste or paste) for applying the first extracted data to a second user interface (e.g., second app execution screen or messenger application execution screen) (512) of a second application (e.g., second app or messenger application).

[0138] In operation 520, a processor (220) according to one embodiment may acquire (or extract) second data corresponding to a second application based on a second user input. A processor (220) according to one embodiment may acquire second data corresponding to a second application to identify features of the second application in order to apply the first extracted data to a second user interface (e.g., AI smart paste or paste). A processor (220) according to one embodiment may extract second metadata of the second application (521). A processor (220) according to one embodiment may extract second image data and / or second text data (text data obtained by performing OCR on the second image data) corresponding to the second user interface (522). The second data according to one embodiment may include second metadata of the second application, second image data and / or second text data (text data obtained by performing OCR on the second image data) corresponding to the second user interface. According to one embodiment, the second data may further include information (e.g., application list and / or application category) regarding applications (or software) installed on the electronic device (201). According to one embodiment, the second data may further include a category of the second application (e.g., messenger category, map category, payment category, or other categories) or user data stored in association with the second application.

[0139] In operation 530, a processor (220) according to one embodiment may obtain second context information by performing context analysis on second data corresponding to a second application through a second artificial intelligence model (e.g., text encoder (531), image encoder (532), and / or multimodal large language model (533)). A processor (220) according to one embodiment may transmit second data including second metadata of the second application, second image data corresponding to the second user interface, and / or second text data (text data obtained by performing OCR on the second image data) as input values ​​to the second artificial intelligence model, and obtain second context information based on context analysis using the second artificial intelligence model.

[0140] A processor (220) according to one embodiment may further utilize (or further transmit to a second artificial intelligence model) application information installed on an electronic device (201) (e.g., app installation information within the terminal) (535), a first extracted image, and first extracted data (e.g., data copied from the first app (e.g., AI smart copied)) (534) when analyzing context. The second context information according to one embodiment may include information for data (e.g., second extracted data) that can be used (or extracted) from the second data or additionally (or additionally) acquired (e.g., data types available in the second application and a third app (information of the third application) that can acquire additional information among the applications included in the electronic device) to apply the first extracted data to the second user interface. A processor (220) according to one embodiment may obtain second extracted data from second data through a second artificial intelligence model (e.g., text encoder (531), image encoder (532), and / or multimodal large language model (533)) based on second context information. The second extracted data according to one embodiment may include data types (e.g., one or more UI components) that are available (or applicable or usable) for the second user interface and / or additional information for modifying the first extracted data to be applicable to the second data. A processor (220) according to one embodiment may identify a third application (e.g., a third app) among the applications of the electronic device (201) to be used for obtaining additional information, and may obtain additional information through background execution (536) of the third app.A processor (220) according to one embodiment can execute a third application in the background using information of a third application capable of obtaining additional information, and can extract (or obtain) additional information to change the first extracted data so that it can be applied to the second data through the third application.

[0141] In operation 540, a processor (220) according to one embodiment may generate content through a third artificial intelligence model (e.g., a text encoder (541), an image encoder (542), a large language model (LLM) (543) and / or a generative artificial intelligence model including a large visual model (544)). A processor (220) according to one embodiment may generate (or acquire) at least one content item comprising text data based on the first data, the second data, and / or the third app data, and image data based on the first data, the second data, and the third app data, through a text encoder (541), an image encoder (541), an LLM (543), and / or an LVM (544) corresponding to the third artificial intelligence model based on the first extracted data and the second extracted data. The at least one content item according to one embodiment is at least one content in a form usable in the second application. It may include an item (e.g., at least one content item that is displayed and accessible in the second user interface of the second application). In one embodiment, the at least one content item may include a plurality of content items. In one embodiment, the plurality of content items may have different data types.

[0142] FIG. 6 is a diagram illustrating operations for arranging multiple content items generated through a third artificial intelligence model according to one embodiment based on priority and displaying them on a display.

[0143] Referring to FIG. 6, in operation 610, a processor (220) according to one embodiment may generate at least one content item through a third artificial intelligence model (e.g., a generative artificial intelligence model). Operation 610 may refer to the description of operation 540 of FIG. 5.

[0144] In operation 620, a processor (220) according to one embodiment can identify the priority of multiple content items when at least one content item includes multiple content items. A processor (220) according to one embodiment can identify the priority of multiple content items in a data order meaningful from a user's perspective (e.g., in order of importance) based on second context information. A processor (220) according to one embodiment can identify the priority of multiple content items in order of high similarity between text summary information for each of the multiple content items and second context information of the second application. For example, the processor (220) can obtain multiple first vector values ​​for each of the multiple content items by embedding text summary information for each of the multiple content items into a vector through a text encoder (622), and obtain a second vector value by embedding second context information of the second application into a vector. The processor (220) can obtain cosine similarity values ​​between each of the first vector values ​​and the second vector value through a similarity analysis module (624) and compare the cosine similarity values ​​to identify the priority of multiple content items in order of highest similarity value. The processor (220) according to one embodiment can sort (or reorder) the order (e.g., display order) of multiple content items according to the priority of multiple content items. The processor (220) according to one embodiment can obtain the reordered multiple content items as the order (e.g., display order) of multiple content items is reordered according to the priority of multiple content items. The 620 operation may be omitted.

[0145] In operation 630, a processor (220) according to one embodiment may display a first main image (632) and at least one content item (634) on a second user interface through a display (260). In one embodiment, if at least one content item (634) includes a plurality of content items, the processor (220) may display the plurality of content items (634) arranged according to the priority of the plurality of content items (634). Each of the plurality of content items (634) may include text, an image, an icon or link information that can be linked to another application, and may include information of different categories (or topics) (description, location information, place information, and / or traffic information).

[0146] In operation 640, a processor (220) according to one embodiment can identify whether a user has selected at least one content item. If at least one content item is selected by user input, the processor (220) according to one embodiment can apply the selected at least one content item to a second application (e.g., forward or input into an input window). The processor (220) according to one embodiment can store the first data, first context information, first main image, first extracted data, second data, second context information, and / or second extracted data obtained from the time a first area is identified based on a first user input to a first user interface of a first application for data tracking until a first main image and at least one content item are displayed on a second user interface as tracking data.

[0147] In operation 650, the processor (220) according to one embodiment may configure a prompt for content regeneration if at least one content item is not selected by the user and proceed to operation 610 to regenerate (or re-acquire) at least one content item through a third artificial intelligence model (e.g., a generative artificial intelligence model) based on the first extracted data and the second extracted data.

[0148] In operation 660, a processor (220) according to one embodiment may perform data tracking when at least one content item is selected by a user. A processor (220) according to one embodiment may track and store data obtained during the data transfer process between the first application, the second application, and the third application after data transfer and display operation from the first application to the second application, and data transfer and display operation from the second application to the third application. A processor (220) according to one embodiment may use (or apply) the data tracked and stored during the data transfer process between the first application, the second application, and the third application in the next application data transfer process. According to one embodiment, the processor (220) may store the history of data converted from the data of the first application to the data of the Nth application, and may detect data changes between the data of the first application and the data of the Nth application. A processor (220) according to one embodiment may reflect the data tracked and stored during the data transfer process from the first application to the N-1st application when data is transferred from the N-1st application to the Nth application.

[0149] In operation 670, the processor (220) according to one embodiment can identify whether a change (or fluctuation) in some data is detected on the tracking path while data is sequentially transmitted from the first application to the Nth application and tracking data exists. The processor (220) according to one embodiment can continue to perform data tracking if there is no change in data.

[0150] In operation 680, the processor (220) according to one embodiment can reflect the content of the changed data in the user interface of the application associated with the changed data if there is a change (or variation) in some data on the tracking path while tracking data exists (it can update and display each user interface of each app).

[0151] FIG. 7a is a diagram illustrating an example of obtaining first extracted data including first image data from first data of a first area of ​​a map application execution screen based on user input to a map application execution screen of a map application according to one embodiment.

[0152] Referring to FIG. 7a, a processor (220) according to one embodiment is <751> As shown above, while the map application screen (701) (e.g., user interface) is displayed, user input (72) (e.g., AI smart select input) for the map application screen (701) can be received (①). A processor (220) according to one embodiment, based on the user input (72) for the map application screen (701), <752> As shown, a first area (710) corresponding to user input (72) of a map application screen (701) can be identified (or selected) (②). A processor (220) according to one embodiment can acquire first data including first image data (712) of the first area (710). The first data according to one embodiment may include at least one of UI component unit data corresponding to the first area (710), UI spatial data corresponding to the first area (710), first metadata corresponding to the first area (710), first image data (712) corresponding to the first area (710), or first text data corresponding to the first area (710).

[0153] A processor (220) according to one embodiment can extract and obtain information regarding one or more independent UI components included in a first area (710) as UI component unit data, such as a button (e.g., map zoom in button, map zoom out button, route search button, or other buttons related to a map application), a figure (e.g., a marker icon for indicating a place, destination, or current location, a route line, a POI (point of interest) icon, or other figures related to a map application), and a menu (e.g., a navigation options (walking, driving, public transport) menu, a place details menu, or other menus related to a map application), unique attributes of each of the one or more UI components (e.g., ID, name, location (x, y coordinates), size, type (button, figure, or menu), characteristics of each of the one or more UI components (e.g., whether they are active and / or visible), and / or subsequent actions, sub-items, and / or functions based on user interaction (e.g., actions) with the one or more UI components. there is.

[0154] A processor (220) according to one embodiment may extract and obtain UI spatial data corresponding to the first area (710). A processor (220) according to one embodiment may extract and obtain spatial relationship information between UI components (e.g., hierarchy and / or depth between UI components) as UI spatial data corresponding to the first area (710). For example, the hierarchy and / or depth between UI components may include the Android view structure or the HTML DOM structure. A processor (220) according to one embodiment may extract and obtain metadata of a map application corresponding to the first area (710). A processor (220) according to one embodiment may extract and obtain metadata of the first area (710) that is not displayed on the screen according to the characteristics of the map application corresponding to the first area (710) (e.g., important data or designated data).For example, if the first area (710) of the map application screen corresponds to a specific location area (e.g., near Gangnam Station), the processor (220) may include metadata including GPS information corresponding to the specific location area (e.g., Gangnam Station GPS information), information on facilities within the specific location area (e.g., road information (or road name) (e.g., road information around Gangnam Station), intersection information (or intersection name) (e.g., intersection information around Gangnam Station), subway information (or subway name) (e.g., Gangnam Station), bus stop information (or bus stop name) (e.g., Gangnam Station stop information), building information (or building name) (e.g., building information around Gangnam Station), shop information (or shop name) (e.g., shop information around Gangnam Station), restaurant information (or restaurant name) (e.g., restaurant information around Gangnam Station), famous restaurant information (or famous restaurant name) (e.g., famous restaurant information around Gangnam Station), or other facility information) (e.g., other facility information around Gangnam Station)), or traffic information associated with the specific location area (e.g., smooth or Metadata including congestion (e.g., traffic information around Gangnam Station) may be included, and additional metadata including other information obtainable in association with a specific location area (e.g., other information obtainable in association with Gangnam Station) may be included. A processor (220) according to one embodiment may obtain first image data (712) by copying (or cropping) image data of a first area (710) corresponding to user input (72) from image data of an application screen (701). A processor (220) according to one embodiment may identify an area within a specified range from the location where user input (72) is received as the first area (710), or identify an area within the range of the most important location (e.g., a landmark, a busy street) from the location where user input (72) is received as the first area (71), and obtain first image data (712) by copying (or cropping) image data of the first area (710).A processor (220) according to one embodiment can perform OCR on the first image data (712) to identify text data and obtain first text data containing text data. For example, the processor (220) can perform OCR on the first image data (712) to identify text data representing road information (or road name) (e.g., road information around Gangnam Station), intersection information (or intersection name) (e.g., intersection information around Gangnam Station), subway information (or subway name) (e.g., Gangnam Station), bus stop information (or bus stop name) (e.g., Gangnam Station stop information), building information (or building name) (e.g., building information around Gangnam Station), shop information (or shop name) (e.g., shop information around Gangnam Station), restaurant information (or restaurant name) (e.g., restaurant information around Gangnam Station), famous restaurant information (or famous restaurant name) (e.g., famous restaurant information around Gangnam Station), or other facility information (e.g., other facility information around Gangnam Station) shown by the first image data (712).

[0155] A processor (220) according to one embodiment may transmit at least one of UI component unit data corresponding to the first area (710), UI spatial data corresponding to the first area (710), first metadata corresponding to the first area (710), first image data (712) corresponding to the first area (710), or first text data corresponding to the first area (710) to a first artificial intelligence model (multimodal AI model) to perform context analysis and obtain first context information. The first context information according to one embodiment may include association relationships between UI components, metadata and action mapping tables (e.g., a mapping table for metadata and actions (e.g., actions, sub-items, and / or functions) connected to UI components), and boundaries of important areas within the first image data (712) (e.g., areas including subway stations and intersections).

[0156] A processor (220) according to one embodiment can recognize the UI component entity name of the first data and normalize the first data. A processor (220) according to one embodiment can preprocess (e.g., generalize, normalize, or standardize) the first data into a form that can be used in at least one other application or other format (e.g., iOS, Android OS, or Web) through a text encoder and a multi-modal large language model (LLM). A processor (220) according to one embodiment can normalize the text contained in the first data to obtain normalized text data, and use the normalized text data to obtain normalized first data corresponding to the first data. For example, text normalization, which normalizes the text contained in the first data, may mean converting text representing UI elements into standard or general text so that UI elements between different platforms (Android, iOS, Web, etc.) can be compared or converted. For example, text normalization may include standardizing uppercase and lowercase letters (e.g., CheckBox → checkbox), removing unnecessary symbols, standardizing abbreviations or syllabaries (e.g., btn → button), translation or multilingual processing, and mapping platform-specific names to common expressions.

[0157] A processor (220) according to one embodiment may perform first data optimization through redefining a major region of interest (③). A processor (220) according to one embodiment may use first context information obtained through a first artificial intelligence model to redefine a region of interest among a first region (712) to extract first major image data (or first important image or first representative image data) (714) (e.g., image data including a subway station and an intersection) that corresponds to (matches) the first context information from the first image data, and obtain partial data (major data or important data or representative data) (e.g., first extracted data) that corresponds to (matches) the first context information from the first data (or normalized first data). A processor (220) according to one embodiment can store (or copy (e.g., AI smart copy)) a first extracted data (e.g., normalized first data) including first key image data (e.g., boundary-edited image data) (714) as a result of optimizing first data from a map application.

[0158] FIG. 7b is a diagram showing an example of applying first extracted data obtained through a map application screen according to one embodiment to a messenger application screen.

[0159] Referring to FIG. 7b, a processor (220) according to one embodiment is <753> As shown above, messages (e.g., "Shall we look for a good restaurant near Gangnam Station?" or "Please take me near the Gangnam Station Line 2 bus stop") can be exchanged with another person (e.g., Yu OO) through the messenger application screen (702) (e.g., messenger application execution screen or messenger application conversation screen) (④). A processor (220) according to one embodiment may receive a second user input (76) (e.g., AI smart paste) to apply first extracted data including first main image data (714) to the messenger application screen (e.g., messenger application execution screen or messenger application conversation screen) (702) while exchanging messages with another person through the messenger application screen (702) (⑤). A processor (220) according to one embodiment may acquire (or extract) second data corresponding to the messenger application based on the second user input (76). A processor (220) according to one embodiment may acquire second data corresponding to a messenger application to identify the characteristics of the messenger application in order to apply (e.g., paste) the first extracted data, which includes the first main image data (714), to the messenger application execution screen. The second data according to one embodiment may include second metadata of the messenger application, second image data corresponding to the messenger application execution screen, and / or second text (text obtained by performing OCR on the second image). The second metadata according to one embodiment may include metadata related to a user exchanging messages, metadata related to a chat room, metadata related to UI components related to the messenger application screen, or other metadata related to the messenger application.According to one embodiment, the second image data corresponding to the messenger application execution screen may include the messenger application execution screen image data at the time of the second user input (76). According to one embodiment, the second text (text obtained by performing OCR on the second image data) may include text data obtained by performing OCR on the messenger application execution screen image data at the time of the second user input (76) (e.g., "Shall we look for a good restaurant near Gangnam Station?" or "Please take me near the Gangnam Station Line 2 bus stop"). According to one embodiment, the second data may further include information about applications (or software) installed on the electronic device (201) (e.g., application list and / or application category). According to one embodiment, the second data may further include the category of the messenger application (e.g., messenger category) or user data stored in association with the messenger application (e.g., conversation history) (e.g., conversation history prior to "Shall we look for a good restaurant near Gangnam Station?" or "Please take me near the Gangnam Station Line 2 bus stop"). A processor (220) according to one embodiment can obtain second context information by performing context analysis on second data corresponding to a messenger application through a second artificial intelligence model (e.g., a multimodal artificial intelligence model). When performing context analysis, the processor (220) according to one embodiment may further utilize application information installed on the electronic device (201) (e.g., app installation information within the terminal), a first extracted image, and first extracted data (e.g., data copied from a map application (e.g., AI smart copied)).According to one embodiment, the second context information may include information for data (e.g., second extracted data) that can be used (or extracted) or additionally (or additionally) obtained from the second data to apply the first extracted data to a messenger application (e.g., data types usable in a messenger application and a third app (information of the third application (e.g., a search application)) that can obtain additional information among the applications included in the electronic device). According to one embodiment, the processor (220) may obtain the second extracted data from the second data through a second artificial intelligence model (e.g., a multimodal artificial intelligence model) based on the second context information. According to one embodiment, the second extracted data may include a data type that is usable (or applicable or usable) on the messenger application screen (e.g., one or more UI components usable on the messenger application screen or a data type that can be displayed on the messenger application screen (e.g., text, image, video, or URL)) and additional information for modifying the first extracted data so that it can be applied to the second data. According to one embodiment, the processor (220) identifies a third application (third app) (e.g., a search application) among the applications of the electronic device (201) to be used for obtaining additional information, and can obtain additional information through background execution of the third app. According to one embodiment, the processor (220) can execute the search application in the background using information of the search application capable of obtaining additional information, and can extract (or obtain) additional information for modifying the first extracted data so that it can be applied to the second data through the search application. According to one embodiment, the processor (220) can create (or obtain) at least one content item (720) based on the first extracted data and the second extracted data.According to one embodiment, at least one content item (720) may be a data type usable in the second application (e.g., one or more UI components usable on the messenger application screen) or a data type displayable on the messenger application screen (e.g., text, image, video, or URL). According to one embodiment, at least one content item (720) may include a plurality of content items (e.g., a content item associated with restaurant information (722), a content item associated with traffic information (724), a content item associated with location information (726)). According to one embodiment, the plurality of content items (722, 724, 726) may differ from each other in data type (e.g., text, image, video, or URL) or category (e.g., restaurant information, traffic information, or location information).

[0160] A processor (220) according to one embodiment can identify the priority of a plurality of content items (722, 724, 726). A processor (220) according to one embodiment can identify the priority of a plurality of content items (722, 724, 726) based on second context information. A processor (220) according to one embodiment can sort (or reorder) the order (e.g., display order) of a plurality of content items (722, 724, 726) according to the priority of the plurality of content items. A processor (220) according to one embodiment through a display (260) <753> As shown above, a first main image (714) and a plurality of content items (722, 724, 726) arranged according to priority (e.g., restaurant information -> traffic information -> location information) can be displayed on the messenger application execution screen (702). Each of the plurality of content items (722, 724, 726) may include text, an image, an icon or link information that can be linked to another application, and may include information of different categories (or topics) (restaurant information, traffic information, or location information).

[0161] A processor (220) according to one embodiment can apply at least one content item (722) and a first main image (714) to a messenger application (e.g., forwarding or inputting into an input window of a messenger application screen (702)).

[0162] A processor (220) according to one embodiment may display at least one content item (722) and a first main image (714) on a messenger application screen (702), and then, when at least one content item (722) is selected by user input, apply the selected at least one content item (722) and the first main image (714) to the messenger application (e.g., forward or input into an input window of the messenger application screen (702). A processor (220) according to one embodiment <754> As such, a result (730) with at least one content item (722) and a first main image (714) applied can be displayed on the messenger application screen (702) (⑥, ⑦).

[0163] FIG. 8 is a diagram illustrating an example of obtaining image data that has been edited from the first image data of a first area according to one embodiment.

[0164] Referring to FIG. 8, a processor (220) according to one embodiment can obtain an image data (850) that has been edited from the first image data (812) when first data including the first image data (812) of the first area (810) is obtained based on a first user input for a movie application screen (801) of the movie application. A processor (220) according to one embodiment can identify at least one cell corresponding to a first image (812) among the cells according to the grid structure of the movie application screen (801) using cell information (e.g., area information) according to the grid structure of the movie application screen (801), and can identify at least one first cell among the at least one cell corresponding to the first image data (812) in which important data (e.g., data corresponding to the first context corresponding to the first data (e.g., UI component, action connected to the UI component, sub-item, and / or function, or first metadata)) is mapped or more than a specified number of data items are mapped, and at least one second cell in which data is not mapped or unimportant data is mapped or less than a specified number of data items are mapped. A processor (220) according to one embodiment can obtain edited image data (850) by editing the image data (852) corresponding to at least one first cell of the first image data (810) so that the image data (854) corresponding to at least one second cell of the first image data (810) becomes a high-quality image. A processor (220) according to one embodiment can divide a screen of a user interface (e.g., a movie application screen (801)) into a two-dimensional grid structure composed of fixed-size cells and perform independent data processing for each cell.A processor (220) according to one embodiment can display important image data that reflects a lot of context in each cell as a high-quality image and display image data that reflects relatively little context as a low-quality image so as to preserve visual information while increasing processing efficiency. A processor (220) according to one embodiment can increase the efficiency of data analysis and processing by selectively displaying high-quality images and low-quality images according to the importance of image data in each cell, rather than displaying the entire user interface screen as a high-quality image.

[0165] FIG. 9 is a diagram showing an example of applying first extracted data, including first main image data obtained through a movie application screen according to one embodiment, to a messenger application screen.

[0166] Referring to FIG. 9, a processor (220) according to one embodiment may receive a first user input (92) (e.g., AI smart select input) while a movie application screen (901) is displayed. A processor (220) according to one embodiment may identify (or select) a first area (910) of the movie application screen (901) based on the first user input (92). A processor (220) according to one embodiment may obtain first image data (912) of the first area (910) and first data including the first image data (912) of the first area (910). The first data including the first image data (912) of the first area (910) according to one embodiment may include at least one of UI component unit data corresponding to the first area (910), UI spatial data corresponding to the first area (910), first metadata corresponding to the first area (910), first image data (912) corresponding to the first area (910), or first text data corresponding to the first area (910).A processor (220) according to one embodiment may extract and obtain information regarding one or more independent UI components included in a first area (910), such as a button (e.g., a seat information button, or other buttons related to a movie application), a figure (e.g., a figure of showtimes and seat information, or other figures related to a movie application), and a menu (e.g., a reservation menu, or other menus related to a movie application), as UI component unit data, unique attributes of each of the one or more UI components (e.g., ID, name, location (x, y coordinates), size, type (button, figure, or menu), characteristics of each of the one or more UI components (e.g., whether they are active and / or visible), and / or subsequent actions, sub-items, and / or functions according to user interaction (e.g., actions) with the one or more UI components. The UI component unit data according to one embodiment includes a button corresponding to 24:00 125 / 130 seats, and JavaScript connected to the button and related to seat assignment. It can include an action that navigates to a function or HTML page.

[0167] A processor (220) according to one embodiment may extract and obtain UI spatial data corresponding to the first area (910). A processor (220) according to one embodiment may extract and obtain spatial relationship information between UI components (e.g., hierarchy and / or depth between UI components) as UI spatial data corresponding to the first area (910). For example, the hierarchy and / or depth between UI components may include the Android view structure or the HTML DOM structure. A processor (220) according to one embodiment may extract and obtain metadata of a movie application corresponding to the first area (910). A processor (220) according to one embodiment may extract and obtain metadata of the first area (910) that is not displayed on the screen (e.g., important data or designated data) depending on the characteristics of the movie application corresponding to the first area (910). A processor (220) according to one embodiment may extract and obtain metadata including information on the screening time and remaining seats of a specific movie when the first area (910) of the movie application screen corresponds to the screening time and remaining seats of a specific movie. A processor (220) according to one embodiment may obtain first image data (912) by copying (or cropping) the image data of the first area (910) corresponding to the user input (92) from the image data of the movie application screen (901). A processor (220) according to one embodiment may identify the area of ​​a specified range from the location where the user input (92) is received as the first area (910), or identify the area of ​​a range containing important data from the location where the user input (92) is received as the first area (910), and obtain first image data (912) by copying (or cropping) the image data of the first area (910).A processor (220) according to one embodiment may perform OCR on the first image data (912) to identify text data and obtain first text data including text data (e.g., 24:00 125 / 130 seats). For example, the processor (220) may perform OCR on the first image data (912) to identify text data indicating the screening time and remaining seats shown by the first image data (912).

[0168] A processor (220) according to one embodiment can extract partial image data (important image data) (e.g., first main image data) (914) (e.g., image data representing 24:00 125 / 130 seats) corresponding to (matching to) the first context information from the first image data (912) through a first artificial intelligence model based on first context information for the first data including the first image data (912), and obtain and store (or copy (e.g., AI smart copy)) partial data (important data) (e.g., first extracted data) (e.g., empty seat information, movie theater information) corresponding to (matching to) the first context information from the first data (or normalized first data).

[0169] A processor (220) according to one embodiment may receive a second user input (94) (e.g., AI smart paste) for applying the first extracted data, which includes the first main image data (914), to a messenger application screen (e.g., a messenger application execution screen or a conversation screen of a messenger application) (902). A processor (220) according to one embodiment may exchange messages (e.g., "Shall we go to the movie theater tonight if there are seats available?" and "I'll look for it on the movie application") with a counterparty (e.g., Yu OO) through the messenger application screen (902) (e.g., a messenger application execution screen or a conversation screen of a messenger application). A processor (220) according to one embodiment may receive a second user input (94) (e.g., AI smart paste) for applying the first extracted data, which includes the first main image data (914), to a messenger application screen (e.g., a messenger application execution screen or a conversation screen of a messenger application) (902) while exchanging messages with a counterparty through the messenger application screen (902).

[0170] A processor (220) according to one embodiment can obtain (or extract) second data corresponding to a messenger application based on the second user input (94).

[0171] A processor (220) according to one embodiment may acquire second data corresponding to a messenger application to identify the features of a messenger application in order to apply (e.g., paste) first extracted data including first main image data to a messenger application execution screen.

[0172] According to one embodiment, the second data may include second metadata of the messenger application, second image data corresponding to the execution screen of the messenger application, and / or second text (text obtained by performing OCR on the second image). According to one embodiment, the second metadata may include metadata related to users exchanging messages, metadata related to chat rooms, metadata related to UI components related to the messenger application screen, or other metadata related to the messenger application. According to one embodiment, the second image data corresponding to the execution screen of the messenger application may include image data of the execution screen of the messenger application at the time of the second user input (94). According to one embodiment, the second text (text obtained by performing OCR on the second image data) may include text data obtained by performing OCR on the image data of the execution screen of the messenger application at the time of the second user input (95) (e.g., "Shall we go to the movie theater at night if there are seats available?" and "I'll look for it on the movie application"). The second data according to one embodiment may further include information about applications (or software) installed on the electronic device (201) (e.g., application list and / or application category). The second data according to one embodiment may further include a category of a messenger application (e.g., messenger category) or user data stored in association with a messenger application (e.g., conversation history) (e.g., conversation history prior to "Shall we go to the movie theater at night if there are seats available?" and "I'll look for it on the movie application").

[0173] A processor (220) according to one embodiment can identify second context information corresponding to a messenger application by performing context analysis on second data corresponding to a messenger application through a second artificial intelligence model (e.g., a multimodal artificial intelligence model). The second context information according to one embodiment may include information for data (e.g., second extracted data) that can be used (or extracted) from the second data or additionally (or additionally) acquired in order to apply the first extracted data to the messenger application (e.g., data types usable in the messenger application and a third app (information of the third application (e.g., a search application)) that can acquire additional information among the applications included in the electronic device). A processor (220) according to one embodiment can acquire second extracted data from the second data through a second artificial intelligence model (e.g., a multimodal artificial intelligence model) based on the second context information. According to one embodiment, the second extracted data may include a data type that is usable (or applicable or available) on the messenger application screen (e.g., one or more UI components usable on the messenger application screen or a data type that can be displayed on the messenger application screen (e.g., text, image, video, or URL)) and additional information for modifying the first extracted data to be applicable to the second data. According to one embodiment, the processor (220) may generate (or acquire) at least one content item (930) based on the first extracted data and the second extracted data. According to one embodiment, at least one content item (920) may be a data type usable in the second application (e.g., one or more UI components usable on the messenger application screen or a data type that can be displayed on the messenger application screen (e.g., text, image, video, or URL).According to one embodiment, at least one content item (930) may include a plurality of content items (e.g., a content item associated with empty seat information (932), a content item associated with movie theater information (934)). According to one embodiment, the plurality of content items (932, 934) may have different data types (e.g., text, image, video, or URL) or categories (e.g., empty seat information, movie theater information).

[0174] A processor (220) according to one embodiment may apply a first main image (914) and at least one content item (922, 924) to a messenger application (e.g., forwarding or inputting into an input window of a messenger application screen (902)) and display them on a messenger application screen (902).

[0175] FIG. 10 is a diagram showing an example of data transfer and display between a calendar application, a map application, and a messenger application according to one embodiment.

[0176] Referring to FIG. 10, in one embodiment, a processor (220) can track and store data obtained during the data transfer and display process between the calendar application, the map application, and the messenger application when the operation of transferring and displaying data (e.g., first extracted data including first main image data (1010)) from the calendar application screen (1001) to the map application screen (1002) and then transferring and displaying data (e.g., second extracted data including second main image data (1024)) from the map application screen (1002) to the messenger application screen (1003) is performed.

[0177] A processor (220) according to one embodiment can display a calendar application screen (1001), a map application screen (1002), and a messenger application screen (1003) through a display (260).

[0178] A processor (220) according to one embodiment may receive a first user input (1011) (e.g., AI smart select input) on a calendar application screen (1001) (①). A processor (220) according to one embodiment may identify (or select) at least one UI component (1010) of the calendar application screen (1001) based on the first user input (1011).

[0179] A processor (220) according to one embodiment may collect widget data corresponding to the location where the first user input (1010) is received and obtain at least one UI component (1010), which is an interactive screen component included in the widget data. The at least one UI component (1010) may include a menu component, a button component, a checkbox component, a slider component, and / or an image component, and may include actions, sub-items, and / or functions associated (or connected) with each component. For example, the at least one UI component may include a button, a checkbox, a slider, a text field, a picture, a dropdown menu, etc., and in the case of a web page, <figure> , <menu>It can include HTML Tag-unit components representing inserted content or menus.

[0180] A processor (220) according to one embodiment can identify at least one UI component (1010) including "10 Gangnam Station Exit 0 Appointment" and "16 Seocho-daero Nearby Dinner" based on the location where user input (1011) is received.

[0181] A processor (220) according to one embodiment may acquire first data including first image data of at least one UI component (1010) (or an image of one or more UI components). The first data according to one embodiment may include at least one of UI component unit data corresponding to at least one UI component (1010), UI spatial data corresponding to at least one UI component (1010), first metadata corresponding to at least one UI component (1010), first image data corresponding to at least one UI component (1010), or first text data corresponding to at least one UI component (1010). A processor (220) according to one embodiment can extract and obtain information regarding buttons (e.g., a button including “10 Gangnam Station Exit 0 Appointment” and a button including “16 Seocho-daero Nearby Dinner”) as UI component unit data, unique attributes of the buttons (e.g., ID, name, location (x, y coordinates), size, type (button)), characteristics of each button (e.g., whether it is active and / or visible), and / or subsequent actions (action), sub-items, and / or functions based on user interaction with the buttons (e.g., actions).

[0182] A processor (220) according to one embodiment can extract and obtain UI spatial data corresponding to at least one UI component (1010). A processor (220) according to one embodiment can extract and obtain spatial relationship information between UI components (e.g., hierarchy and / or depth between UI components) as UI spatial data corresponding to at least one UI component (1010). For example, the hierarchy and / or depth between UI components may include the Android view structure or the HTML DOM structure. A processor (220) according to one embodiment can extract and obtain metadata of a calendar application corresponding to at least one UI component (1010). A processor (220) according to one embodiment can extract and obtain metadata that is not displayed on the screen (e.g., important data or designated data) depending on the characteristics of the calendar application. A processor (220) according to one embodiment may obtain a first image data by copying (or cropping) the image data of at least one UI component (1010) from the image data of a calendar application screen (1001). A processor (220) according to one embodiment may perform OCR on the first image data to identify text data and obtain a first text data including text data (e.g., “10 Gangnam Station Exit 9 appointment” and “16 company dinner near Seocho-daero”). For example, the processor (220) may perform OCR on the first image data to identify text data shown by the first image data (1012).

[0183] A processor (220) according to one embodiment can extract partial image data (important image data) (e.g., first key image data) (1012) (e.g., image data representing “10 Gangnam Station Exit 0 appointment” and “16 Seocho-daero nearby dining”) corresponding to (matching to) the first context information from the first image data (1012) through a first artificial intelligence model based on first context information for first data including first image data, and acquire and store (or copy (e.g., AI smart copy)) partial data (important data) (e.g., first extracted data) (e.g., “10 Gangnam Station Exit 0 appointment” and “16 Seocho-daero nearby dining”) corresponding to (matching to) the first context information from the first data (or normalized first data).

[0184] A processor (220) according to one embodiment may receive a second user input (e.g., an input for dragging (or dragging and dropping) the first image data toward the map application screen (1002)) for applying the first extracted data including the first main image data (1012) to the map application screen (e.g., a map application execution screen) (1002) (②).

[0185] A processor (220) according to one embodiment can acquire (or extract) second data including second image data (1022) of a second area in a map application screen (1002) based on a second user input (3). The feature of acquiring (or extracting) second data including second image data (1022) of a second region according to one embodiment may be performed in the same or similar manner as the method of acquiring (or extracting) first data including first image data (712) described in FIG. 7a. A processor (220) according to one embodiment may acquire second context information based on the second data through a first artificial intelligence model and perform optimization of the second data using the second context information to acquire second extracted data including a second main image (1024). The feature of acquiring second extracted data including a second main image (1024) according to one embodiment may be performed in the same or similar manner as the method of acquiring (or extracting) first extracted data including the first main image data (714) by performing first data optimization described in FIG. 7a.

[0186] A processor (220) according to one embodiment can use (or apply) tracked stored data (e.g., first data, first context information, first extracted data, second data, second context information, second extracted data) during the data transfer process between a calendar application and a map application during the data transfer process between a map application and a messenger application. According to one embodiment, when a processor (220) performs a data transmission and / or display operation from a calendar application screen (1001) to a map application screen (1002) using data (e.g., first data and first extracted data) (e.g., "December 10th Gangnam Station Exit 9 appointment" and "dinner party near 16 Seocho-daero"), and then performs a data transmission or display operation from the map application screen (1002) to a messenger application screen (1003) using data (e.g., second data and second extracted data) (e.g., map data around Gangnam Station Exit 9), at least one content item (e.g., "December 10th Gangnam Station Exit 9 appointment" and map data near Gangnam Station Exit 9 or Seocho-daero") using at least a portion of the data obtained during the data transmission and / or display process between the calendar application, the map application, and the messenger application (e.g., "December 10th Gangnam Station Exit 9 appointment" and map data near Gangnam Station Exit 9 or Seocho-daero") is displayed on the messenger application screen (1003). Depending on the choice to apply at least one content item that can be created and created to the message application screen (1003), a message (1026) (e.g., "December 10th, coming down from Gangnam-daero at Gangnam Station Exit 9, turn right at the first alley") that reflects data obtained during the data transfer process between the calendar application, map application, and messenger application can be displayed on the message application screen (1003) (⑤).

[0187] FIG. 11 is a diagram illustrating an example in which a change in some of the tracking data occurs during data transmission and display between a calendar application, a map application, and a messenger application according to one embodiment.

[0188] Referring to FIG. 11, a processor (220) according to one embodiment can identify that the schedule information has been changed from the 16th to the 17th and the location information has been changed from Gangnam-daero to Sinnonhyeon Station as a message (1126) that changes the location and schedule (e.g., "But let's change the schedule for the 16th to the 17th and change the location to Sinnonhyeon Station") is obtained through the message application screen (1003) after displaying a message (1026) that changes the location and schedule (e.g., "But let's change the schedule for the 16th to the 17th and change the location to Sinnonhyeon Station") on the map application screen (1002) according to one embodiment (⑦). A processor (220) according to one embodiment may change the schedule for the 16th (e.g., 16th -> 17th) to the schedule for the 17th through a calendar application according to a change in the schedule (e.g., 16th -> 17th) and display at least one UI component (1110) including the changed schedule for the 17th on a calendar application screen (1001). A processor (220) according to one embodiment may further perform an update to change the route or location of a map application according to the schedule changed in the calendar application.

[0189] FIG. 12 is a diagram illustrating an example of providing information of an application capable of applying first data, including first image data of a first area obtained through a user interface of an application according to one embodiment.

[0190] Referring to FIG. 12, a processor (220) according to one embodiment may display a calendar application screen (1201) through a display (260) in accordance with the execution of a calendar application. A processor (220) according to one embodiment may obtain first data including first image data of a first area (1210) based on a first user input (1205) on the calendar application screen (1201). A processor (220) according to one embodiment may execute an AI assistant (or an AI assistant function or program) and display an AI assistant execution screen (1250). An AI assistant according to one embodiment may include a conversational AI assistant. A processor (220) according to one embodiment may input text (e.g., a query or phrase) (e.g., "What is the purpose of this text field and how can it be converted?") to obtain information of at least one application capable of applying (or utilizing) first data including first image data corresponding to a first area (1210) through an AI assistant execution screen (1250). A processor (220) according to one embodiment may provide (display or output) information of at least one application capable of applying (or utilizing) the first data including the first image data corresponding to the first area (1210) among the applications included in the electronic device (201) in response to an input (e.g., “This field is for inputting a schedule and can be used in a note app, map app, browser app, etc.”). A processor (220) according to one embodiment may provide information through an AI assistant regarding whether the first data including the first image data corresponding to the first area (1210) of the first area (1210) includes data obtained through data transfer and display operations between previous applications.A processor (220) according to one embodiment may provide information indicating which data of which previous application was converted into which data and included in the first data when the first data, which includes first image data corresponding to the first area (1210) through an AI assistant, includes data obtained through data transfer and / or display operations between previous applications.

[0191] FIG. 13 is a diagram illustrating an example of providing information indicating the previous usage history of a first data including first image data corresponding to a first area obtained through a user interface of an application according to one embodiment.

[0192] Referring to FIG. 13, a processor (220) according to one embodiment may display usage history (or history) information (e.g., AI smart select briefing) (1360) for first data (e.g., "December 10 Gangnam Station Exit 9 appointment" and "dinner near 17 Seocho-daero") when acquiring first data including first image data corresponding to a first area (1310) based on a first user input (1305) on a calendar application screen (1301). The data usage history information (1360) according to one embodiment may include the application name, date, and purpose (or action) in which the first data including the first image data of the first area (1310) was used. A processor (220) according to one embodiment may display the usage history (or history) of the first data including the first image of the first area (1310) as a timeline. A processor (220) according to one embodiment may execute an application that uses first data, including first image data of a first area (1310) of the selected time point, when a specific time point (date or time) is selected by a user from the timeline of data usage history (or history) information (1360). A processor (220) according to one embodiment may display a screen (1305) of a map application according to the first usage history (1362) when a first usage history (1362) (e.g., map, location, traffic information addition) is selected from the data usage history (or history) information (1360).A processor (220) according to one embodiment can display on the display (260) again a map application screen (1305) that includes a location area (1352) (e.g., Gangnam Station Exit 9 or an area near Seocho-daero) associated with the first data that was displayed on the display (260) at the time of the first usage history (1362), the first data (e.g., "December 10 Gangnam Station Exit 9 appointment" and "17 Seocho-daero nearby dinner") that was applied to the map application.

[0193] A processor (220) according to one embodiment may display a messenger application screen (1307) according to the second usage history (1364) (e.g., Messenger - schedule change from 16th to 17th) when a second usage history (1364) (e.g., Messenger - schedule change from 16th to 17th) is selected among the data usage history (or history) information (1360). A processor (220) according to one embodiment may display a messenger application screen (1307) again on the display (260) containing a message (1372) (e.g., "December 10th Gangnam Station Exit 9 appointment" and "17 Seocho-daero nearby dinner") associated with the first data that was displayed on the display (260) at the time of the second usage history (1364), when the first data (e.g., "December 10th Gangnam Station Exit 9, coming down from Gangnam-daero, turn right at the first alley") is applied to the messenger application.

[0194] A processor (220) according to one embodiment may acquire a first image of a first area representing text data of a part of song lyrics and first data corresponding to the first area based on user input during an internet search through a web browser application screen. A processor (220) according to one embodiment may execute a music streaming application to display a music streaming application screen, search for a singer and title corresponding to the song lyrics based on user input for applying the first image of a first area representing text data of a part of song lyrics and first data corresponding to the first area to the music streaming application screen, and enable music to be played through the music streaming application screen using the search results.

[0195] A processor (220) according to one embodiment may acquire a first image of a first area corresponding to an image of an internet article and first data corresponding to a first area based on user input (e.g., copy) while displaying an internet article through a web browser application screen. A processor (220) according to one embodiment may display content with a title written by summarizing the article content in a poster format over an image of an internet article through an image editing application screen based on user input (e.g., paste) for applying the first image of a first area corresponding to an image of an internet article and the first data corresponding to a first area to an image of an internet article to an image editing application screen. A processor (220) according to one embodiment may convert an image of an internet article to fit the posting format of an SNS application screen based on user input (e.g., paste) for applying the first image of a first area corresponding to an image of an internet article and the first data corresponding to a first area to an SNS application screen and display the converted content (e.g., an image of an internet article and a post summarizing the main contents of the internet article) on an SNS application screen. A processor (220) according to one embodiment analyzes the page structure of a document viewer application (e.g., HTML UI structure, or DOM Tree structure) based on user input (e.g., paste) for applying a first image of a first area corresponding to an image of an internet article and a first data corresponding to a first area to a document viewer application (e.g., PDF application or Excel application), and converts the first image of a first area corresponding to an image of an internet article and the first data corresponding to a first area to match the page structure of the document viewer application screen, and can display the converted content (e.g., content in which text and images are converted to an appropriate resolution while having the same document structure (e.g., table of contents, header / footer, margins)) on the document viewer application screen.

[0196] A processor (220) according to one embodiment may acquire a first image of a first area corresponding to specific conversation content and first data corresponding to a first area based on user input (e.g., copy) for specific conversation content while displaying conversation history through a messenger application screen. A processor (220) according to one embodiment may convert the first image of a first area corresponding to specific conversation content and the first data corresponding to a first area corresponding to specific conversation content into content (e.g., place, restaurant, or travel route) applicable to a map application screen based on user input (e.g., paste) for applying the first image of a first area corresponding to specific conversation content and the first data corresponding to a first area to a map application screen according to the context corresponding to the messenger application, and display the converted content on the map application screen. A processor (220) according to one embodiment can convert the first image of the first area corresponding to the specific conversation content and the first data corresponding to the first area corresponding to the specific conversation content into content (e.g., event title, schedule, location, place) applicable to the calendar application screen based on user input (e.g., paste) for applying the first image of the first area corresponding to the specific conversation content and the first data corresponding to the first area corresponding to the first area to the calendar application screen according to the context corresponding to the calendar application, and display the converted content on the calendar application screen.

[0197] A processor (220) according to one embodiment may obtain a first image of a first area corresponding to a specific menu button and first data corresponding to a first area based on user input (e.g., copy) for a specific menu button in a settings menu list. A processor (220) according to one embodiment may convert the first image of the first area corresponding to a specific menu button and the first data corresponding to a first area into content (e.g., guide information indicating a method to access a specific menu and an explanation of a specific menu) that can be applied to a messenger application screen based on user input (e.g., paste) for applying the first image of the first area corresponding to a specific menu button and the first data corresponding to a first area to a messenger application screen, and display the converted content on the messenger application screen.

[0198] A processor (220) according to one embodiment can obtain a first image and first data (e.g., the entire state of the conversation content and the conversation content) corresponding to the conversation content based on a user's voice command (e.g., smart copy) regarding the conversation content during a conversation using voice with an AI assistant through an AI assistant application screen. A processor (220) according to one embodiment can convert the first image and first data corresponding to the conversation content into content (e.g., text data organized around a schedule or the topic of the content the user inquired about) applicable to the note application screen based on a user command (e.g., paste) for applying the first image and first data corresponding to the conversation content to the note application screen, and display the converted content on the note application screen.

[0199] A processor (220) according to one embodiment can obtain a first image and first data corresponding to an area of ​​a specific restaurant alley based on user input (e.g., copy) regarding an area of ​​a specific restaurant alley through a map application screen. A processor (220) according to one embodiment can convert the first image and first data corresponding to an area of ​​a specific restaurant alley into content (e.g., a list of restaurants and information, review data included in a specific restaurant alley (street)) applicable to a web browser application screen based on a user command (e.g., paste) for applying the first image and first data corresponding to an area of ​​a specific restaurant alley to a web browser application screen, and can display the converted content on the web browser application screen. When at least some of the content (e.g., a specific restaurant) among the content displayed on the web browser application screen is selected, the processor (220) according to one embodiment can update and display related data or display a route to the specific restaurant on the map application screen according to the information of the selected content (e.g., a specific restaurant).

[0200] FIG. 14 is a generative artificial intelligence (AI) system (1400) according to one embodiment.

[0201] Referring to FIG. 14, a generative AI system (1400) may include a User Interface (1410), an AI Framework (1420), a Generative AI Model (1430), a Knowledge Repository (1440), and an Application / Service Module (1450). These components may be operated on one or more of an electronic device (101), an external electronic device (102 or 104), or a server (108). For example, the User Interface (1410) and the AI ​​Framework (1420) may be operated on the electronic device (101), and the Knowledge Repository (1440) and the Generative AI Model (1430) may be operated on the server (108).

[0202] According to one embodiment, the User Interface (1410) may receive user input (e.g., user query). User input may be received in the form of text, images, voice (e.g., natural language), video, menu selection, or a combination thereof. The User Interface (1410) may include various context information (e.g., running application or user location) related to the generative artificial intelligence system (1400) at the time the user input is received, in addition to or instead of the user input. The User Interface (1410) may provide the user input or the context information to the AI ​​Framework (1420) and provide the result of processing therefrom to the user, for example, through the AI ​​Framework (1420). According to one embodiment, in addition to user input, the electronic device may provide context information obtained using information included on the screen to the AI ​​Framework (1420). The result may be provided in the form of text, images, voice, video, an action requested by the user (e.g., launching a specified function or app), or a combination thereof.

[0203] According to one embodiment, the AI ​​Framework (1420) can identify (e.g., estimate) a user intent based on at least part of user input or context information received from the User Interface (1410), control each of the relevant modules (e.g., 1421, 1423, or 1425) to perform a function or action corresponding to the identified user intent, and coordinate collaboration between two or more modules. The AI ​​Framework (1420) may include a Prompt Design Module (1421), an API / Plug-in Management Module (1423), and an Output Modification Module (1425), as illustrated in FIG. 14.

[0204] According to one embodiment, the Prompt Design Module (1421) can generate a prompt to be input to the Generative AI Model (1430) based at least partially on user input or context information received from the User Interface (1410). For example, the Prompt Design Module (1421) can generate a prompt using user preferences, a prompt library, or prompt examples stored in the Knowledge Repository (1440) based at least partially on user input or context information.

[0205] According to one embodiment, the API / Plug-in Management Module (1423) may communicate, for example, via an API, with various resources (e.g., Knowledge Repository (1440)) that provide said additional information when there is a request for said additional information in relation to user input. Additionally or generally, when a specified action (e.g., function, app, or service) is performed in response to said user input, the API / Plug-in Management Module (1423) may request the Application / Service Module (1450) to perform said specified action via a corresponding API. The API / Plug-in Management Module (1423) may provide information obtained from the Knowledge Repository (1440), the Application / Service Module 1450, or another external resource to the Prompt Design Module (1421). That obtained information may be used by the Prompt Design Module (1421) to generate a prompt together with the user input or provided to a generative AI model (1430).

[0206] According to one embodiment, the Output Modification Module (1425) can fine-tune the results obtained through the Generative AI Model (1430) as at least part of the response to user input (e.g., user query). For example, the Output Modification Module (1425) can determine whether the content of the response obtained through the Generative AI Model (1430) is appropriate as a response to a request made by the user input. For example, the Output Modification Module (1425) can determine the degree of relevance, degree of bias (e.g., political or social bias), or degree of harmfulness (e.g., sexual or profanity) of the difference between the response obtained through the Generative AI Model (1430) and the user input. Additionally or generally, the Output Modification Module (1425) can request that additional AI processing be performed on the obtained response, or provide the user with a hint to avoid unwanted output. For example, additional prompts can be generated through the Prompt Design Module to obtain a response again through the Generative AI Model (1430).

[0207] According to one embodiment, the Generative AI Model (1430) may form at least part of an artificial intelligence neural network and may include a model that generates images or a model that generates language. The image generation model may include, for example, a generative adversarial network (GAN), a variational autoencoder (VAE), or a Diffusion-based model using a VAE and a Transformer. The language generation model may include, for example, a large language model (LLM), a large multimodal model (LMM), a large vision model (LVM), or a large action model (LAM). The LAM may automatically generate actions for an environment (e.g., a robot, a car, an electronic device (101), or a program (140)). Additionally, for at least some AI models (e.g., LLM), there may be a low-rank adaptation (LoRA) adaptor fine-tuned for, for example, a specific task or a specific situation.

[0208] FIG. 15 illustrates an AI Framework (1420) having on-device AI processing capabilities according to one embodiment. In this case, the AI ​​Framework (1420) may generate and learn a response to the user input using resources within the device, instead of sending the user input received through a User Interface (1410) operating on the same device (e.g., electronic device (101)) to a Generative AI Model (1430) operating on an external device (e.g., server (108)), or additionally. Referring to FIG. 15, the AI ​​Framework (1420) may include a Cross-Application Action Module (1510), a Personal Data Managing Module (1530), an On-device AI Model (1550), and an Orchestration Module (1570).

[0209] According to one embodiment, the Cross-Application Action Module (1510) determines one or more additional applications required for the operation of an executed application (e.g., an assistant app) and may link or suggest operations between the app and at least one additional application, or between multiple additional applications. For example, the Cross-Application Actions Module (1510) may execute one or more additional applications to be used to respond to a user request through the assistant app sequentially or at least partially and simultaneously. Additionally, the Cross-Application Action Module (1510) may communicate with the additional applications so that the result of the execution of one additional application (e.g., content) can be shared with other additional applications.

[0210] According to one embodiment, the Personal Data Managing Module (1530) may provide personal information (e.g., schedule, contact, or message information) about a user of the application (e.g., assistant app) or the additional application running on the device (e.g., electronic device 101) or other related individuals (e.g., family or friends) to another module of the AI ​​Framework (1420) or a related module (e.g., Generative AI Model 1430) running on another device.

[0211] According to one embodiment, the on-device AI model (1550) may include at least one model among one or more AI models (e.g., GAN, VAE, LLM, LMM, LVM, or LAM) operated on an external device (e.g., server (108)) or a corresponding lightweight AI model. Additionally, for said model or said lightweight model, there may be, for example, a LoRA adaptor.

[0212] According to one embodiment, the Orchestration Module (1570) may select one or more AI models to be used to obtain a response to user input (e.g., user query). For example, the Orchestration Module (1570) may select one or more AI models from among an On-Device AI Model (1550), an AI model operating on an external device (e.g., server (108)) (e.g., Generative AI Model (1430)), or a third AI model (not shown) operating on another external device. When multiple AI models are selected, the Orchestration Module (1570) may communicate with the selected models or devices so that the operation between the selected AI models and the processing of the results thereof can be coordinated between the relevant models or devices.

[0213] According to one embodiment, two or more modules of a generative AI system (1400) (e.g., Cross-Application Action Module (1510) and Orchestration Module (370)) may be implemented as a single module to maintain the same functionality. Various variations are possible.

[0214] According to various embodiments of the present disclosure, upon receiving a message, an artificial intelligence (AI) model can be used to obtain and provide a response message that reflects a mode according to the user's situation, and actions based on the received message and the response message can be automatically performed, which can be convenient.

[0215] The electronic device according to the various embodiments disclosed in this document may be a device of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0216] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, each of phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as “first,” “second,” or “first” or “second” may be used simply to distinguish a component from another component and do not limit the components in any other aspect (e.g., importance or order). Where any (e.g., first) component is referred to as “coupled” or “connected” to another (e.g., second) component, with or without the terms “functionally” or “communicationally,” it means that said component may be connected to said other component directly (e.g., wired), wirelessly, or through a third component.

[0217] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0218] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0219] In a non-transient storage medium storing commands according to one embodiment, the commands are configured to cause the electronic device (101, 201) to perform at least one operation when executed by the electronic device, wherein the at least one operation may include an operation of identifying a first area of ​​the first user interface based on a first user input for the first user interface while displaying a first user interface of a first application on the display of the electronic device. The at least one operation may include an operation of acquiring first data corresponding to the first area. The first data may include at least one of UI component unit data corresponding to the first area, UI spatial data corresponding to the first area, first metadata corresponding to the first area, first image data corresponding to the first area, or first text data corresponding to the first area. The at least one operation may include an operation of acquiring first extracted data including first key image data from the first data through an artificial intelligence model based on first context information for the first data. The at least one operation may include an operation of displaying a second user interface of a second application through the display. The at least one operation may include an operation of obtaining second data corresponding to the second application based on a second user input for applying the first extracted data to the second user interface. The second data may include second metadata of the second application, second image data corresponding to the second user interface, and / or second text data. The at least one operation may include an operation of obtaining second extracted data extracted from the second data through the artificial intelligence model based on second context information regarding the second data.The second extracted data may include one or more UI components applicable to the second user interface and additional information for modifying the first extracted data to be applicable to the second data. The at least one operation may include an operation of obtaining at least one content item through the artificial intelligence model based on the first extracted data and the second extracted data. The at least one operation may include an operation of displaying the first main image data and the at least one content item on the second user interface through the display.

[0220] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0221] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations among the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.< / menu> < / figure> < / menu> < / figure> < / menu> < / figure>

Claims

1. In an electronic device (101, 201), Display(160, 260); Memory for storing commands (130, 230); and It includes at least one processor (130, 230), and When the above commands are executed individually or collectively by the at least one processor, the electronic device, Identifying a first area of ​​the first user interface based on a first user input for the first user interface while displaying the first user interface of the first application on the display, and A first data corresponding to the first area is obtained, and the first data includes at least one of UI component unit data corresponding to the first area, UI spatial data corresponding to the first area, first metadata corresponding to the first area, first image data corresponding to the first area, or first text data corresponding to the first area. Based on first context information regarding the first data, first main image data and first extracted data are obtained from the first data through an artificial intelligence model, and Displays the second user interface of the second application on the above display, and Based on a second user input for applying the first extracted data to the second user interface, second data corresponding to the second application is obtained, and the second data includes second metadata of the second application, second image data corresponding to the second user interface and / or second text data, and Based on second context information regarding the second data, second extracted data extracted from the second data is obtained through the artificial intelligence model, wherein the second extracted data includes one or more UI components applicable to the second user interface and additional information for modifying the first extracted data so that it can be applied to the second data. Based on the first extracted data and the second extracted data, at least one content item is obtained through the artificial intelligence model, and An electronic device that displays the first main image data and the at least one content item on the second user interface through the display.

2. In Paragraph 1, When the above commands are executed individually or collectively by the at least one processor, the electronic device, Normalize the text included in the first extracted data above to obtain normalized text data, and An electronic device that obtains normalized first extracted data corresponding to the first extracted data using the normalized text data.

3. In Paragraph 1 or 2, When the above commands are executed individually or collectively by the at least one processor, the electronic device, When the above at least one content item includes a plurality of content items, the priority of the plurality of content items is identified, and An electronic device that displays the plurality of content items based on the above priority.

4. In any one of paragraphs 1 through 3, When the above commands are executed individually or collectively by the at least one processor, the electronic device, Identifying the application among the applications of the electronic device to be used for obtaining the additional information, and An electronic device that enables the acquisition of the above additional information through the above application.

5. In any one of paragraphs 1 through 4, When the above commands are executed individually or collectively by the at least one processor, the electronic device, Based on the third user input to the second user interface, third main image data and third extracted data for third data corresponding to the second area of ​​the second user interface are obtained, and Displays a third user interface of a third application on the above display, and Based on a fourth user input for applying the third extracted data to the third user interface, the fourth extracted data for the fourth data corresponding to the third area of ​​the third user interface is obtained, and Based on the first extracted data, the second extracted data, the third extracted data, and the fourth extracted data, at least one other content item is obtained through the artificial intelligence model, and An electronic device that displays the third main image data and the at least one other content item on the third user interface through the display.

6. In any one of paragraphs 1 through 5, When the above commands are executed individually or collectively by the at least one processor, the electronic device, An electronic device that, when a first area of ​​the first user interface is identified, displays information of at least one other application capable of applying the first data through the display.

7. In any one of paragraphs 1 through 6, When the above commands are executed individually or collectively by the at least one processor, the electronic device, An electronic device that displays information indicating the previous usage history of the first data when the first area of ​​the first user interface is identified.

8. In any one of paragraphs 1 through 7, When the above commands are executed individually or collectively by the at least one processor, the electronic device, An electronic device that performs optical character recognition (OCR) on the first image data corresponding to the first area to obtain text corresponding to the first image data corresponding to the first area.

9. In any one of paragraphs 1 through 8, An electronic device in which one or more of the above UI components include a menu component, a button component, a checkbox component, a slider component, and / or an image component, and the UI component includes an action, a sub-item, and / or a function associated with the above UI component.

10. In any one of paragraphs 1 through 9, The first metadata is included in the first application and is an electronic device comprising a description, location information, place information, and / or traffic information associated with the first area.

11. A method for transferring and displaying data between applications using an artificial intelligence model in an electronic device (101, 201), An operation of identifying a first area of ​​a first user interface based on a first user input for the first user interface while displaying a first user interface of a first application on a display (160, 260) of the electronic device; An operation of acquiring first data corresponding to the first area, wherein the first data comprises at least one of UI component unit data corresponding to the first area, UI spatial data corresponding to the first area, first metadata corresponding to the first area, first image data corresponding to the first area, or first text data corresponding to the first area; An operation of obtaining first extracted data including first key image data from the first data through the artificial intelligence model based on first context information regarding the first data; The operation of displaying the second user interface of the second application through the above display; An operation of acquiring second data corresponding to the second application based on a second user input for applying the first extracted data to the second user interface, wherein the second data includes second metadata of the second application, second image data and / or second text data corresponding to the second user interface; An operation of obtaining second extracted data extracted from the second data through the artificial intelligence model based on second context information regarding the second data, wherein the second extracted data includes one or more UI components applicable to the second user interface and additional information for modifying the first extracted data so that it can be applied to the second data; An operation of acquiring at least one content item through the artificial intelligence model based on the first extracted data and the second extracted data; and A method including the operation of displaying the first main image data and the at least one content item to the second user interface through the display.

12. In Paragraph 11, An operation to obtain normalized text data by normalizing the text included in the first extracted data; and A method comprising the operation of obtaining normalized first extracted data corresponding to the first extracted data using the normalized text data.

13. In Paragraph 11 or 12, An operation to identify the priority of the plurality of content items when the above at least one content item includes a plurality of content items; and A method including the operation of displaying the plurality of content items based on the above priority.

14. In any one of paragraphs 11 through 13, An operation to identify an application among the applications of the electronic device to be used for acquiring the additional information; A method including the operation of obtaining the above additional information through the above application.

15. In a non-transient storage medium storing instructions, said instructions are set to cause said electronic device (101, 201) to perform at least one operation when said electronic device is executed, said at least one operation being, said at least one operation An operation of identifying a first area of ​​a first user interface based on a first user input for the first user interface while displaying a first user interface of a first application on a display (160, 260) of the electronic device; An operation of acquiring first data corresponding to the first area, wherein the first data comprises at least one of UI component unit data corresponding to the first area, UI spatial data corresponding to the first area, first metadata corresponding to the first area, first image data corresponding to the first area, or first text data corresponding to the first area; An operation of obtaining first extracted data including first key image data from the first data through an artificial intelligence model based on first context information regarding the first data; The operation of displaying the second user interface of the second application through the above display; An operation of acquiring second data corresponding to the second application based on a second user input for applying the first extracted data to the second user interface, wherein the second data includes second metadata of the second application, second image data and / or second text data corresponding to the second user interface; An operation of obtaining second extracted data extracted from the second data through the artificial intelligence model based on second context information regarding the second data, wherein the second extracted data includes one or more UI components applicable to the second user interface and additional information for modifying the first extracted data so that it can be applied to the second data; An operation of acquiring at least one content item through the artificial intelligence model based on the first extracted data and the second extracted data; and A storage medium comprising the operation of displaying the first main image data and the at least one content item to the second user interface through the display.