Electronic device, method, and non-transitory computer-readable recording medium
Patent Information
- Application Number
- PCT/KR2026/003383
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-28
- Filing Date
- 2026-03-03
- Publication Date
- 2026-10-01
Smart Images

Figure KR2026003383_01102026_PF_FP_ABST
Abstract
Description
Electronic device, method and non-transient computer-readable recording medium
[0001] The present disclosure relates to an electronic device, a method, and a non-transient computer-readable recording medium for performing at least one operation based on a user request.
[0002] With the recent advancement of artificial intelligence technology, research and applications of artificial intelligence agents are actively underway. An AI agent can be an intelligent software module capable of making independent judgments based on given inputs or performing tasks through interaction with users. AI agents are being applied in various fields, such as natural language processing, image analysis, recommendation systems, and conversational interfaces, and can improve user convenience by providing customized results based on user needs.
[0003] Artificial intelligence agents can take the form of a multi-agent system structure in which multiple agents cooperate or share roles to operate, and each agent is specialized in a specific function, enabling them to handle complex problems more efficiently. For example, a multi-agent system may include individual agents and a higher-level agent that manages and coordinates the individual agents.
[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0005] According to one embodiment of the present disclosure, an electronic device may be provided. The electronic device may include a memory comprising at least one storage medium in which at least one instruction is stored, and at least one processor capable of executing said at least one instruction. The at least one processor may control a method of operation of the electronic device by executing said at least one instruction.
[0006] According to one embodiment, instructions, when executed individually or collectively by at least one processor, cause an electronic device to: identify at least one operation corresponding to a first user request; determine a first category based on a first operation included in the at least one operation, wherein the first category is included in at least one category for classifying applications; identify whether the first operation can be performed using a predetermined set of instructions; and, based on identifying whether the first operation cannot be performed using the predetermined set of instructions: determine a first application among at least one application included in the first category or a category similar to the first category, based on at least one of the first operation, user input, or personalized data; determine a first operation process corresponding to the first operation based on the first operation and the first application; wherein the first operation process includes instructions for performing one or more operations sequentially or in parallel, and cause the first operation process to be executed through the first application.
[0007] Based on identifying that the first operation cannot be performed using the predetermined set of instructions: among at least one application included in the first category or a category similar to the first category, the first application is determined based on at least one of the first operation, user input, or personalized data, and a first operation process corresponding to the first operation is determined based on the first operation and the first application, wherein the first operation process includes instructions for performing one or more operations sequentially or in parallel, and may cause the first operation process to be executed through the first application.
[0008] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may: determine a second application among at least one application included in the first category based on at least one of the first operation, the user input, or the personalized data, and cause the first operation to be performed through the second application using the determined instruction set, based on identifying that the first operation can be performed using the determined instruction set.
[0009] According to one embodiment, when instructions are executed individually or collectively by at least one processor, the electronic device may: execute the first operation process through the first application and then store the first operation process in the memory, and based on the fact that at least one operation corresponding to a second user request includes the first operation, identify that the first operation can be performed using the predetermined set of instructions and cause the first operation process to be executed through the first application using the stored first operation process.
[0010] According to one embodiment, instructions, when executed individually or collectively by at least one processor, may cause an electronic device to: analyze a screen for executing the first application based on identifying that the first operation cannot be performed using the instruction set; identify at least one first element for the first operation process; and, by interacting with the at least one first element, cause the first operation process to be executed. According to one embodiment, the screen may be displayed on a physical display or may be virtually rendered without being displayed.
[0011] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: identify the first element using an artificial intelligence model, wherein the artificial intelligence model is a model trained to interpret the meaning of a visual element included on a screen, and the visual element may include at least one of text or UI.
[0012] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: identify the first category using an upper agent, transmit the first category information to a lower agent, determine the first application or the second application using the lower agent, and perform the first operation through the first application or execute the first operation process through the second application. The upper agent and the lower agent may include a software program that automatically performs one or more operations based on user input.
[0013] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: transmit the progress of the first operation or the first operation process to the upper agent using the lower agent, and transmit to the lower agent data among the personalized data necessary for the progress of the first operation or the first operation process using the upper agent.
[0014] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may, based on identifying that the first operation cannot be performed using the instruction set: if the first application for executing the first operation process cannot be determined from among the applications stored in the memory, access an application distribution platform through the communication circuit, identify at least one third application downloadable from the application distribution platform capable of executing the first operation process, provide information about the at least one third application, or cause the at least one third application to be downloaded.
[0015] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: determine the priority of at least one application among the at least one application included in the first category based on at least one of the first operation, the user input, or the personalized data, wherein the personalized data includes the execution frequency of the applications, and determine the application with the highest priority among the plurality of applications as the first application.
[0016] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: generate a prompt for inputting into an artificial intelligence model using a first artificial intelligence model based on the user request entered in natural language and the personalized data, and identify the first operation using a second artificial intelligence model based on the prompt. The second artificial intelligence model may include an artificial intelligence model for natural language analysis.
[0017] According to one embodiment, the second artificial intelligence model may include a text analysis module, an application management module, and an action template management module. When the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: identify the first action using the text analysis module based on the prompt, identify at least one application for the first action using the application management module based on the prompt, obtain one or more action templates using the action template management module based on the prompt, the first action, and the at least one application, wherein the action template includes actions to be performed sequentially or in parallel and information necessary to perform said actions, and cause the first action to be performed based on said action template.
[0018] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to: identify the first operation based on analyzing the meaning of a sentence included in the prompt using the text analysis module; identify the highest priority application among applications included in a category suitable for performing the first operation using the application management module; obtain one or more operation templates using the operation template management module based on the prompt, the first operation, and the highest priority application; and perform the first operation by searching for necessary elements on the screen and interacting with specific elements based on the one or more operation templates.
[0019] According to one embodiment of the present disclosure, a method of operating an electronic device may be provided. The method of operating an electronic device may include at least one operation. The at least one operation may include: an operation of identifying at least one operation corresponding to a first user request; an operation of determining a first category based on a first operation included in the at least one operation, wherein the first category is included in at least one category for classifying applications; an operation of identifying whether the first operation can be performed using a predetermined set of instructions; an operation of determining a first application based on at least one of the first operation, user input, or personalized data among at least one application included in the first category or a category similar to the first category, based on identifying that the first operation cannot be performed using the predetermined set of instructions; an operation of determining a first operation process corresponding to the first operation based on the first operation and the first application, wherein the first operation process includes instructions for performing one or more operations sequentially or in parallel; and / or an operation of executing the first operation process through the first application.
[0020] According to one embodiment, a method of an electronic device may further include: an action of determining a second application among at least one application included in the first category based on at least one of the first operation, the user input, or the personalized data; and an action of performing the first operation through the second application using the determined set of instructions, based on identifying that the first operation can be performed using the determined set of instructions.
[0021] According to one embodiment, a method of an electronic device may further include: an action of storing the first action process in the memory after executing the first action process through the first application; an action of identifying that the first action can be performed using the predetermined set of instructions based on at least one action corresponding to a second user request including the first action; and an action of executing the first action process through the first application using the stored first action process.
[0022] According to one embodiment, the operation of executing the first operation process through the first application may include: an operation of identifying at least one first element for the first operation process by analyzing a screen resulting from the execution of the first application; and an operation of executing the first operation process by interacting with the at least one first element.
[0023] According to one embodiment, the operation of identifying at least one first element may include the operation of identifying the first element using an artificial intelligence model. The artificial intelligence model is a model trained to interpret the meaning of a visual element included on a screen, and the visual element may include at least one of text or UI.
[0024] According to one embodiment, a method of an electronic device may further include: an operation of identifying the first category using an upper agent; an operation of transmitting the first category information to a lower agent; an operation of determining the first application or the second application using the lower agent; and an operation of performing the first operation through the first application or executing the first operation process through the second application. The upper agent and the lower agent may include a software program that automatically performs one or more operations based on user input.
[0025] According to one embodiment, a non-transient computer-readable recording medium may be provided for storing instructions that, when executed by at least one processor, cause the at least one processor to perform set operations. The operations include: an operation of identifying at least one operation corresponding to a first user request; an operation of determining a first category based on a first operation included in the at least one operation, wherein the first category is included in at least one category for classifying applications; an operation of identifying whether the first operation is executable using a predetermined set of instructions; and, based on identifying that the first operation is not executable using the predetermined set of instructions: an operation of determining a first application among at least one application included in the first category or a category similar to the first category, based on at least one of the first operation, user input, or personalized data; and an operation of determining a first operation process corresponding to the first operation based on the first operation and the first application, wherein the first operation process includes instructions for performing one or more operations sequentially or in parallel. and may include at least one of the operation of executing the first operation process through the first application.
[0026] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0027] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment.
[0028] FIG. 2 is a block diagram of a generative artificial intelligence (AI) system according to one embodiment.
[0029] FIG. 3 is a block diagram of an AI framework according to one embodiment.
[0030] FIG. 4 is a drawing for explaining an artificial intelligence agent system according to one embodiment of the present disclosure.
[0031] FIG. 5 illustrates the configuration of an artificial intelligence agent according to one embodiment of the present disclosure.
[0032] FIG. 6 illustrates personalized data according to one embodiment of the present disclosure.
[0033] FIG. 7a is a flowchart illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0034] FIG. 7b is a drawing for explaining the operation according to one embodiment of the present disclosure.
[0035] FIG. 7c is a drawing for explaining the operation according to one embodiment of the present disclosure.
[0036] FIGS. 8a to 8c illustrate a screen displayed by an electronic device according to embodiments of the present disclosure.
[0037] FIGS. 9a to 9c illustrate a screen displayed by an electronic device according to embodiments of the present disclosure.
[0038] FIGS. 10a and FIGS. 10b illustrate a screen displayed by an electronic device according to embodiments of the present disclosure.
[0039] FIGS. 11a and FIGS. 11b illustrate a screen displayed by an electronic device according to embodiments of the present disclosure.
[0040] FIG. 12 is a drawing for explaining the operation of an electronic device according to one embodiment of the present disclosure.
[0041] FIG. 13a is a drawing for explaining the operation using a mixer of an electronic device according to one embodiment of the present disclosure.
[0042] FIG. 13b is a drawing for illustrating an operation template according to one embodiment of the present disclosure.
[0043] FIGS. 14a to 14g illustrate the screens of an electronic device according to embodiments of the present disclosure.
[0044] FIG. 15 illustrates an operation using a screen analysis and interaction system of an electronic device according to one embodiment of the present disclosure.
[0045] FIG. 16 is a drawing for explaining the operation of an electronic device according to one embodiment of the present disclosure.
[0046] FIG. 17 illustrates the operation of an electronic device using a processor and a display according to one embodiment of the present disclosure over time.
[0047] FIG. 18 is a flowchart illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0048] FIG. 19 illustrates the operation of a processor and a display module of an electronic device according to one embodiment of the present disclosure over time.
[0049] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and should be understood to include various modifications, equivalents, or substitutions of the embodiments described herein, rather than being limited to the embodiments described herein. The present disclosure is capable of various modifications by those skilled in the art without departing from the gist of the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.
[0050] The purposes and effects of the present disclosure are not limited to those mentioned in the drawings and the following related description, and various modifications may be made within the technical scope of the present disclosure. The effects according to the embodiments of the present disclosure mentioned below are merely illustrative and are not limited thereto; depending on various modifications, different or additional effects may be realized.
[0051] In the following drawings and related descriptions, functions, configurations, technical terms, and technical details well known in the art to which this disclosure pertains may be omitted. This is intended to convey the essentials of this disclosure more clearly and concisely by minimizing unnecessary detailed descriptions.
[0052] In the drawings, each block of the flowcharts and combinations of the flowcharts may be performed by at least one instruction. The instruction may be loaded into a processor of a computer or other programmable data processing equipment to generate means for performing the functions described in the drawings. The instruction may also provide steps for performing the functions described in the drawings by being executed on a computer or other programmable data processing equipment.
[0053] Meanwhile, various elements and regions in the drawings are depicted schematically, and the technical concept of the present disclosure is not limited by the relative sizes, spacing, or arrangements depicted in the attached drawings. The electronic device of the present disclosure is not limited to the configuration and / or operation shown in the drawings and may include all other configurations capable of performing the same or similar functions.
[0054] The individual components depicted in the drawings are not required to be implemented in a physically separate form, but are shown separately to aid in the description and understanding of the present disclosure. The present disclosure may be implemented in a form in which the individual components shown in the drawings are merged, modified, or have some components deleted and / or added. Each component may perform functions in conjunction with one another while existing in physically separated locations via a network or communication link.
[0055] Likewise, the operations depicted in the drawings are illustrative to aid in the description and understanding of the present disclosure, and the present disclosure may be modified by merging, changing the order of, or deleting and / or adding parts of the operations shown in the drawings. For example, two or more operations shown consecutively in the drawings may be performed simultaneously, in reverse order as necessary, repeatedly, or omitted depending on the actual situation.
[0056] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0057] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), secure processing unit (SPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0058] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence is performed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network (DQN), or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0059] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0060] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146). One or more related applications may form a service configured to handle a series of user requests.
[0061] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0062] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0063] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0064] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0065] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0066] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0067] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0068] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0069] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0070] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) may be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0071] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0072] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., local area network (LAN) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0073] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. The NR access technology can support enhanced mobile broadband (Embb), massive machine type communications (mMTC), or ultra-reliable and low-latency communications (URLLC). The wireless communication module (192) can support high-frequency bands (e.g., mmWave bands) to achieve high data transmission rates, for example. The wireless communication module (192) can support various technologies for securing performance in high-frequency bands, for example, beamforming, multiple-input and multiple-output (massive MIMO), full-dimensional MIMO (FD-MIMO), array antennas, analog beamforming, or large-scale antennas. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., electronic device (104)), or a network system (e.g., a second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0074] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0075] According to one embodiment, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0076] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, general purpose input and output (GPIO), serial peripheral interface (SPI), or mobile industry processor interface (MIPI)) and exchange signals (e.g., commands or data) with each other.
[0077] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least a part of the requested function or service, or additional functions or services related to the request, and transmit the result of the execution to the electronic device (101).
[0078] The electronic device (101) may process the above results as they are or additionally and provide them as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services, for example, by using distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199). The electronic device (101) may be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0079] FIG. 2 is a generative artificial intelligence (AI) system (200) according to one embodiment. Referring to FIG. 2, the generative AI system (200) may include a user interface (210), an AI framework (220), a generative AI model (230), a knowledge repository (240), and an application / service module (250). These components may be operated on one or more of an electronic device (101), an external electronic device (102 or 104), or a server (108). For example, the user interface (210) and the AI framework (220) may be operated on the electronic device (101), and the knowledge repository (240) and the generative AI model (230) may be operated on the server (108).
[0080] According to one embodiment, a user interface (210) may receive user input (e.g., user query). User input may be received in the form of text, images, voice (e.g., natural language), video, menu selection, or a combination thereof. The user interface (210) may include various context information (e.g., running application or user location) related to the generative artificial intelligence system (200) at the time the user input is received, in addition to or instead of the user input. The user interface (210) may provide the user input or the context information to the AI framework (220) and provide the result of processing therefrom to the user, for example, through the AI framework (220). According to one embodiment, in addition to user input, the electronic device may provide context information obtained using information included on the screen to the AI framework (220). The result may be provided in the form of text, images, voice, video, an action requested by the user (e.g., execution of a specified function or app), or a combination thereof.
[0081] According to one embodiment, the AI framework (220) can identify (e.g., estimate) a user intent based on at least some user input or context information received from a user interface (210), control each of the relevant modules (e.g., 221, 223, or 225) to perform a function or action corresponding to the identified user intent, and coordinate collaboration between two or more modules. The AI framework (220) may include a prompt design module (221), an API / plugin management module (223), and an output modification module (225), as illustrated in FIG. 2.
[0082] According to one embodiment, the prompt design module (221) can generate a prompt to be input to a generative AI model (230) based at least partially on user input or context information received from a user interface (210). For example, the prompt design module (221) can generate a prompt using user preferences, a prompt library, or prompt examples stored in a knowledge repository (240) based at least partially on user input or context information.
[0083] According to one embodiment, an API / plug-in management module (223) may communicate, for example, via an API, with various resources (e.g., a knowledge repository (240)) that provide said additional information when there is a request for said additional information in relation to user input. Additionally or alternatively, when a specified action (e.g., a function, app, or service) is performed in response to said user input, the API / plug-in management module (223) may request the application / service module (250) to perform said specified action via a corresponding API. The API / plug-in management module (223) may provide information obtained from the knowledge repository (240), the application / service module 250, or another external resource to the prompt design module (221). The obtained information may be used by the prompt design module (221) to generate a prompt together with the user input or provided to a generative AI model (230).
[0084] According to one embodiment, the output processing module (225) can fine-tune the results obtained through the generative AI model (230) as at least part of the response to user input (e.g., user query). For example, the output processing module (225) can determine whether the content of the response obtained through the generative AI model (230) is appropriate as a response to a request made by the user input. For example, the output processing module (225) can determine the degree of relevance, degree of bias (e.g., political or social bias), or degree of harmfulness (e.g., sexual or profanity) of the difference between the response obtained through the generative AI model (230) and the user input. Additionally or generally, the output processing module (225) can request that additional AI processing be performed on the obtained response, or provide the user with a hint to avoid unwanted output. For example, the response can be obtained again through the generative AI model (230) by generating additional prompts through the prompt design module.
[0085] According to one embodiment, the generative AI model (230) may form at least part of an artificial intelligence neural network and may include a model that generates images or a model that generates language. The image generation model may include, for example, a generative adversarial network (GAN), a variational autoencoder (VAE), or a Diffusion-based model using a VAE and a Transformer. The language generation model may include, for example, a large language model (LLM), a large multimodal model (LMM), a large vision model (LVM), or a large action model (LAM). The LAM may automatically generate actions for an environment (e.g., a robot, a car, an electronic device (101), or a program (140)). Additionally, for at least some AI models (e.g., LLM), there may be a low-rank adaptation (LoRA) adapter fine-tuned for, for example, a specific task or a specific situation.
[0086] FIG. 3 illustrates an AI framework (220) having on-device AI processing capabilities according to one embodiment. In this case, the AI framework (220) may generate and learn a response to the user input using resources within the device, instead of sending the user input received through a user interface (210) operating on the same device (e.g., electronic device (101)) to a generative AI model (230) operating on an external device (e.g., server (108)), or additionally. Referring to FIG. 3, the AI framework (220) may include a cross-application action module (310), a personal data management module (330), an on-device AI model (350), and an orchestration module (370).
[0087] According to one embodiment, the cross-application action module (310) determines one or more additional applications required for the operation of an executed application (e.g., an assistant app) and may connect or suggest operations between the app and at least one additional application, or between a plurality of additional applications. For example, the cross-application action module (310) may execute one or more additional applications to be used to respond to a user request through the assistant app sequentially or at least partially and simultaneously. Additionally, the cross-application action module (310) may communicate with the additional applications so that the result of the execution of one additional application (e.g., content) can be shared with other additional applications.
[0088] According to one embodiment, the personal data management module (330) may provide personal information (e.g., schedule, contact, or message information) about a user of the application (e.g., assistant app) or the additional application running on the device (e.g., electronic device 101) or other related individuals (e.g., family or friends) to another module of the AI framework (220) or a related module (e.g., generative AI model (230)) running on another device.
[0089] According to one embodiment, the on-device AI model (350) may include at least one model among one or more AI models (e.g., GAN, VAE, LLM, LMM, LVM, or LAM) operated on an external device (e.g., server (108)) or a corresponding lightweight AI model. Additionally, for said model or said lightweight model, there may be, for example, a LoRA adapter.
[0090] According to one embodiment, the orchestration module (370) may select one or more AI models to be used to obtain a response to user input (e.g., user query). For example, the orchestration module (370) may select one or more AI models from an on-device AI model (350), an AI model operating on an external device (e.g., server (108)) (e.g., generative AI model (230)), or a third AI model (not shown) operating on another external device. When multiple AI models are selected, the orchestration module (370) may communicate with the selected models or devices so that the operation between the selected AI models and the processing of the results thereof can be coordinated between the relevant models or devices.
[0091] According to one embodiment, two or more modules of a generative AI system (200) (e.g., a cross-application action module (310) and an orchestration module (370)) may be implemented as a single module to maintain the same functionality. Various variations are possible.
[0092] The 'artificial intelligence model' in the present disclosure below may be identical or similar to the configuration of the generative AI model (230) of FIG. 2 or the on-device AI model (350) of FIG. 3, or may include an artificial intelligence model included in the generative AI model (230) of FIG. 2 or the on-device AI model (350) of FIG. 3. The 'artificial intelligence system' in the following description may be the generative artificial intelligence system (200) of FIG. 2.
[0093] FIG. 4 is a drawing for explaining an artificial intelligence agent system (400) according to one embodiment of the present disclosure.
[0094] Referring to FIG. 4, according to one embodiment, an artificial intelligence agent system (400) may include or be composed of at least one of a query filtering model (410), a higher-level AI agent (420), or a lower-level AI agent (430). The artificial intelligence agent system is based on a multi-agent structure and enables each lower-level AI agent to efficiently distribute and perform tasks required to set relationship information regarding an object by using a higher-level AI agent (420) that manages each of the lower-level AI agents included in the lower-level AI agent (430). Once relationship information regarding an object is set, the artificial intelligence agent system can provide convenience functions by providing augmented content corresponding to the registered object for which relationship information regarding the object has been set.
[0095] According to one embodiment, a query filtering model (410) may receive a query entered through an input window (e.g., a search window), analyze the query, and output a filtered result. For example, a user may search for an object in the main search window of an electronic device (e.g., the electronic device (101) of FIG. 1) (e.g., displayed by a slide-up or slide-down gesture input while running a home application) and / or in the search window of an application (e.g., a gallery application, a messenger application, and / or a social media application). The query may be input data in the form of natural language entered by the user. The query may be text data. The query filtering model (410) may receive queries in various languages. The query filtering model (410) may take image data and / or text data as input and output text data. The query filtering model (410) may filter the query using an artificial intelligence model and output the filtered text data. The query filtering model (410) may filter text corresponding to an object using an artificial intelligence model. The output text data may include words or sentences in the form of natural language. The query filtering model (410) may receive image data or text data as input, or receive both different types of data simultaneously. The query filtering model (410) may receive text data describing the image data along with the image data. Text data describing the image data may be obtained through image captioning. The query filtering model (410) may receive text data converted from voice data.
[0096] According to one embodiment, the query filtering model (410) may analyze the query independently or analyze the query based on image data input along with the query, and correct errors in the query. The query filtering model (410) may include a large language model (LLM). The LLM may perform pre-training using a vast amount of corpus data consisting of images and text. For example, the LLM may include, but is not limited to, GPT, LLaMA, and / or Bard.
[0097] According to one embodiment, an upper AI agent (420) can manage lower AI agents. The upper AI agent (420) can assign tasks to lower AI agents so that lower AI agents can set relationship information regarding objects, and can collect data output by each of the lower AI agents. The lower AI agents may be AI agents (431, 432) included in the lower AI agent (430).
[0098] According to one embodiment, the upper AI agent (420) can manage the AI agents (431, 432) within a circular structure in which the AI agents (431, 432) update information. The upper AI agent (420) can receive a command agenda received from a user, a rule controlling the lower AI agent (430), and / or output data of the lower AI agent (430).
[0099] According to one embodiment, the electronic device (101) can correct a command proposal in the form of natural language received from a user into a more specific and clear form using an artificial intelligence model (e.g., a large language model). An upper artificial intelligence agent (420) can receive the corrected command proposal and determine at least one corresponding action. The upper artificial intelligence agent (420) can divide and assign the command to each lower artificial intelligence agent (e.g., a first artificial intelligence agent (431), a second artificial intelligence agent (432)) suitable for performing each of the at least one action.
[0100] According to one embodiment, the upper artificial intelligence agent (420) may, as needed, access data (e.g., personalized data, authentication information) within the electronic device (101) to which the lower artificial intelligence agent (430) has restricted access, and provide the necessary data to the lower artificial intelligence agent (430).
[0101] According to one embodiment, a rule for controlling a subordinate artificial intelligence agent (430) may include, for example, a termination condition indicating that the setting of relationship information regarding an object has ended. A rule for controlling a subordinate artificial intelligence agent (430) may be set in advance. A rule for controlling a subordinate artificial intelligence agent (430) may be set by user input. The output data of the subordinate artificial intelligence agent (430) may include the output data of the artificial intelligence agents (431, 432) included in the subordinate artificial intelligence agent (430).
[0102] According to one embodiment, the upper AI agent (420) can determine the progress of the work performed by the AI agents (431, 432) based on the output data of the lower AI agent (430). Based on the output data of the lower AI agent (430), if the termination condition of the work performed by the AI agents (431, 432) is not satisfied, the upper AI agent (420) can output a command prompt to each of the AI agents (431, 432). If the termination condition of the work performed by the AI agents (431, 432) is satisfied, the upper AI agent (420) can output the result obtained by the AI agents (431, 432) (e.g., relationship information regarding an object) in text form.
[0103] According to one embodiment, the upper AI agent (420) may be based on a multi-agent architecture responsible for task instructions and progress tracking to manage AI agents (431, 432). The upper AI agent (420) may formulate a plan to process tasks and collect data required for each operation cycle using data required for the task and a model trained through pre-training. The upper AI agent (420) may obtain the progress of the task for each cycle of the operation plan of the AI agents (431, 432) and provide feedback thereon.
[0104] According to one embodiment, while the upper AI agent (420) and the lower AI agent (430) are performing a task based on a user request, the electronic device (101) may provide a user interface that indicates that the task based on the user request is in progress through a display (e.g., the display module (160) of FIG. 1). When the upper AI agent (420) and the lower AI agent (430) complete the task based on the user request, the electronic device (101) may provide a user interface that allows monitoring of the task results.
[0105] According to one embodiment, the upper AI agent (420) can check whether the AI agents (431, 432) have satisfied the task completion conditions as previously learned (or as specified by the user). If the AI agents (431, 432) have not completed the task, the upper AI agent (420) can assign the task to some or all of the AI agents (431, 432). When the AI agents (431, 432) complete the task, the upper AI agent (420) updates the information based on the output data of the AI agents (431, 432) and can repeat this action until the entire task is completed. If the upper AI agent (420) determines that the AI agents (431, 432) have not satisfied the task completion conditions, it can establish a new plan (e.g., acquiring other information about the object and image clustering using the acquired other information) and have the AI agents (431, 432) re-execute the task.
[0106] According to one embodiment, the subordinate AI agent (430) may include one or more AI agents. For example, the subordinate AI agent (430) may include at least one of a first AI agent (431) or a second AI agent (432), or may be composed of these. The first AI agent (431) may manage images stored in the user's personalized data (e.g., images in a gallery application). The second AI agent (432) may manage information regarding objects stored in the user's personalized data (e.g., personalized data related to objects).
[0107] According to one embodiment, each artificial intelligence agent (431 and 432) can perform tasks independently depending on the type of application or data type installed in the electronic device (101). When using an artificial intelligence agent that independently performs specialized functions, processing various types of data can be relatively effective compared to using an artificial intelligence model that performs only a specific purpose.
[0108] According to one embodiment, each of the artificial intelligence agents (431 and 432) can perform learning using a predefined dataset. In this case, the first and second artificial intelligence agents (431 and 432) can adaptively perform special tasks according to the input data. The lower artificial intelligence agent (430) can contextually activate the artificial intelligence agents (431, 432) and provide contextually appropriate training data pairs to perform tasks for different applications and / or functions in accordance with user requests transmitted through the upper artificial intelligence agent (420).
[0109] According to one embodiment, each agent included in the sub-artificial intelligence agent (430) may correspond to a detailed category for classifying applications. For example, each sub-agent may be an agent for operation through an application of the corresponding category. For example, the first artificial intelligence agent (431) may correspond to the first category (e.g., travel-related application category) and the second artificial intelligence agent (432) may correspond to the second category (e.g., photography-related application category). Although the sub-artificial intelligence agent (430) of FIG. 4 is illustrated as including two artificial intelligence agents, this is merely an example, and the sub-artificial intelligence agent (430) of the present disclosure may include three or more artificial intelligence agents.
[0110] According to one embodiment, a lower AI agent (430) can determine one of the applications belonging to a specific category based on a command divided and assigned by an upper AI agent (420) and perform at least one operation through the determined application. For example, the upper AI agent (420) transmits a command to perform a first operation to a first AI agent (431) corresponding to a first category, and the first AI agent (431) can analyze the first operation and determine a first application to perform the first operation within the first category. The first AI agent (431) can perform the first operation through the first application.
[0111] According to one embodiment, a subordinate artificial intelligence agent (430) can manage raw data received by an electronic device (101). A second artificial intelligence agent (432) can manage all raw data that can be obtained through a communication module (e.g., Wi-Fi and / or Bluetooth) and / or an external electronic device (e.g., a wearable device) while a user performs a specific task through the electronic device (101) or obtains and / or generates specific information using an application. The second artificial intelligence agent (432) can efficiently manage data through a data processing unit (DPU) instead of a central processing unit (CPU).
[0112] According to one embodiment, a subordinate AI agent (430) may receive a data extraction request signal. The subordinate AI agent (430) may receive a data extraction request signal when a query regarding an object is detected. In response to the receipt of the data extraction request signal, the subordinate AI agent (430) may extract all data associated with the query regarding the object and / or images from the user's personalized data.
[0113] According to one embodiment, a subordinate artificial intelligence agent (430) may use a large language model (LM) to extract data from the user's personalized data. The LLM may be one of the modules that manage the user's personalized data. The subordinate artificial intelligence agent (430) may transmit words entered by the user to the LLM in the form of a prompt. The LLM may obtain, for example, images, text, voice, time information, location information, and / or health information from the user's personalized data.
[0114] According to one embodiment, since the upper AI agent (420) can refer to the data acquired by the lower AI agent (430) while operating based on a circular structure, the data acquired by the lower AI agent (430) can be stored and managed in the form of a JSON file for easy management. The lower AI agent (430) can extract information about objects from the user's personalized data in the form of a JSON file. For example, the data acquired by the lower AI agent (430) can be stored and managed in the following file format.
[0115] "Image"=[DCIM / users / sub_img_1.jpg, DCIM / users / sub_img2.jpg, … ]
[0116] "Text"={"Kakao": 2024_10_30_talk.txt, 2024_11_02_talk.txt
[0117] "Instagram": 2024_10_29_dm.txt, 2024_11_13_dm.txt}
[0118] “GeoDB”=[…] / users / 20241013_loc.db, … / users / 20241017_loc.db]
[0119] According to one embodiment, if the data acquired by the subordinate artificial intelligence agent (430) is managed in the form of a JSON file, it may be easy to modify the original data stored in the user's personalized data. Since the data acquired by the subordinate artificial intelligence agent (430) is data used exclusively when setting relationship information regarding an object, it is not necessary to include the original data. For example, the memory usage required when setting relationship information regarding an object may be reduced.
[0120] According to one embodiment, a lower AI agent (430) can transmit acquired data to a higher AI agent (420). The higher AI agent (420) can analyze the received data (e.g., classify data types) and transmit the data along with work details to each of the lower AI agents (431 and 432). The first AI agent (431) and the second AI agent (432) can, for example, sequentially analyze the data they manage in a Chain-of-Thought (CoT) manner and perform the work instructed by the higher AI agent (420).
[0121] According to one embodiment, the supervising AI agent may be referred to as a supervisory AI agent. According to one embodiment, the subordinate AI agent may be referred to as a subordinate AI agent.
[0122] According to one embodiment, the electronic device (101) may additionally include a personal DB updater (not shown) and / or an input box manager (not shown). The personal DB updater receives data acquired by a subordinate artificial intelligence agent (430) and can update (or update) relationship information regarding objects. The input box manager can control an input box (e.g., a search box and / or a chat box). The input box manager can receive a query entered into the input box.
[0123] FIG. 5 illustrates the configuration of an artificial intelligence agent (500) according to one embodiment of the present disclosure.
[0124] In FIG. 5, the artificial intelligence agent (500) may include or be composed of at least one of a configurator module (510), a memory module (520), a perception module (530), an execution module (540), or a critique-reflection module (550). For example, the artificial intelligence agent (500) may be a management artificial intelligence agent (the parent artificial intelligence agent (420) in FIG. 4). The artificial intelligence agent (500) may be the first artificial intelligence agent (431) and / or the second artificial intelligence agent (432) in FIG. 4.
[0125] According to one embodiment, the configuration module (510) can set the role that the artificial intelligence agent (500) will perform within the device. The configuration module (510) can set initial conditions that cause the artificial intelligence agent (500) to perform an operation.
[0126] According to one embodiment, the memory module (520) may provide a function to store and retrieve data change and / or generation processes. The memory module (520) may acquire data flow processes within the artificial intelligence agent (500) as data, thereby enabling the artificial intelligence agent (500) to share data with other artificial intelligence agents.
[0127] According to one embodiment, the recognition module (530) observes the situation regarding data that changes during the operation of the artificial intelligence agent (500) and can obtain the situation regarding the changing data as text data.
[0128] According to one embodiment, the execution module (540) can substantially execute the action that the artificial intelligence agent (500) has decided to execute.
[0129] According to one embodiment, the feedback module (550) can update model parameters on its own. The artificial intelligence agent (500) can converse with other artificial intelligence agents (500). Through conversation, the feedback module (550) can provide feedback to the artificial intelligence agent (500) to identify problems that occur during operation and to ensure that the operation is executed smoothly. For example, a management artificial intelligence agent (the parent artificial intelligence agent (420) in FIG. 4) can receive output data from the artificial intelligence agents (431, 432) and determine whether the termination condition is satisfied.
[0130] FIG. 6 illustrates personalized data (600) according to one embodiment of the present disclosure.
[0131] Referring to FIG. 6, according to one embodiment, personalized data (600) may include image data (610), audio data (620), operation history (630) of an electronic device (e.g., electronic device (101) of FIG. 1), user preference data (640) and / or user context data (650), and may include all other user-related data, although not illustrated.
[0132] According to one embodiment, the image data (610) may include image or video data associated with the user, for example, a photograph taken by the user, a captured screen image and / or real-time video data obtained through an external camera.
[0133] According to one embodiment, the audio data (620) may include the user's voice, ambient sounds and / or voice interaction data with the electronic device (101), and may include, for example, voice commands, call recordings and / or background noise data.
[0134] According to one embodiment, the operation history (630) of the electronic device may include a record of user interaction with the electronic device (101) and / or a record of operations performed by the electronic device (101). For example, the operation history (630) of the electronic device may include usage history by application, usage frequency by application, payment history, touch input patterns, gestures and / or activation history of specific functions. A record of operations performed by the electronic device (101) may correspond to a record of commands executed by the electronic device (101).
[0135] According to one embodiment, the operation history (630) of the electronic device may be dynamically updated. For example, the electronic device (101) may receive a specific user request and perform at least one operation sequentially or simultaneously, and the electronic device (101) may store a record of the performed operation in personalized data (600). For example, the electronic device (101) may perform a specific operation process consisting of commands, and the electronic device (101) may store a set of executed commands in personalized data (600).
[0136] According to one embodiment, user preference data (640) may include data related to the user's tendencies and customized settings. User preference data (640) may include, for example, frequently used applications, preferred content types, frequently used functions, preference information entered by the user (text input regarding direct preference / dislike, or input such as SNS likes / dislikes from which preference / dislike can be inferred indirectly), and / or personalized recommendation records.
[0137] According to one embodiment, user context data (650) may include environment and situation information related to the user, and may include context data such as the user's current location, schedule, weather information, date, time, biometric data collected from a wearable device and / or whether the user is walking or driving.
[0138] According to one embodiment, an electronic device (101) requires user authentication (e.g., login) to access personalized data (600), and can access the personalized data (600) based on the user performing authentication. After the user performs authentication, the electronic device (101) can also obtain and use personalized data (600) stored in an external cloud. User authentication can be performed in various ways, such as direct input of authentication information, biometric recognition (e.g., fingerprint recognition, facial recognition), and voice recognition.
[0139] FIG. 7a is a flowchart illustrating the operation of an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment of the present disclosure.
[0140] Referring to FIG. 7a, in operation 710, the electronic device (101) can identify at least one operation (or task) corresponding to the first user request. For example, the electronic device (101) can receive a user request such as “book a hotel that fits my March travel schedule” and identify “an operation to check the travel schedule (place and date) from user personalization data” and “an operation to book a hotel stay that fits the schedule.” The electronic device (101) can identify at least one operation from a user request in natural language format, for example, by using an artificial intelligence model that analyzes natural language.
[0141] According to one embodiment, the electronic device (101) can utilize user requests by further refining them (e.g., prompt engineering). For example, as the electronic device (101) performs LLM-based prompt engineering, it can modify the user request entered in natural language into a prompt in a natural language form that can give more specific commands than the natural language request.
[0142] According to one embodiment, the electronic device (101) can perform prompt engineering using an artificial intelligence agent system (e.g., the artificial intelligence agent system (400) of FIG. 4). For example, the electronic device (101) can obtain a prompt by modifying a filtered user request more specifically and clearly using a first artificial intelligence model (e.g., a query filtering model (410)) based on user requests and personalized data entered in natural language, and input the prompt into a second artificial intelligence model (e.g., a higher artificial intelligence agent (420)) to identify a first action. For example, if the user request is "recommend lunch nearby," the electronic device (101) can generate a more specific and personalized prompt, such as "recommend a Korean restaurant within a 10-minute walk from the current location that is open at 13:00 today and has a rating of 4.5 or higher," by reflecting location information, past restaurant selection history, preferred food types, and meal times among the user's personalized data.
[0143] According to one embodiment, operation 710 may be performed using a mixer (e.g., the mixer (1320) of FIG. 13a). For example, the electronic device (101) may identify at least one operation corresponding to a first user request using the mixer. The mixer may be a set of applications or functions of applications for performing at least one operation. The operation of the electronic device (101) corresponding to embodiments in which the electronic device (101) uses the mixer is described in detail in the description of FIG. 13a.
[0144] In operation 720, the electronic device (101) can determine a first category among at least one category for classifying applications based on a first operation included in at least one operation. For example, the electronic device (101) can determine a “user schedule-related application category” based on the “operation of checking a travel schedule from user personalized data.”
[0145] In operation 730, the electronic device (101) can identify whether the first operation can be performed using a defined instruction set. The defined instruction set may include, for example, an application programming interface (API), a software instruction set, a script function definition, and / or external system interaction instructions. The electronic device (101) can perform at least one operation by executing at least a portion of the defined instruction set (e.g., by executing on a processor (e.g., the processor (120) of FIG. 1 or through the artificial intelligence agent system (400) of FIG. 4).
[0146] According to one embodiment, the electronic device (101) may store commands for automatically performing at least one operation as a predetermined set of commands. For example, the electronic device (101) may store commands in advance for "an operation of querying user calendar data corresponding to a specific date, parsing the data into a structured text format (e.g., JSON or XML), and returning it."
[0147] According to one embodiment, a defined instruction set may include instructions for executing a specific application itself, as well as instructions for performing operations to execute a specific function of the specific application. Even for the same application, the electronic device (101) may manage a related instruction set (instructions included in the defined instruction set) by distinguishing it in units of multiple executable functions of the application. The defined instruction set may include, for example, a first instruction set corresponding to (or mapped to) a first function of a first application, a second instruction set corresponding to a second function of the first application, and may not include an instruction set corresponding to a third function of the first application.
[0148] According to one embodiment, based on the electronic device (101) distinguishing and managing instruction sets for each function of each application, the result of determining operation 730 may differ by function even for the same application. For example, if the first operation is an operation that executes the first function of the first application and a first instruction set corresponding to the first function of the first application is defined, the electronic device (101) can identify that the first operation can be performed using the defined instruction set. For example, if the first operation is an operation that executes the third function of the first application and a instruction set corresponding to the third function of the first application is not defined, the electronic device (101) can identify that the first operation cannot be performed using the defined instruction set.
[0149] According to one embodiment, operation 730 may be expressed as an operation that identifies whether the first operation can be performed using at least one application included in the first category. Operation 730 may also be expressed as an operation that identifies whether the first operation can be performed using an artificial intelligence agent (e.g., the artificial intelligence agent (500) of FIG. 5).
[0150] In operation 740, the electronic device (101) may determine a first application included in a first category or a similar category based on identifying that the first operation cannot be performed using a defined set of instructions. The electronic device (101) may determine a first application among the applications included in the first category or a similar category based on at least one of the first operation, user input, or personalized data. The first application may be an application determined to be most suitable for performing the first operation. The electronic device (101) may not determine a first application based on determining that there is no application suitable for performing the first operation.
[0151] According to one embodiment, a category similar to the first category may include a category to which an application belongs that is not the first category but provides a function similar to the first operation or provides a function that can be used complementarily to achieve the purpose of the first operation. For example, if the first category is a schedule management category, the similar category may include a memo category and a calendar category. For example, if the first category is a message category, the similar category may include an SNS category and a messenger category. For example, if the first category is an accommodation reservation category, the similar category may include a travel-related category.
[0152] According to one embodiment, an electronic device (101) can determine the priority of at least one application included in a first category based on a first action, user input, and / or personalized data (e.g., application execution frequency) among at least one application included in a first category. The electronic device (101) can determine the application with the highest priority as the first application. For example, the electronic device (101) can determine the application with the highest usage frequency among at least one application included in the first category as the first application. For example, the electronic device (101) can analyze the content of the first action and, if the first action is "generating a summary based on specific keywords within a PDF document," select a document analysis application (e.g., an AI-based document summary app) that provides specialized functions for the first action as the first application, rather than a document viewer app with the highest usage frequency. For example, the electronic device (101) can determine a specific app as the first application based on user input selecting a specific app.
[0153] According to one embodiment, the electronic device (101) may recommend or download a suitable application based on determining that a first application for executing a first operation process cannot be determined from among the applications stored in memory (e.g., memory (130) of FIG. 1) or that there is no application suitable for performing the first operation in memory (130). For example, the electronic device (101) may access an application distribution platform (e.g., Google Play Store) via wireless communication based on the inability to determine the first application, identify a third application capable of executing the first operation process on the application distribution platform, and recommend or download the third application. The electronic device (101) may, for example, recommend the third application and download the third application based on user input approving the download.
[0154] According to one embodiment, operation 740 can be performed using a mixer. For example, the electronic device (101) can determine the first application using a mixer. The operation of the electronic device (101) corresponding to embodiments in which the electronic device (101) uses a mixer is described in detail in the description of FIG. 13a.
[0155] In operation 750, the electronic device (101) may determine a first operation process based on the first operation and the first application. The first operation process may include instructions for performing one or more operations sequentially or in parallel. For example, if the first application is a calendar application and the first operation is “checking a schedule for a specific date,” the first operation process may consist of a series of instructions including (i) accessing calendar data included in personalized data (e.g., personalized data (600) in FIG. 6), (ii) searching for a schedule corresponding to a specific date, (iii) obtaining the searched schedule as text data, and (iv) displaying the text data to the user. The electronic device (101) may determine the first operation process using, for example, an artificial intelligence model.
[0156] According to one embodiment, the first operation process may include instructions for analyzing and interacting with a screen displayed through a display (e.g., the display module (160) of FIG. 1). For example, the first operation process may include instructions for the electronic device (101) to execute a first application when executed by a processor (e.g., the processor (120) of FIG. 1), to capture and analyze the displayed screen (e.g., analyze using an artificial intelligence model), and to identify and interact with elements (e.g., user interface) for the first operation (e.g., touch, click, scroll, text input).
[0157] According to one embodiment, operation 750 can be performed using a mixer. For example, the electronic device (101) can determine the first operation process using a mixer. The operation of the electronic device (101) corresponding to embodiments in which the electronic device (101) uses a mixer is described in detail in the description of FIG. 13a.
[0158] In operation 760, the electronic device (101) can execute a first operation process through a first application (e.g., in the processor (120) of FIG. 1). For example, if the first operation process is an operation process for checking user schedule data, the electronic device (101) can perform a process of querying schedule data through a calendar application and then providing the schedule data (e.g., outputting it through a display) by executing commands that constitute the first operation process.
[0159] According to one embodiment, the electronic device (101) may store the first operation process in memory (e.g., memory (130) of FIG. 1) after executing the first operation process. For example, the electronic device (101) may store the instructions constituting the first operation process as a predetermined set of instructions. Accordingly, after performing operations 710 to 760, the electronic device (101) may identify the first operation again based on a user request, and thus identify that the first operation can be performed using the predetermined set of instructions. Through this, the electronic device (101) does not have to repeat the computational process to determine the operation process corresponding to the first operation anew each time, thereby saving time and processor resources.
[0160] In operation 770, the electronic device (101) can perform the first operation using a set of instructions based on identifying that the first operation can be performed using a set of instructions. For example, based on the fact that the first operation is an operation that executes the first function of the first application, the electronic device (101) can call and perform the corresponding function by executing a set of instructions corresponding to the first function of the first application that is included in the set of instructions.
[0161] FIG. 7b is a drawing for explaining operation 770 according to one embodiment of the present disclosure in more detail.
[0162] Referring to FIG. 7b, according to one embodiment, operation 770 may include operation 771 and / or operation 772.
[0163] In operation 771, the electronic device (101) may determine a second application that is included in the first category. The second application may be an application for executing a set of instructions. Instead of being determined in operation 771, the second application may be predetermined.
[0164] In operation 772, the electronic device (101) can perform the first operation through the second application using a defined set of instructions. For example, the electronic device (101) can perform the first operation by executing the second application by calling a predefined API and controlling the functions of the second application.
[0165] FIG. 7c is a drawing for explaining operation 760 according to one embodiment of the present disclosure in more detail.
[0166] Referring to FIG. 7c, according to one embodiment, operation 760 may include operation 761 and / or operation 762.
[0167] In operation 761, the electronic device (101) can identify a first element for a first operation process by analyzing a screen (e.g., a screen that may be displayed through the display of the electronic device (101)) resulting from the execution of a first application. The first element may include, for example, at least one user interface element. For example, the electronic device (101) can recognize at least one element for executing a first operation process by analyzing a screen resulting from the execution of a “hotel reservation application.” For example, the electronic device (101) can recognize a user interface element containing “rating” text and a search window user interface element on the screen to execute an operation process for “searching for high-rated hotels.”
[0168] In operation 762, the electronic device (101) can execute a first operation process by interacting with the first identified element. For example, the electronic device (101) can enter a specific regional hotel for a specific date into a recognized search window user interface and interact with a search condition setting user interface containing the text “Rating” to set a search condition of a certain rating or higher (e.g., 9.0 points or higher) and perform a search operation.
[0169] According to one embodiment, the screen may be actually displayed on a physical display (e.g., the display module (160) of FIG. 1) or may be rendered virtually without being displayed. In the latter case, since the actions of analyzing the screen and interacting with elements on the screen in actions 761 and 762 are performed on a virtual display, the user can simultaneously perform a separate task on the actual display while the actions are in progress. Accordingly, the electronic device (101) can provide a multitasking environment to the user without interruption while executing the first action process in the background.
[0170] According to one embodiment, an electronic device (101) may perform operation 761 and / or operation 762 using an artificial intelligence model for image analysis and user interface interaction. The artificial intelligence model may include a model trained to interpret the meaning of visual elements (e.g., text, user interface) included on a screen. For example, the artificial intelligence model may recognize the content of text included on the screen and identify the function or meaning represented by the text to identify whether it is an element for performing a first operation process. For example, the electronic device (101) may use the artificial intelligence model to recognize a button or input window containing keywords such as "search," "reservation," or "share," and execute a first operation process through interaction with the elements.
[0171] FIGS. 8a through 8c illustrate a screen displayed by an electronic device (e.g., the electronic device (101) of FIG. 1) according to embodiments of the present disclosure.
[0172] Referring to FIG. 8a, according to one embodiment, an electronic device (101) may display a first user interface that guides actions that can be performed by category through a display (e.g., the display module (160) of FIG. 1). The actions may be, for example, the actions of FIG. 4. The first user interface may include a user interface related to the travel category (810), a user interface related to the finance category (820), a user interface related to the meal category (830), and a user interface related to the exercise category (840) in the example of FIG. 8a, and each category may correspond to a category that classifies applications.
[0173] “Application” may be referred to as “app” below. “x-related user interface” may be referred to as “x” below to the extent that the meaning is clear. For example, “travel category-related user interface (810)” may be referred to as “travel (810)”.
[0174] According to one embodiment, the electronic device (101) can display the screen of FIG. 8b based on user input through a travel category (810).
[0175] Referring to FIG. 8b, the electronic device (101) may display a second user interface. The second user interface may include a request input user interface (850), a user interface related to an airline ticket category app (860), a user interface related to an accommodation category app (870), and a user interface related to a schedule category app (880).
[0176] According to one embodiment, the request input user interface (850) may include an interactive interface for a user to input a request. For example, a user may input a request by touching the request input user interface (850) and inputting text (e.g., input via keyboard or input via voice).
[0177] According to one embodiment, the electronic device (101) can classify the apps into a plurality of categories by analyzing the characteristics of the stored apps. For example, the electronic device (101) can classify the first app and the second app into the airline ticket category, classify the third app and the fourth app into the accommodation category, and classify the fifth app into the schedule category. The user interface (860) related to the airline ticket category app may include the first app icon (861) and the second app icon (862). The user interface (870) related to the accommodation category app may include the third app icon (871) and the fourth app icon (872). The user interface (880) related to the schedule category app may include the fifth app icon (881).
[0178] According to one embodiment, the user interface (860) related to the flight ticket category app, the user interface (870) related to the accommodation category app, and the user interface (880) related to the schedule category app may each include a user interface (865, 875, and 885) for adding apps. The electronic device (101) may add other apps to each pre-classified category based on user input through the user interface (865, 875, and 885) for adding apps. For example, based on the user adding a sixth app through the user interface (885), the electronic device (101) may add the sixth app to an existing classified schedule category and add and display the sixth app icon in the user interface (880) related to the schedule category app.
[0179] According to one embodiment, the electronic device (101) can delete apps in addition to adding them. For example, the electronic device (101) may delete the sixth app from a certain category based on user input.
[0180] According to one embodiment, the electronic device (101) may include a task completion notification text (890), a first task completion notification user interface (891), a second task completion notification user interface (892), and a third task completion notification user interface (893) based on performing at least one action in response to a user request. The task completion notification text (890) may include, for example, text such as “Task completed. Check details.” The first task completion notification (891) may include information notifying the completion of a first action (e.g., flight reservation action) included in at least one action. The second task completion notification (892) may include information notifying the completion of a second action (e.g., hotel reservation action). The third task completion notification (893) may include information notifying the completion of a third action (e.g., schedule management action).
[0181] FIGS. 9a to 9c illustrate a screen displayed by an electronic device (101) according to embodiments of the present disclosure.
[0182] According to one embodiment, the electronic device (101) can interact with the user in real time and provide application-related information (e.g., recommended application information) in response to a request entered by the user in real time.
[0183] FIG. 9a may be a screen in which a first request (910) is entered into a user request input window and an electronic device (101) displays a user interface (911) related to an airline category in accordance with the first request (910) (“book the earliest afternoon flight to New York on the 12th of this month”). According to one embodiment, the user may add, delete, or modify the request entered into the user request input window.
[0184] FIG. 9b may be a screen displayed by an electronic device (101) based on the user additionally inputting a second request (920) (“Find a hotel to stay at for 10 days. Also find a decent restaurant within a 5-minute walk of the accommodation and make a reservation”) following the first request (910) on the screen of FIG. 9a. According to one embodiment, the electronic device (101) may display a user interface related to accommodation categories (922) and a user interface related to restaurant categories (921) in addition to the previously displayed airline (911) based on the user input of the additional second request (920).
[0185] FIG. 9c may be a screen displayed by an electronic device (101) based on the user clearing the first request (910) and the second request (920) from the screen of FIG. 9b and instead entering a third request (930) (“Book the earliest afternoon flight to London on the 20th of this month and get a taxi to the airport before I’m late”). According to one embodiment, the electronic device (101) may display a user interface (931) related to the flight (911) and taxi categories based on the third request (930).
[0186] FIGS. 10a and FIGS. 10b illustrate a screen displayed by an electronic device (e.g., the electronic device (101) of FIG. 1) according to embodiments of the present disclosure.
[0187] Referring to FIG. 10a, according to one embodiment, an electronic device (101) can display a first request input window (1010), a first operation completion notification user interface (1015), a second request input window (1020), and a second operation completion notification user interface (1025) through a display (e.g., the display module (160) of FIG. 1). The electronic device (101) may be, for example, a foldable device. That is, the first request input window (1010) and the first operation completion notification user interface (1015) may be displayed at the top of the folding line of the display module (160), and the second request input window (1020) and the second operation completion notification user interface (1025) may be displayed at the bottom of the folding line.
[0188] According to one embodiment, the first operation completion notification (1015) may include information indicating that the electronic device (101) has completed an operation (e.g., delivery reservation) corresponding to a user request through the first request input window (1010). The second operation completion notification (1025) may include information indicating that the electronic device (101) has completed an operation (e.g., taxi reservation) corresponding to a user request through the second request input window (1020).
[0189] According to one embodiment, FIG. 10b may be a screen in which only the lower part of the screen of FIG. 10a is changed to a different user interface.
[0190] Referring to FIG. 10b, according to one embodiment, an electronic device (101) may analyze a user request through a first request input window (1010) and display a related action recommendation user interface (1030) that recommends actions that can be performed in addition to actions that directly correspond to the user's request. The related action recommendation user interface (1030) may include a first action recommendation (1031), a second action recommendation (1032), and a third action recommendation (1033).
[0191] According to one embodiment, the electronic device (101) may perform a corresponding action based on user input selecting or approving at least some of the action recommendations. For example, the electronic device (101) may perform the actions corresponding to the first action recommendation, such as “organizing seminar lecture materials” and “sending emails to relevant parties,” based on user input approving the first action recommendation (1031).
[0192] FIGS. 11a and FIGS. 11b illustrate a screen displayed by an electronic device (e.g., the electronic device (101) of FIG. 1) according to embodiments of the present disclosure.
[0193] Referring to FIG. 11a, according to one embodiment, an electronic device (101) may display a user request input window (1110) and a notification user interface (1120) during operation execution. The notification during operation execution (1120) may include information indicating that an operation corresponding to a user request through the user request input window (1110) is currently being performed. That an operation corresponding to a user request is currently being performed may mean that at least one of the operations corresponding to the user request has started and at least some of the operations corresponding to the user request have not yet been completed. The notification during operation execution (1120) may display, for example, icons of running applications (e.g., flight reservation app, calendar app, hotel reservation app).
[0194] Referring to FIG. 11b, according to one embodiment, an electronic device (101) may display an operation completion notification user interface (1130) indicating that an operation corresponding to a user request through a user request input window (1110) has been completed. An operation corresponding to a user request has been completed may mean that all of the operations corresponding to the user request have been completed. The operation completion notification (1130) may include, for example, information on the completed operation for each application that was displayed in the operation execution notification (1120) (e.g., in the case of a hotel reservation operation, the reservation date and time and the check-in / check-out time).
[0195] FIG. 12 is a drawing for explaining the operation of an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment of the present disclosure.
[0196] Referring to FIG. 12, according to one embodiment, an electronic device (101) can transmit a received natural language user request (1210) to an LLM (1220). The LLM (1220) can analyze the natural language user request (1210), identify the user's intent, and output a prompt (1230) that is improved into a more specific form. For example, if the user request (1210) is "Help prepare for the meeting," the LLM (1220) can output a prompt (1230) such as "Check the list of attendees, summarize the meeting materials, and reserve the meeting room for the meeting scheduled for 2:00 PM today."
[0197] According to one embodiment, an electronic device (101) may transmit a prompt (1230) to an artificial intelligence agent system (400). The artificial intelligence agent system (400) may be, for example, the artificial intelligence agent system (400) of FIG. 4, or may be replaced by any artificial intelligence model (e.g., a large action model) trained to convert into device actions based on user requests. The artificial intelligence agent system (400) may obtain an application list (1245) to perform at least one action based on the prompt (1230).
[0198] According to one embodiment, the electronic device (101) can determine an operation process (1240) from a prompt (1230) using an artificial intelligence agent system (400). The operation of determining the operation process (1240) may be, for example, as shown in FIG. 7a.
[0199] FIG. 13a is a drawing for explaining the operation using a mixer (1320) of an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment of the present disclosure.
[0200] According to one embodiment, the mixer (1320) may include a text analysis module (1321), an application management module (1322), and an action template management module (1323).
[0201] According to one embodiment, the prompt (1310) may be input into the text analysis module (1321) of the mixer (1320). The prompt (1310) may be a prompt obtained by an artificial intelligence model (e.g., LLM) analyzing and improving a natural language user request. The prompt (1310) may be a prompt obtained in a more specific natural language form through LLM-based prompt engineering, for example, as described in operation 710 of FIG. 7a.
[0202] According to one embodiment, the text analysis module (1321) may be an artificial intelligence model (e.g., LLM) and may determine the user's intent by analyzing the prompt (1310). The text analysis module (1321) may be trained, for example, using data pairs consisting of a large natural language corpus and a user intent that matches it. The text analysis module (1321) may derive the user intent by analyzing the implied meaning of various natural languages. The user intent may be, for example, an intent to perform at least one action corresponding to the user request in action 710 of FIG. 7a. For example, the user intent may include an application category for performing at least a part of at least one action (e.g., a first category of action 720 of FIG. 7a).
[0203] According to one embodiment, the text analysis module (1321) may analyze only the input natural language, but may also analyze actions that are inferred to have been missed by the user based on the input natural language. For example, based on the user inputting "Is there a meeting tomorrow morning?", the text analysis module (1321) may analyze auxiliary actions such as viewing meeting materials or checking meeting room reservation status in addition to the request to check the schedule derived directly from the input natural language. The text analysis module (1321) may analyze the text and output at least one action corresponding to the prompt (1310) and / or at least one application category corresponding to the at least one action as a result.
[0204] According to one embodiment, the text analysis module (1321) can classify and output the results of analyzing the prompt (1310) (e.g., actions directly derived from the input natural language and / or actions inferred to have been missed by the user based on the input natural language) by category. For example, the text analysis module (1321) can output user intent analysis results such as Table 1 below from the natural language text “Find and book a hotel or other accommodation with a rating of 9 or higher and 4 stars or higher in New York from November 19 to 21, 2024.”
[0205] - Location: New York - Action: Book accommodation - Details: High rating (9 or higher) - Additional actions (inference): Book flight, add calendar event
[0206] For example, in Table 1, the location item, action item, and detail item may be items for actions derived directly from the natural language of the prompt (1310). For example, the additional action item may be items for actions that the user missed or is inferred to perform later based on the natural language of the prompt (1310).
[0207] According to one embodiment, the application management module (1322) can select a necessary app among the apps stored in the electronic device (101) based on the analysis result output by the text analysis module (1321) (e.g., the need to execute a specific function of a specific application). The application management module (1322) can access the category data of the application within the electronic device (101) to search for an application that matches the category provided by the text analysis module (1321). To this end, the application management module (1322) can internally identify and / or store the category of the application and / or the function(s) that the application can perform when the application is installed.
[0208] According to one embodiment, the application management module (1322) may select an application for performing a major operation among at least one operation as a top priority application. Once the selection of the top priority application is complete, the application management module (1322) may select an application corresponding to a lower priority category to provide additional functions. The application management module (1322) may transmit the selected application list information to the operation template management module (1323).
[0209] According to one embodiment, when the application management module (1322) selects an application, it can select an application corresponding to a major action as the highest priority application by setting the 'action' category as the highest priority among the categories identified by the text analysis module (1321). After the selection of the highest priority application is completed, the application management module (1322) can also provide additional functions by selecting an application corresponding to a category other than the 'action' category (e.g., schedule management category, information provision category, environment settings, etc., and other additional support categories).
[0210] According to one embodiment, the operation of the application management module (1322) may correspond to the operation 740 of FIG. 7. For example, the electronic device (101) may use the application management module (1322) to select a top priority application in a first category or a similar category.
[0211] According to one embodiment, the action template management module (1323) may receive a selected application list along with the user's natural language input and / or a prompt (1310) that improves the user's natural language input using an artificial intelligence model. The action template management module (1323) may generate an item of action in text form and output it as an action template (1330). The action template (1330) may be divided into steps of the action, and the mixer (1320) may output one or more action templates (1330).
[0212] FIG. 13b is a drawing for explaining an operation template (1330) according to one embodiment of the present disclosure.
[0213] Referring to FIG. 13b, according to one embodiment, the action template (1330) may include a plan (1331) item, a current action (1332) item and / or a next action (1333) item.
[0214] According to one embodiment, the plan (1331) may include information on the operation to be performed through the template of the current stage. The current action (1332) may mean an operation to be performed immediately in the electronic device (101), and information on the location to be found to perform the operation in the electronic device (101) may be included together with the coordinates of the display (e.g., the display module (160) of FIG. 1). The next action (1333) may include an operation to be performed after the current action (1332) and other data required to perform the operation.
[0215] In FIG. 13b, the “box size” included in the next action (1333) may refer to a search range that can be arbitrarily used within the search area to detect a place search box on the screen. Although a smaller box size allows for more accurate operation, searching the entire screen area with a small box size takes a long time, so the electronic device (101) may set an appropriate value based on user input or automatically.
[0216] According to one embodiment, the process of performing an action based on an action template (1330) may vary depending on the number of templates and / or the method of template generation generated by the action template management module (e.g., the action template management module (1323) of FIG. 13a). For example, when an application is added or deleted based on user input during the process of selecting an application to perform an action, the process of generating the template by the action template management module (1323) may be initialized and performed again from the beginning. For example, the total execution time may increase as the process is repeated or additional actions are performed due to user intervention, but the accuracy of template generation (the degree to which it matches the user's intention) may increase, thereby providing more precise work results.
[0217] FIGS. 14a through 14g illustrate screens of an electronic device (e.g., the electronic device (101) of FIG. 1) according to embodiments of the present disclosure. FIGS. 14a through 14g may be screens displayed on a display during the process of analyzing and interacting with the screen to perform an action corresponding to the action template (1330) of FIG. 13b and / or other action templates based on a user's request, for example, as in action 760 of FIG. 7c.
[0218] According to one embodiment, the electronic device (101) can run a hotel reservation application to display the screen of FIG. 14a. The operation of running the hotel reservation application may be, for example, an operation corresponding to “1. Run a hotel reservation application installed on the device” in the plan (1331) of the operation template (1330) of FIG. 13b.
[0219] According to one embodiment, the electronic device (101) can identify a location selection user interface (1411), a date selection user interface (1412), a room and number of people selection user interface (1413), and a search button (1414) by analyzing the screen of FIG. 14a based on the contents of the plan (1331). For example, the electronic device (101) can identify the location selection user interface (1411) based on “2. Detecting the location search bar of the application” in the plan (1331), and can perform a search by entering “New York” through the location selection user interface (1411) and using user input through the search button (1414) based on “3. Entering ‘New York’ into the location search bar” and “4. Touching the search button next to the location search bar.”
[0220] According to one embodiment, screen analysis can be performed, for example, using an artificial intelligence model for image analysis (e.g., LVM, large vision model). The artificial intelligence model for image analysis can acquire, for example, optical character recognition (OCR), object segmentation information, depth information, and optical flow information.
[0221] According to one embodiment, the electronic device (101) identifies information regarding the location, date, and number of people based on a user request to reserve a hotel, then interacts with a location selection user interface (1411), a date selection user interface (1412), and a room and number of people selection user interface (1413) to set appropriate search conditions, and can perform a search by interacting with a search button (1414).
[0222] According to one embodiment, the electronic device (101) may display a search results screen based on interaction through a search button (1414) on the screen of FIG. 14a. The screen of FIG. 14b may include a user interface for filtering search results according to a rating on the search results screen.
[0223] According to one embodiment, the electronic device (101) can identify the “Customer Rating” text (1421), the rating 9 or higher selection user interface (1422), and the completion button (1423) by analyzing the screen of FIG. 14b based on the contents of the plan (1331). For example, the electronic device (101) can identify that the current screen is a screen for filtering search results based on ratings by analyzing the meaning of the “Customer Rating” text (1421) on the screen based on “7. Detect Rating Order Button” in the plan (1331), and can interact with the completion button (1423) while selecting the rating 9 or higher selection user interface (1422) based on “8. Touch Rating Order Button”.
[0224] According to one embodiment, FIG. 14c and FIG. 14d may be screens displayed during the process in which an electronic device (101) performs operations included in a first operation template. The operations included in the first operation template may include, for example, a hotel information checking operation, a hotel detailed information checking operation, a hotel determination operation, and a room selection operation.
[0225] According to one embodiment, an electronic device (101) can identify a first hotel information user interface (1430), a second hotel information user interface (1440), and a third hotel information user interface (1450) in order to perform a hotel information verification operation included in a first operation template by analyzing the screen of FIG. 14c, which is a screen resulting from filtering search results. Each hotel information user interface may include information such as a hotel name, rating, and cost. For example, the first hotel information user interface (1430) may include the name of the first hotel, “The Bowery Hotel,” along with first hotel rating information (1431) and first hotel cost information (1432).
[0226] According to one embodiment, an electronic device (101) may display the screen of FIG. 14d based on interaction through a first hotel information user interface (1431) to perform a hotel detail information verification operation included in a first operation template, and may analyze the screen to recognize rating information (1461), amenities and services (1462), cost (1463), and / or room selection buttons (1464). To perform a hotel determination and room selection operation, the electronic device (101) may recognize the text of rating information (1461), amenities and services (1462), and cost (1463), and reconfirm the information to determine whether an appropriate hotel has been selected. For example, if it is identified that the user is looking for a hotel with a rating of 9 points or higher, pet-friendly, and a cost of 1.5 million won or less per night based on user requests and / or user personalization data, the electronic device (101) may determine that “The Bowery Hotel” is a suitable hotel based on the text of rating information (1461) (9.4 points), amenities and services (1462) (pet-friendly), and cost (1463) (1,004,463 won / night). If it is determined that it is a suitable hotel, the electronic device (101) may display the screen of FIG. 14e based on interaction through the room selection button (1464).
[0227] According to one embodiment, FIGS. 14e, FIGS. 14f, and FIGS. 14g may be screens displayed during the process in which an electronic device (101) performs operations included in a second operation template. The operations included in the second operation template may include, for example, a payment option selection operation, a payment information verification operation, a detail reconfirmation operation, a name input operation, and a payment information input operation.
[0228] According to one embodiment, an electronic device (101) can recognize a payment information-related user interface (1470) by analyzing the screen of FIG. 14e to perform a payment option selection operation and / or a payment information verification operation included in a second operation template. The electronic device (101) can perform a payment information verification operation by analyzing the payment information-related user interface (1470) and recognizing the “Payment Option Selection” text (1471), other payment information (1472), cost information (1473), and payment button (1474). The electronic device (101) can perform a payment option selection operation by identifying from the “Payment Option Selection” text (1471) that the current screen is a screen in the pre-payment stage displaying payment information, and by recognizing information such as discount coupon information or payment method from the other payment information (1472).
[0229] According to one embodiment, the electronic device (101) may display a payment progress screen based on interaction via the payment button (1474) on the screen of FIG. 14e. The payment progress screen may include various information, which may all be displayed at a single point in time, but may not all be displayed at a single point in time and may require moving the screen by scrolling to view all the various information. The screens of FIG. 14f and FIG. 14g may be screens that display various information related to the payment progress.
[0230] According to one embodiment, the electronic device (101) can recognize rating information (1461) and check-in / check-out information (1482) by analyzing the screen of FIG. 14f to perform a detail reconfirmation operation included in the second operation template. The electronic device (101) can prevent the booking of accommodation on an incorrect date or with an incorrect standard (a rating lower than the standard) by reconfirming the rating information (1461) and check-in / check-out information (1482).
[0231] According to one embodiment, the electronic device (101) can recognize the reservation name information input window (1491) and the payment card information input window (1492) by analyzing the screen of FIG. 14g to perform a name input operation and a payment information input operation included in the second operation template. The electronic device (101) can access personalized data (e.g., personalized data (600) of FIG. 6) and automatically input the user's name information and payment card information without manual input by the user to proceed with payment.
[0232] According to one embodiment, if there are multiple operation templates, the electronic device (101) may receive all operation templates sequentially or simultaneously and perform the above-described process sequentially or simultaneously for each operation template until all operations included in each operation template are completed. For example, after the electronic device (101) has performed all operations included in the operation template (1330) of FIG. 13b, it may perform all operations included in another operation template using screen analysis and interaction.
[0233] FIG. 15 is a diagram illustrating the operation of an electronic device (e.g., the electronic device (101) of FIG. 1) using a screen analysis and interaction system (1530) according to one embodiment of the present disclosure. The screen analysis and interaction system (1530) may be configured to determine an action(s) and an application to perform the action(s) in response to a user request, and then acquire an action process (or action template) and perform an action to execute the action process (or action template). For example, the screen analysis and interaction system (1530) may be configured to perform action 750 and / or action 760 of FIG. 7a. The screen analysis and interaction system (1530) may be configured to be included in an artificial intelligence agent (e.g., a first artificial intelligence agent (431)) corresponding to a specific application included in a subordinate artificial intelligence agent (e.g., a subordinate artificial intelligence agent (430) of FIG. 4). The screen analysis and interaction system (1530) may include, for example, at least a part of the configuration of the artificial intelligence agent (500) of FIG. 5 (e.g., execution module (540)).
[0234] Referring to FIG. 15, according to one embodiment, a screen analysis and interaction system (1530) may include a screen analysis module (1531), a motion matching module (1533), and a screen control module (1534).
[0235] According to one embodiment, since the screen displayed on the display of the electronic device (101) may change each time the electronic device (101) performs an operation (e.g., an operation included in an operation template), the electronic device (101) may capture the screen and transmit the screen capture image (1510) to the screen analysis module (1531). The screen analysis module (1531) may be, for example, an artificial intelligence model for image analysis (e.g., LVM), and may acquire OCR, object segmentation information, depth information, and light flow information because it recognizes elements of the screen based on computer vision, for example. The screen analysis module (1531) may analyze the screen capture image (1510) to acquire screen analysis data (1532) and transmit it to the operation matching module (1533).
[0236] According to one embodiment, an electronic device (101) operation template (1520) can be transmitted to an operation matching module (1533). The operation template (1520) may be, for example, the operation template (1330) of FIG. 13b.
[0237] According to one embodiment, the action matching module (1533) can identify whether there is text on the screen capture image (1510) that matches an action in the action template (1520) based on OCR information. At this time, even if the text does not match exactly as it is, the action matching module (1533) can identify it as matching text if it has a similar meaning by correcting it based on the meaning of the text. Based on the fact that the action template (1520) contains the text "Submit," the action matching module (1533) can identify a button on the screen capture image (1510) that contains semantically similar text, such as "Complete," "Upload," or "Confirm," based on screen analysis data (1532), and determine that the button corresponds to the "Submit" action.
[0238] According to one embodiment, the action matching module (1533) may transmit the matching result to the screen control module (1534). The matching result may include the locations of elements on the screen (e.g., user interfaces) that are matched with actions on the action template (1520).
[0239] According to one embodiment, the screen control module (1534) can execute an operation process (1540) by interacting with elements on the screen to perform at least one operation. For example, the screen control module (1534) can execute each step command of the operation process (1540) to perform operations such as clicking a user interface button at a specific coordinate on the screen or entering specific text into an input window.
[0240] According to one embodiment, the output result of the screen analysis and interaction system (1530) may be provided, for example, in the form of a user interface (e.g., task completion notification text (890), operation completion notification user interface (1130)) as in FIG. 8c or FIG. 11b, by an electronic device (101) using an artificial intelligence agent (e.g., artificial intelligence agent (500) of FIG. 5). The user interface provided by the artificial intelligence agent may be interactive with the user and may provide details of the operation performed.
[0241] According to one embodiment, the operation of the electronic device (101) described in FIGS. 14a to 14g may be an operation using a screen analysis and interaction system (1530). For example, in FIG. 14a, the electronic device (101) may execute an operation process (1540) by (i) inputting the screen capture image (1510) of FIG. 14a into a screen analysis module (1531) to obtain screen analysis data (1532), (ii) inputting the screen analysis data (1532) and an operation template (1520) into an operation matching module (1533) to match the positions of elements on the screen (e.g., “detecting a place search bar,” “entering ‘New York’ in the place search bar,” “touching a search button”) included in the operation template (1520), and (iii) interacting with the elements on the screen using a screen control module (1534).
[0242] FIG. 16 is a drawing for explaining the operation of an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment of the present disclosure.
[0243] Referring to FIG. 16, according to one embodiment, an electronic device (101) can input a natural language-based user request (1610) (e.g., a user request entered in the form of voice or text) into an LLM (1220) as in FIG. 12. The electronic device (101) can convert the user request (1610) entered in voice into a text type using a speech-to-text artificial intelligence model. The LLM (1220) can analyze the user request (1610) that was entered in text type or converted into text type after being entered in voice and generate a more specific prompt (e.g., the prompt (1230) of FIG. 12) (e.g., generated through prompt engineering as described in operation 710 of FIG. 7) and input it into an artificial intelligence agent system (400) as in FIG. 4.
[0244] For example, based on a user request such as “If the weather is nice tomorrow, text my friends to play soccer,” LLM (1220) can generate a prompt by changing the “If the weather is nice” part to “weather suitable for playing outdoor sports (soccer) with humidity, temperature, and no rain” and supplementing the “text my friends” part to “send a message via a messenger application or SNS application to friends I have recently scheduled a new event with or with whom I have played soccer regularly.”
[0245] According to one embodiment, the artificial intelligence agent system (400) receives a prompt from the LLM (1220), which is first received by the upper artificial intelligence agent (420), and the upper artificial intelligence agent (420) can control at least one application by issuing a command to each lower artificial intelligence agent (e.g., first artificial intelligence agent (431), second artificial intelligence agent (432)) capable of controlling applications of the required category.
[0246] According to one embodiment, the artificial intelligence agent system (400) can exchange data with the screen analysis and interaction system (1530) of FIG. 15. For example, when each subordinate artificial intelligence agent generates a series of action templates and transmits them to the screen analysis and interaction system (1530), the screen analysis module (1531) can capture and analyze the screen and match the screen analysis data (1532) with the action process received from each subordinate artificial intelligence agent in the action matching module (1533), and accordingly, the screen can be controlled by the screen control module (1534).
[0247] According to one embodiment, the electronic device (100) can execute a main operation application among a plurality of applications, and then capture the execution screen of the application and input it into a screen analysis module (1531). The screen analysis module (1531) can identify elements (e.g., user interface) included in the input screen image and obtain information regarding the content, type, and location on the screen of the identified elements. For example, the screen analysis module (1531) can obtain information about elements included in the screen, such as a 'reservation button (x-coordinate: 360, y-coordinate: 127)' and a 'date selection button (x-coordinate: 130~135, y-coordinate: 155~164)', based on the execution screen of an airline ticket reservation application.
[0248] According to one embodiment, there may be cases where correction is required for the data output by the screen analysis module (1531). For example, the screen analysis module (1531) is trained to identify the function of a specific button based on the word 'reservation', but if a function with the same meaning is displayed as a similar expression such as 'pay' or 'confirm', it may not be able to accurately classify the function of the similar expression button.
[0249] According to one embodiment, the screen analysis module (1531) may refer to and / or learn a predefined data correspondence table, such as homonyms and synonyms. Through this, the screen analysis module (1531) can reduce errors by identifying and / or interpreting elements of the same or similar functions by checking whether information (e.g., text information) of elements on the screen matches information included in the data correspondence table.
[0250] According to one embodiment, the lower artificial intelligence agent (431 or 432) may transmit additional information generated or required during the application operation performed by each to the upper artificial intelligence agent (420) in real time. The additional information may include, for example, the execution status of the application, data regarding the task currently being processed (e.g., metadata), and / or data related to external input (e.g., response results to the input).
[0251] According to one embodiment, the upper artificial intelligence agent (420) can analyze information received from the lower artificial intelligence agent (431 or 432) to adjust the operation schedule (e.g., order of operation among multiple lower artificial intelligence agents), priority, accuracy of operation, execution conditions and / or resource allocation information of the lower artificial intelligence agent (431 or 432).
[0252] According to one embodiment, based on the existence of data that the lower AI agent (431 or 432) cannot access among the additional information requested by the lower AI agent (431 or 432), the upper AI agent (420) may request the data on its behalf or receive it from an external system and transmit it to the lower AI agent (431 or 432). The data may be data requiring security authorization, such as user authentication information, for example.
[0253] According to one embodiment, after performing an operation using an artificial intelligence agent system (400) and a screen analysis and interaction system (1530), the electronic device (101) can transmit the result to a UI (user interface) generation module (1620).
[0254] According to one embodiment, the electronic device (101) can generate an operation result (1630), which is a user interface including visual elements that indicate the operation result, using a UI generation module (1620). For example, the electronic device (101) can generate an operation result (1630), which is a user interface for notifying the upper artificial intelligence agent (320) that the work of the lower artificial intelligence agents (431 or 432) has been completed and for checking the operation result of each lower artificial intelligence agent (431 or 432).
[0255] According to one embodiment, a predefined template may be used to generate the result of the operation (1630); however, since the electronic device (101) according to various embodiments of the present disclosure performs operations by utilizing various applications of various combinations of application categories, it may always be difficult to predefine an appropriate template. Accordingly, the UI generation module (1620) includes a generative artificial intelligence model, and said generative artificial intelligence model can generate a user interface.
[0256] According to one embodiment, an artificial intelligence model included in a UI generation module (1620) can generate an action execution result (1630) and display the action execution result (1630) through a display of an electronic device (101) (e.g., a display module (160) of FIG. 1). For example, each lower artificial intelligence agent (431 or 432) can transmit the action execution result (1630) generated by the UI generation module (1620) to a higher artificial intelligence agent (420), and the higher artificial intelligence agent (420) can display the action execution result (1630).
[0257] FIG. 17 illustrates the operation of a processor (120) and a display module (160) of an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment of the present disclosure over time.
[0258] Referring to FIG. 17, in operation 1710, the electronic device (101) can display a user request input window through the display module (160). The user can input a user request through the user request input window.
[0259] In operation 1720, the electronic device (101) can perform prompt engineering on the processor (120). For example, the processor (120) can use the LLM (1220) to analyze the user request and modify it into a more specific and clear prompt.
[0260] In operation 1730, the electronic device (101) can perform an operation using a higher-level artificial intelligence agent (e.g., the higher-level artificial intelligence agent (420) of FIG. 4) in the processor (120). The higher-level artificial intelligence agent (420) can analyze the input prompt and divide the tasks to be performed by each lower-level artificial intelligence agent (e.g., the first artificial intelligence agent (431) and the second artificial intelligence agent (432) of FIG. 4). For example, as described in FIG. 4, the higher-level artificial intelligence agent (420) can receive a command proposal that corrects the user's natural language request (e.g., corrects it more specifically using LLM), determine at least one corresponding operation, and then divide and assign the command to each lower-level artificial intelligence agent (e.g., the first artificial intelligence agent (431) and the second artificial intelligence agent (432)) suitable for performing each of the at least one operation. At this time, the higher-level artificial intelligence agent (420) can provide data within the electronic device (101) to which access by the lower-level artificial intelligence agent (430) is restricted as needed.
[0261] In operation 1740, the electronic device (101) can perform an operation using a lower artificial intelligence agent (e.g., the lower artificial intelligence agent (430) of FIG. 4) in the processor (120). The lower artificial intelligence agent (430) may include one or more artificial intelligence agents (e.g., the first artificial intelligence agent (431), the second artificial intelligence agent (432)) and can perform an operation assigned by the upper artificial intelligence agent (420) using a defined application(s).
[0262] According to one embodiment, each sub-artificial intelligence agent (431, 432) may correspond to a detailed category for classifying applications and may perform operations independently according to the type or data type of the applications. For example, each sub-artificial intelligence agent (431, 432) may be an agent for operations through applications of a corresponding category. For example, each sub-artificial intelligence agent (431, 432) may determine one of the applications belonging to a specific category based on a command divided and assigned by the upper artificial intelligence agent (420) and perform at least one operation through the determined application.
[0263] According to one embodiment, a lower AI agent (430) can transmit acquired data to a higher AI agent (420). The higher AI agent (420) can analyze the received data (e.g., classify data types) and transmit the data along with work details to each of the lower AI agents (431 and 432).
[0264] According to one embodiment, while the lower AI agents (431, 432) are performing an assigned action (or task), the upper AI agent (420) can obtain the progress of each of the lower AI agents (431, 432)'s tasks and provide feedback thereon.
[0265] According to one embodiment, the upper AI agent (420) can check whether the AI agents (431, 432) have satisfied the task completion conditions as previously learned (or as specified by the user). Based on the fact that the AI agents (431, 432) have not completed the task, the upper AI agent (420) can assign the task to some or all of the AI agents (431, 432) to complete. When the AI agents (431, 432) complete the task, the upper AI agent (420) updates the information based on the output data of the AI agents (431, 432) and can repeat this action until the entire task is completed. If the upper AI agent (420) determines that the AI agents (431, 432) have not satisfied the task completion conditions, it can establish a new plan (e.g., acquiring other information about the object and image clustering using the acquired other information) and have the AI agents (431, 432) re-execute the task.
[0266] In operation 1750, the electronic device (101) may display a user interface indicating that the operation is in progress through a display module (160). Operation 1750 may be performed while operation 1730 and / or operation 1740 is being performed. Through this, the user can monitor that the requested operation is being performed.
[0267] In operation 1760, the electronic device (101) can generate a user interface in the processor (120) after the operation of the upper AI agent and the lower AI agent is completed. In operation 1770, the electronic device (101) can display the generated operation result notification user interface through the display module (160). The operation result notification user interface may include information regarding a list of operations performed, a list of applications used, and / or details of the operations. The operation result notification user interface may include information regarding the operation results of each lower agent (431, 432) and the overall operation results. The user interface of operation 1760 and / or operation 1770 may be, for example, a user interface generated by the UI generation module (1620) of FIG. 16.
[0268] FIG. 18 is a flowchart illustrating the operation of a mixer (e.g., the mixer (1320) of FIG. 13a) of an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment of the present disclosure. The mixer may include a text analysis module (1321), an application management module (1322), and an operation template management module (1323).
[0269] Referring to FIG. 18, in operation 1810, the electronic device (101) can obtain a prompt based on a natural language user request. The prompt can be obtained, for example, using an LLM.
[0270] In operation 1820, the electronic device (101) can identify a first operation using a text analysis module (1321). For example, the text analysis module (1321) can analyze the semantic structure of a sentence included in a prompt and identify an operation corresponding to the request.
[0271] In operation 1830, the electronic device (101) can identify at least one application using the application management module (1322) based on the prompt. The at least one application may be, for example, the highest priority application among applications included in a category suitable for performing the first operation.
[0272] In operation 1840, the electronic device (101) may obtain one or more operation templates using the operation template management module (1323) based on a prompt, a first operation, and at least one application. The operation templates may be in the form of, for example, the operation template (1330) of FIG. 13b.
[0273] In operation 1850, the electronic device (101) may perform a first operation based on an operation template (e.g., the operation template (1330) of FIG. 13b). For example, the electronic device (101) may search for necessary elements on the screen according to the contents of the operation template and interact with specific elements. The operation template may include, for example, operations to be performed sequentially or in parallel and information necessary to perform said operations.
[0274] In operation 1860, the electronic device (101) may notify the result of the completion of the first operation based on the completion of the first operation. For example, the electronic device (101) may display a user interface containing information that the first operation has been completed through a display (e.g., the display module (160) of FIG. 1).
[0275] FIG. 19 illustrates the operation of an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment of the present disclosure using a processor (120) and a display module (160) over time. FIG. 19 may illustrate the operation flow of an electronic device (101) using a mixer (e.g., the mixer (1320) of FIG. 13a).
[0276] Referring to FIG. 19, in operation 1910, the electronic device (101) may request a natural language query from the processor (120), and in operation 1920, the electronic device (101) may display the query through the display module (160). The query may be displayed, for example, as a user interface in the form of an input window.
[0277] In operation 1930, the electronic device (101) receives a user request based on a query from the processor (120) and can generate a prompt based on the user request. For example, the processor (120) can obtain a prompt that corrects the user request into more specific natural language using LLM based on the natural language user request.
[0278] In operation 1940, the electronic device (101) can perform an operation using an application management module (e.g., the application management module (1322) of FIG. 13a) in the processor (120). For example, the processor (120) can use the application management module (1322) to select at least one application to perform an action intended by the user based on a prompt.
[0279] In operation 1960, the electronic device (101) can perform an operation using an operation template management module (e.g., the operation template management module (1323) of FIG. 13a) in the processor (120). For example, the processor (120) can generate one or more operation templates, such as the operation template (1330) of FIG. 13b, using the operation template management module (1323).
[0280] In operation 1950, the electronic device (101) may display selected application(s) through the display module (160) while operation 1940 and / or operation 1960 is in progress. For example, the electronic device (101) may display applications of at least one category as shown in FIGS. 9a through 9c. The electronic device (101) may add or exclude applications in each category based on user input (e.g., touch input) associated with the displayed screen.
[0281] In operation 1970, the electronic device (101) can perform at least one operation based on an operation template (1330) output from the operation template management module (1323) in the processor (120). For example, the electronic device (101) can search for necessary elements on the screen according to the contents of the operation template (1330) and interact with specific elements.
[0282] In operation 1980, the electronic device (101) can generate a user interface to notify the processor (120) of the operation result after performing all the operations of the operation template (1330). For example, the electronic device (101) can generate the operation result (1630) using the UI generation module (1620) of FIG. 16.
[0283] In operation 1990, the electronic device (101) can display an operation result notification user interface through a display module (160). The operation result notification user interface may be, for example, the operation completion notification user interface (1130) of FIG. 11b.
[0284] According to one embodiment, the electronic device may include at least one processor comprising a processing circuit; and a memory comprising at least one storage medium for storing instructions. When the instructions are executed individually or collectively by the at least one processor, the electronic device may: identify at least one operation corresponding to a first user request; determine a first category based on a first operation included in the at least one operation, the first category being included in at least one category for classifying applications; identify whether the first operation is executable using a predetermined set of instructions; and identify whether the first operation is not executable using the predetermined set of instructions, thereby causing the electronic device to: determine a first application among at least one application included in the first category or a category similar to the first category, based on at least one of the first operation, user input, or personalized data; determine a first operation process corresponding to the first operation based on the first operation and the first application; and the first operation process includes instructions for performing one or more operations sequentially or in parallel, and cause the first operation process to be executed through the first application.
[0285] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may: determine a second application among at least one application included in the first category based on at least one of the first operation, the user input, or the personalized data, and cause the first operation to be performed through the second application using the determined instruction set, based on identifying that the first operation can be performed using the determined instruction set.
[0286] According to one embodiment, when instructions are executed individually or collectively by at least one processor, the electronic device may: execute the first operation process through the first application and then store the first operation process in the memory, and based on the fact that at least one operation corresponding to a second user request includes the first operation, identify that the first operation can be performed using the predetermined set of instructions and cause the first operation process to be executed through the first application using the stored first operation process.
[0287] According to one embodiment, instructions, when executed individually or collectively by at least one processor, may cause an electronic device to: analyze a screen for executing the first application based on identifying that the first operation cannot be performed using the instruction set; identify at least one first element for the first operation process; and, by interacting with the at least one first element, cause the first operation process to be executed. According to one embodiment, the screen may be displayed on a physical display or may be virtually rendered without being displayed.
[0288] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: identify the first element using an artificial intelligence model, wherein the artificial intelligence model is a model trained to interpret the meaning of a visual element included on a screen, and the visual element may include at least one of text or UI.
[0289] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: identify the first category using an upper agent, transmit the first category information to a lower agent, determine the first application or the second application using the lower agent, and perform the first operation through the first application or execute the first operation process through the second application. The upper agent and the lower agent may include a software program that automatically performs one or more operations based on user input.
[0290] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: transmit the progress of the first operation or the first operation process to the upper agent using the lower agent, and transmit to the lower agent data among the personalized data necessary for the progress of the first operation or the first operation process using the upper agent.
[0291] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may, based on identifying that the first operation cannot be performed using the instruction set: if the first application for executing the first operation process cannot be determined from among the applications stored in the memory, access an application distribution platform through the communication circuit, identify at least one third application downloadable from the application distribution platform capable of executing the first operation process, provide information about the at least one third application, or cause the at least one third application to be downloaded.
[0292] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: determine the priority of at least one application among the at least one application included in the first category based on at least one of the first operation, the user input, or the personalized data, wherein the personalized data includes the execution frequency of the applications, and determine the application with the highest priority among the plurality of applications as the first application.
[0293] According to one embodiment, when the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: generate a prompt for inputting into an artificial intelligence model using a first artificial intelligence model based on the user request entered in natural language and the personalized data, and identify the first operation using a second artificial intelligence model based on the prompt. The second artificial intelligence model may include an artificial intelligence model for natural language analysis.
[0294] According to one embodiment, the second artificial intelligence model may include a text analysis module, an application management module, and an action template management module. When the instructions are executed individually or collectively by at least one processor, the electronic device may be caused to: identify the first action using the text analysis module based on the prompt, identify at least one application for the first action using the application management module based on the prompt, obtain one or more action templates using the action template management module based on the prompt, the first action, and the at least one application, wherein the action template includes actions to be performed sequentially or in parallel and information necessary to perform said actions, and cause the first action to be performed based on said action template.
[0295] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to: identify the first operation based on analyzing the meaning of a sentence included in the prompt using the text analysis module; identify the highest priority application among applications included in a category suitable for performing the first operation using the application management module; obtain one or more operation templates using the operation template management module based on the prompt, the first operation, and the highest priority application; and perform the first operation by searching for necessary elements on the screen and interacting with specific elements based on the one or more operation templates.
[0296] According to one embodiment, a method of an electronic device may include: an action of identifying at least one action corresponding to a first user request; an action of determining a first category based on a first action included in the at least one action, wherein the first category is included in at least one category for classifying applications; an action of identifying whether the first action can be performed using a predetermined set of instructions; an action of determining a first application based on at least one of the first action, user input, or personalized data among at least one application included in the first category or a category similar to the first category, based on identifying that the first action cannot be performed using the predetermined set of instructions; an action of determining a first action process corresponding to the first action based on the first action and the first application, wherein the first action process includes instructions for performing one or more actions sequentially or in parallel; and an action of executing the first action process through the first application.
[0297] According to one embodiment, a method of an electronic device may further include: an action of determining a second application among at least one application included in the first category based on at least one of the first operation, the user input, or the personalized data; and an action of performing the first operation through the second application using the determined set of instructions, based on identifying that the first operation can be performed using the determined set of instructions.
[0298] According to one embodiment, a method of an electronic device may further include: an action of storing the first action process in the memory after executing the first action process through the first application; an action of identifying that the first action can be performed using the predetermined set of instructions based on at least one action corresponding to a second user request including the first action; and an action of executing the first action process through the first application using the stored first action process.
[0299] According to one embodiment, the operation of executing the first operation process through the first application may include: an operation of identifying at least one first element for the first operation process by analyzing a screen resulting from the execution of the first application; and an operation of executing the first operation process by interacting with the at least one first element.
[0300] According to one embodiment, the operation of identifying at least one first element may include the operation of identifying the first element using an artificial intelligence model. The artificial intelligence model is a model trained to interpret the meaning of a visual element included on a screen, and the visual element may include at least one of text or UI.
[0301] According to one embodiment, a method of an electronic device may further include: an operation of identifying the first category using an upper agent; an operation of transmitting the first category information to a lower agent; an operation of determining the first application or the second application using the lower agent; and an operation of performing the first operation through the first application or executing the first operation process through the second application. The upper agent and the lower agent may include a software program that automatically performs one or more operations based on user input.
[0302] According to one embodiment, a non-transient computer-readable recording medium storing instructions that cause at least one processor to perform set operations when executed by at least one processor, wherein the operations may include: an operation of identifying at least one operation corresponding to a first user request; an operation of determining a first category based on a first operation included in the at least one operation, wherein the first category is included in at least one category for classifying applications; an operation of identifying whether the first operation can be performed using a set of instructions; an operation of determining a first application based on at least one of the first operation, user input, or personalized data among at least one application included in the first category or a category similar to the first category, based on identifying that the first operation cannot be performed using the set of instructions; an operation of determining a first operation process corresponding to the first operation based on the first operation and the first application, wherein the first operation process includes instructions for performing one or more operations sequentially or in parallel; and an operation of executing the first operation process through the first application.
[0303] Meanwhile, the various embodiments described above may be implemented as software containing instructions stored on a device-readable storage medium, included in a computer program product in the form of a device-readable storage medium (e.g., flash memory, SSD, HDD, optical disc, magnetic tape, etc.), or distributed online through an application store, website, or cloud service. Additionally, they may be implemented within a recording medium readable by a computer or similar device using software, hardware, firmware, or a combination thereof.
[0304] Each component according to the various embodiments described above may be composed of a single or multiple entities, and some auxiliary components may be omitted or additionally included. Some components may be integrated into a single entity to perform the same or similar functions as those performed by each corresponding component prior to integration.
[0305] The operations according to the various embodiments described above may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
Claims
1. In an electronic device, At least one processor including a processing circuit; and Memory including at least one storage medium for storing instructions; and When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Identify at least one action corresponding to a first user request; A first category is determined based on a first operation included in at least one of the above operations, and the first category is included in at least one category for classifying applications; Identifying whether the above first operation can be performed using a predetermined set of instructions; Based on identifying that the above first operation cannot be performed using the above-determined instruction set: Among at least one application included in the first category or a category similar to the first category, a first application is determined based on at least one of the first operation, user input, or personalized data; Based on the first operation and the first application, a first operation process corresponding to the first operation is determined, and the first operation process includes instructions for performing one or more operations sequentially or in parallel; Causing the execution of the first operation process through the first application, Electronic device.
2. In Paragraph 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, Based on identifying that the above first operation can be performed using the above-determined instruction set: Among at least one application included in the first category above, a second application is determined based on at least one of the first operation, the user input, or the personalized data; Causing the first operation to be performed through the second application using the above-determined set of instructions, Electronic device.
3. In Paragraph 1 or 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: After executing the first operation process through the first application, the first operation process is stored in the memory; Based on the fact that at least one action corresponding to a second user request includes the first action, causing the first action process to be executed through the first application using the stored first action process, Electronic device.
4. In any one of paragraphs 1 through 3, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, Based on identifying that the above first operation cannot be performed using the above instruction set: Analyzing the screen resulting from the execution of the first application to identify at least one first element for the first operation process; By interacting with the above at least one first element, causing the first operation process to be executed, Electronic device.
5. In Paragraph 4, The above screen is displayed on a physical display or virtually rendered, Electronic device.
6. In Paragraph 4 or 5, When the above instructions are executed individually or collectively by the at least one processor, they cause the electronic device to identify the first element using an artificial intelligence model, and The above artificial intelligence model includes a model trained to interpret the meaning of visual elements included on a screen, and The above visual element includes at least one of text or a user interface, Electronic device.
7. In Paragraph 1 or 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Identify the above first category using a parent agent; Transmit the above first category information to the sub-agent; Determining the first application or the second application using the above sub-agent; Causing to perform the first operation through the first application or to execute the first operation process through the second application, The above-mentioned upper agent and the above-mentioned lower agent include a software program that automatically performs one or more actions based on user input. Electronic device.
8. In Paragraph 7, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Using the above-mentioned sub-agent, the progress of the above-mentioned first operation or the above-mentioned first operation process is transmitted to the above-mentioned upper-mentioned agent; Causing the transmission of data necessary for the progress of the first action or the first action process among the personalization data to the lower agent using the upper agent, Electronic device.
9. In any one of paragraphs 1 through 8, Including a communication circuit; further When the above instructions are executed individually or collectively by the at least one processor, the electronic device identifies that the first operation cannot be performed using the instruction set: Based on the inability to determine a first application for executing the above first operation process from among the applications stored in the memory, access the application distribution platform through the communication circuit; Identifying at least one third application downloadable from the application distribution platform capable of executing the above first operation process; Providing information about the at least one third application or causing the at least one third application to be downloaded, Electronic device.
10. In any one of paragraphs 1 through 9, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Among at least one application included in the first category, the priority of the applications is determined based on at least one of the first operation, the user input, or the personalized data, and the personalized data includes the execution frequency of the applications; Causing the application with the highest priority among the plurality of applications above to be determined as the first application, Electronic device.
11. In any one of paragraphs 1 through 10, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on the first user request entered in natural language and the personalized data, a prompt for inputting into the artificial intelligence model is generated using the first artificial intelligence model; Based on the above prompt, causing to identify the first operation using a second artificial intelligence model, wherein the second artificial intelligence model includes an artificial intelligence model for natural language analysis, Electronic device.
12. In Paragraph 11, The above-mentioned second artificial intelligence model includes a text analysis module, an application management module, and an action template management module, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on the above prompt, the first operation is identified using the text analysis module; Based on the above prompt, at least one application for the first operation is identified using the application management module; Based on the above prompt, the above first operation, and the above at least one application, one or more operation templates are obtained using the above operation template management module, and the one or more operation templates include operations to be performed sequentially or in parallel and information necessary to perform said operations; Causing to perform the first operation based on the above one or more operation templates, Electronic device.
13. In Paragraph 11, The above-mentioned second artificial intelligence model includes a text analysis module, an application management module, and an action template management module, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Identifying the first operation based on analyzing the meaning of the sentence included in the prompt using the text analysis module above; Using the above application management module, identify the application with the highest priority among the applications included in the category suitable for performing the first operation; Based on the above prompt, the above first operation, and the above highest priority application, one or more operation templates are obtained using the above operation template management module, and the one or more operation templates include operations to be performed sequentially or in parallel and information necessary to perform said operations; Causing to perform the first action by searching for necessary elements on the screen based on the one or more action templates above and interacting with specific elements, Electronic device.
14. In a method of an electronic device, An action that identifies at least one action corresponding to a first user request; Determining a first category based on a first operation included in at least one of the above operations, wherein the first category is an operation included in at least one category for classifying applications; An operation to identify whether the above first operation can be performed using a predetermined set of instructions; Based on identifying that the above first operation cannot be performed using the above-determined instruction set: An operation to determine a first application among at least one application included in the first category or a category similar to the first category, based on at least one of the first operation, user input, or personalized data; Based on the first operation and the first application, a first operation process corresponding to the first operation is determined, and the first operation process includes instructions for performing one or more operations sequentially or in parallel; and An operation to execute the first operation process through the first application; comprising method.
15. In Paragraph 14, Based on identifying that the above first operation can be performed using the above-determined instruction set: Among at least one application included in the first category above, an operation of determining a second application based on at least one of the first operation, the user input, or the personalized data; and The operation of performing the first operation through the second application using the above-determined set of instructions; further comprising method.