Method for analyzing user utterance, electronic device supporting same, and storage medium

The electronic device processes user queries through multiple models to determine similarity and generate appropriate responses, addressing the challenge of accurate user intent identification in voice recognition systems, enhancing service quality.

WO2025244436A1PCT designated stage Publication Date: 2025-11-27SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/006964
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-11
Filing Date
2025-05-22
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Existing voice recognition systems struggle to accurately identify user intent from user voice and provide appropriate content services, necessitating improved technology for high-quality voice recognition services.

Method used

An electronic device equipped with multiple models processes user queries, determines similarity between output information from these models, and provides answers based on the identified similarity, utilizing a client module, intelligent server, and service server to generate and execute plans for user inputs.

Benefits of technology

Enhances the accuracy of voice recognition by processing user queries through multiple models, ensuring appropriate responses are generated and delivered, thereby improving the quality of voice recognition services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025006964_27112025_PF_FP_ABST
    Figure KR2025006964_27112025_PF_FP_ABST
Patent Text Reader

Abstract

According to one embodiment, an electronic device comprises a memory for storing instructions, and a processor, wherein the instructions, when executed by the processor, may cause the electronic device to: acquire a first user query; input, to each of a plurality of models, input information corresponding to the first user query; provide a first answer corresponding to the first user query on the basis of first output information output from at least one first model among the plurality of models; acquire second output information output from at least one second model among the plurality of models; identify a similarity between the first output information and the second output information; and provide, on the basis of the identified similarity, a second answer acquired on the basis of the first output information and the second output information.
Need to check novelty before this filing date? Find Prior Art

Description

Method for analyzing user speech, electronic device supporting same, and storage medium

[0001] Embodiments of the present disclosure relate to a method for analyzing user speech, an electronic device supporting the same, and a storage medium.

[0002] For many people living in the modern world, portable digital communication devices have become an essential part of their lives. Consumers want to access a variety of high-quality services anytime, anywhere, using these devices.

[0003] A voice recognition service may be a service that provides consumers with various content services in response to received user voices, based on a voice recognition interface implemented in portable digital communication devices. To provide voice recognition services, portable digital communication devices may be equipped with technologies for recognizing and analyzing human language (e.g., automatic speech recognition, natural language understanding, natural language generation, machine translation, conversational systems, question-and-answering, or voice recognition / synthesis).

[0004] In order to provide high-quality voice recognition services to consumers, implementation of technology that accurately identifies user intent from user voice and implementation of technology that provides appropriate content services corresponding to identified user intent may be required.

[0005] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.

[0006] According to one embodiment of the present disclosure, an electronic device may include a memory storing instructions and a processor. The instructions, when executed by the processor, may cause the electronic device to obtain a first user query. The instructions, when executed by the processor, may cause the electronic device to input input information corresponding to the first user query into each of a plurality of models. The instructions, when executed by the processor, may cause the electronic device to provide a first answer corresponding to the first user query based on first output information output from at least one first model among the plurality of models. The instructions, when executed by the processor, may cause the electronic device to obtain second output information output from at least one second model among the plurality of models. The instructions, when executed by the processor, may cause the electronic device to determine a similarity between the first output information and the second output information. The above instructions, when executed by the processor, may cause the electronic device to provide a second answer obtained based on the first output information and the second output information based on the identified similarity.

[0007] According to one embodiment of the present disclosure, a method may include an operation of obtaining a first user query. The method may include an operation of inputting input information corresponding to the first user query into each of a plurality of models. The method may include an operation of providing a first answer corresponding to the first user query based on first output information output from at least one first model among the plurality of models. The method may include an operation of obtaining second output information output from at least one second model among the plurality of models. The method may include an operation of checking a similarity between the first output information and the second output information. The method may include an operation of providing a second answer obtained based on the first output information and the second output information based on the checked similarity.

[0008] According to one embodiment of the present disclosure, a storage medium storing computer-executable instructions may cause the electronic device to perform at least one operation when executed by a processor of the electronic device. The at least one operation may include obtaining a first user query. The at least one operation may include inputting input information corresponding to the first user query into each of a plurality of models. The at least one operation may include providing a first answer corresponding to the first user query based on first output information output from at least one first model among the plurality of models. The at least one operation may include obtaining second output information output from at least one second model among the plurality of models. The at least one operation may include checking a similarity between the first output information and the second output information. The at least one operation may include providing a second answer obtained based on the first output information and the second output information based on the checked similarity.

[0009] According to one embodiment of the present disclosure, the means for solving the problem are not limited to the above-described means, and means for solving the problem that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present invention pertains from this specification and the attached drawings.

[0010] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.

[0011] FIG. 1 is a block diagram illustrating an integrated intelligence system according to one embodiment of the present disclosure.

[0012] FIG. 2 is a diagram showing a form in which relationship information between concepts and actions is stored in a database according to one embodiment of the present disclosure.

[0013] FIG. 3 is a diagram illustrating a user terminal displaying a screen for processing voice input received through an intelligent app according to one embodiment of the present disclosure.

[0014] FIG. 4 is a block diagram of an electronic device within a network environment according to one embodiment of the present disclosure.

[0015] FIG. 5 is a diagram illustrating an example of a configuration of an electronic device within a network environment according to one embodiment of the present disclosure.

[0016] FIG. 6 is a diagram illustrating an example of a configuration of an electronic device within a network environment according to one embodiment of the present disclosure.

[0017] FIG. 7 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present disclosure to provide an answer corresponding to a user query.

[0018] FIG. 8 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present disclosure to provide an answer corresponding to a user query based on whether a first time interval has elapsed.

[0019] FIG. 9 is an exemplary diagram illustrating an operation of an electronic device according to one embodiment of the present disclosure to provide an answer corresponding to a user query based on whether a first time interval has elapsed.

[0020] FIG. 10 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present disclosure to provide an answer corresponding to a user query based on whether the user query corresponds to a multi-intent query.

[0021] FIG. 11 is an exemplary diagram illustrating an operation of an electronic device according to one embodiment of the present disclosure to provide an answer corresponding to a user query based on a multi-intent query.

[0022] FIG. 12 is a flowchart illustrating an operation of providing an answer corresponding to a user query based on an electronic device executing a first application according to one embodiment of the present disclosure.

[0023] FIG. 13 is an exemplary diagram illustrating an operation of providing an answer corresponding to a user query based on an electronic device executing a first application according to one embodiment of the present disclosure.

[0024] FIG. 14 is a flowchart illustrating an operation of providing an answer corresponding to a user query based on an electronic device executing a first application according to one embodiment of the present disclosure.

[0025] FIG. 15a and FIG. 15b are drawings for explaining an example of the configuration of a conversation system included in an electronic device according to one embodiment of the present disclosure.

[0026] FIG. 16 is an exemplary diagram illustrating an operation of providing an answer corresponding to a user query based on an electronic device executing a first application according to one embodiment of the present disclosure.

[0027] FIG. 17 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present disclosure to provide an answer corresponding to a user query based on whether a similarity between first output information and second output information exceeds a threshold value.

[0028] FIG. 18A and FIG. 18B are exemplary diagrams illustrating an operation of an electronic device according to one embodiment of the present disclosure to provide an answer corresponding to a user query based on whether the similarity between first output information and second output information exceeds a threshold value.

[0029] FIG. 19A and FIG. 19B are exemplary diagrams illustrating an operation of an electronic device according to one embodiment of the present disclosure to provide an answer corresponding to a user query based on whether the similarity between first output information and second output information exceeds a threshold value.

[0030] FIG. 20 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present disclosure to provide an answer corresponding to a user query based on providing first output information as input to a second model.

[0031] FIG. 21A and FIG. 21B are exemplary diagrams illustrating an operation of an electronic device according to one embodiment of the present disclosure to provide an answer corresponding to a user query based on providing first output information as input to a second model.

[0032] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0033] FIG. 1 is a block diagram illustrating an integrated intelligence system according to various embodiments.

[0034] Referring to FIG. 1, an integrated intelligence system (10) of one embodiment may include a user terminal (100), an intelligent server (200), and a service server (300).

[0035] The user terminal (100) of one embodiment may be a terminal device (or electronic device) that can connect to the Internet, and may be, for example, a mobile phone, a smart phone, a personal digital assistant (PDA), a laptop computer, a television, white goods, a wearable device, a head mounted display (HMD), or a smart speaker.

[0036] According to the illustrated embodiment, the user terminal (100) may include a communication interface (110) (e.g., a communication module (490) of FIG. 4), a microphone (120) (e.g., an input module (450) of FIG. 4), a speaker (130) (e.g., an audio output module (455) of FIG. 4), a display (140) (e.g., a display module (460) of FIG. 4), a memory (150) (e.g., a memory (430) of FIG. 4), or a processor (160) (e.g., a processor (420) of FIG. 4). The above-listed components may be operatively or electrically connected to each other.

[0037] The communication interface (110) of one embodiment may be configured to connect to an external device and transmit and receive data. The microphone (120) of one embodiment may receive sound (e.g., user speech) and convert it into an electrical signal. The speaker (130) of one embodiment may output the electrical signal as sound (e.g., voice). The display (140) of one embodiment may be configured to display an image or video. The display (140) of one embodiment may also display a graphical user interface (GUI) of an app (or application program) being executed.

[0038] The display (140) of one embodiment may be configured to display an image or video. The display (140) of one embodiment may also display a graphical user interface (GUI) of a running app (or application program). The display (140) of one embodiment may receive touch input via a touch sensor. For example, the display (140) may receive text input via a touch sensor in an on-screen keyboard area displayed within the display (140).

[0039] In one embodiment, the memory (150) may store a client module (151), a software development kit (SDK) (153), and a plurality of apps (155) (e.g., an application (446) of FIG. 4). The client module (151) and the SDK (153) may constitute a framework (or solution program) for performing general-purpose functions. In addition, the client module (151) or the SDK (153) may constitute a framework for processing user input (e.g., voice input, text input, touch input).

[0040] The plurality of apps (155) stored in the memory (150) of one embodiment may be programs for performing a designated function. According to one embodiment, the plurality of apps (155) may include a first app (155_1) and a second app (155_3). According to one embodiment, each of the plurality of apps (155) may include a plurality of operations for performing a designated function. For example, the apps may include an alarm app, a message app, and / or a schedule app. According to one embodiment, the plurality of apps (155) may be executed by the processor (160) to sequentially execute at least some of the plurality of operations.

[0041] The processor (160) of one embodiment can control the overall operation of the user terminal (100). For example, the processor (160) can be electrically connected to a communication interface (110), a microphone (120), a speaker (130), and a display (140) to perform a designated operation.

[0042] The processor (160) of one embodiment may also execute a program stored in the memory (150) to perform a designated function. For example, the processor (160) may execute at least one of the client module (151) or the SDK (153) to perform the following operations for processing user input. The processor (160) may control the operations of multiple apps (155), for example, through the SDK (153). The operations described below as operations of the client module (151) or the SDK (153) may be operations executed by the processor (160). The processor (160) may control the operations of the user terminal (100) and / or the electronic device of FIGS. 1 to 21 by executing instructions stored in the memory (150). For example, the processor (160) may correspond to multiple processors that collectively perform multiple operations by dividing them among the processors.

[0043] The client module (151) of one embodiment can receive user input. For example, the client module (151) can receive a voice signal corresponding to a user utterance detected through the microphone (120). Alternatively, the client module (151) can receive a touch input detected through the display (140). Alternatively, the client module (151) can receive a text input detected through a keyboard or a visual keyboard. In addition, the client module (151) can receive various types of user input detected through an input module included in the user terminal (100) or an input module connected to the user terminal (100). The client module (151) can transmit the received user input to the intelligent server (200). The client module (151) can transmit status information of the user terminal (100) to the intelligent server (200) together with the received user input. The status information can be, for example, execution status information of an app.

[0044] The client module (151) of one embodiment can receive a result corresponding to the received user input. For example, the client module (151) can receive a result corresponding to the received voice input when the intelligent server (200) can produce a result corresponding to the received user input. The client module (151) can display the received result on the display (140). In addition, the client module (151) can output the received result as audio through the speaker (130).

[0045] The client module (151) of one embodiment can receive a plan corresponding to the received user input. The client module (151) can display the results of executing multiple operations of the app according to the plan on the display (140). For example, the client module (151) can sequentially display the results of executing multiple operations on the display and output audio through the speaker (130). The user terminal (100) can, for example, display only some results of executing multiple operations (e.g., the result of the last operation) on the display and output audio through the speaker (130).

[0046] According to one embodiment, the client module (151) may receive a request from the intelligent server (200) to obtain information necessary to produce a result corresponding to a user input. According to one embodiment, the client module (151) may transmit the necessary information to the intelligent server (200) in response to the request.

[0047] The client module (151) of one embodiment can transmit result information of executing multiple operations according to a plan to the intelligent server (200). The intelligent server (200) can use the result information to confirm that the received user input has been processed correctly.

[0048] The client module (151) of one embodiment may include a voice recognition module. According to one embodiment, the client module (151) may recognize voice inputs that perform limited functions through the voice recognition module. For example, the client module (151) may execute an intelligent app to process voice inputs to perform organic actions through designated inputs (e.g., "Wake up!").

[0049] An intelligent server (200) according to one embodiment may receive information related to a user voice input from a user terminal (100) via a communication network. According to one embodiment, the intelligent server (200) may convert data related to the received voice input into text data. According to one embodiment, the intelligent server (200) may generate a plan for performing a task corresponding to the user voice input based on the text data.

[0050] In one embodiment, the plan may be generated by an artificial intelligence (AI) system. The AI ​​system may be a rule-based system, a neural network-based system (e.g., a feedforward neural network (FNN) or a recurrent neural network (RNN)), or a combination of the foregoing or another AI system. In one embodiment, the plan may be selected from a set of predefined plans or may be generated in real time in response to a user request. For example, the AI ​​system may select at least one plan from a plurality of predefined plans.

[0051] An intelligent server (200) of one embodiment may transmit results according to a generated plan to a user terminal (100), or transmit the generated plan to the user terminal (100). According to one embodiment, the user terminal (100) may display results according to the plan on a display. According to one embodiment, the user terminal (100) may display results of executing an operation according to the plan on a display.

[0052] An intelligent server (200) of one embodiment may include a front end (210), a natural language platform (220), a capsule database (230), an execution engine (240), an end user interface (250), a management platform (260), a big data platform (270), or an analytic platform (280).

[0053] The front end (210) of one embodiment can receive user input from a user terminal (100). The front end (210) can transmit a response corresponding to the user input.

[0054] According to one embodiment, the natural language platform (220) may include an automatic speech recognition module (ASR module) (221), a natural language understanding module (NLU module) (223), a planner module (225), a natural language generator module (NLG module) (227), or a text to speech module (TTS module) (229).

[0055] The automatic speech recognition module (221) of one embodiment can convert voice input received from the user terminal (100) into text data. The natural language understanding module (223) of one embodiment can use the text data of the voice input to determine the user's intention. For example, the natural language understanding module (223) can perform syntactic analysis or semantic analysis on user input in the form of text data to determine the user's intention. The natural language understanding module (223) of one embodiment can use linguistic features (e.g., grammatical elements) of morphemes or phrases to determine the meaning of words extracted from the user input, and can match the meaning of the determined words to the intent to determine the user's intent. The natural language understanding module (223) can obtain intent information corresponding to the user's utterance. The intent information may be information indicating the user's intent determined by interpreting text data. The intent information may include information indicating an action or function that the user intends to execute using the device.

[0056] The planner module (225) of one embodiment can generate a plan using the intent and parameters determined by the natural language understanding module (223). According to one embodiment, the planner module (225) can determine a plurality of domains necessary to perform a task based on the determined intent. The planner module (225) can determine a plurality of operations included in each of the plurality of domains determined based on the intent. According to one embodiment, the planner module (225) can determine parameters necessary to execute the determined plurality of operations or result values ​​output by the execution of the plurality of operations. The parameters and the result values ​​can be defined as concepts of a specified format (or class). Accordingly, the plan can include a plurality of operations and a plurality of concepts determined by the user's intent. The planner module (225) can determine the relationship between the plurality of operations and the plurality of concepts in a stepwise (or hierarchical) manner. For example, the planner module (225) can determine the execution order of a plurality of actions based on the user's intention based on a plurality of concepts. In other words, the planner module (225) can determine the execution order of a plurality of actions based on parameters required for the execution of the plurality of actions and results output by the execution of the plurality of actions. Accordingly, the planner module (225) can generate a plan including association information (e.g., ontology) between the plurality of actions and the plurality of concepts. The planner module (225) can generate the plan using information stored in a capsule database (230) in which a set of relationships between concepts and actions is stored.

[0057] The natural language generation module (227) of one embodiment can convert specified information into text format. The information converted into text format may be in the form of natural language utterances. The text-to-speech conversion module (229) of one embodiment can convert information in text format into information in speech format.

[0058] According to one embodiment, some or all of the functions of the natural language platform (220) may also be implemented in the user terminal (100).

[0059] The capsule database (230) may store information on the relationships between a plurality of concepts and actions corresponding to a plurality of domains. According to one embodiment, a capsule may include a plurality of action objects (or action information) and concept objects (or concept information) included in a plan. According to one embodiment, the capsule database (230) may store a plurality of capsules in the form of a concept action network (CAN). According to one embodiment, the plurality of capsules may be stored in a function registry included in the capsule database (230).

[0060] The capsule database (230) may include a strategy registry that stores strategy information necessary for determining a plan corresponding to a voice input. The strategy information may include reference information for determining a single plan when there are multiple plans corresponding to a user input. According to one embodiment, the capsule database (230) may include a follow-up registry that stores information on follow-up actions for suggesting follow-up actions to a user in a given situation. The follow-up actions may include, for example, follow-up utterances. According to one embodiment, the capsule database (230) may include a layout registry that stores layout information of information output through the user terminal (100). According to one embodiment, the capsule database (230) may include a vocabulary registry that stores vocabulary information included in capsule information. According to one embodiment, the capsule database (230) may include a dialog registry that stores information on dialogue (or interaction) with a user. The capsule database (230) may update stored objects through a developer tool. The developer tool may include, for example, a function editor for updating an action object or a concept object. The developer tool may include a vocabulary editor for updating a vocabulary. The developer tool may include a strategy editor for creating and registering a strategy that determines a plan. The developer tool may include a dialog editor for creating a dialogue with a user.The developer tool may include a follow-up editor that activates follow-up goals and allows editing of follow-up utterances that provide hints. The follow-up goals may be determined based on currently set goals, user preferences, or environmental conditions. In one embodiment, the capsule database (230) may also be implemented within the user terminal (100).

[0061] The execution engine (240) of one embodiment can produce a result using the generated plan. The end user interface (250) can transmit the produced result to the user terminal (100). Accordingly, the user terminal (100) can receive the result and provide the received result to the user. The management platform (260) of one embodiment can manage information used in the intelligent server (200). The big data platform (270) of one embodiment can collect user data. The analysis platform (280) of one embodiment can manage the quality of service (QoS) of the intelligent server (200). For example, the analysis platform (280) can manage the components and processing speed (or efficiency) of the intelligent server (200).

[0062] A service server (300) of one embodiment may provide a service designated to a user terminal (100). In one embodiment, the service designated to the user terminal (100) may include CP service A (301), CP service B (302), or CP service C. For example, CP service A (301) may be a food ordering service. CP service B (302) may be a hotel reservation service. CP service A (301) or CP service B (302) is not limited to the examples described above. According to one embodiment, the service server (300) may be a server operated by a third party. The service server (300) of one embodiment may provide information for generating a plan corresponding to the received user input to the intelligent server (200). The provided information may be stored in the capsule database (230). In addition, the service server (300) may provide result information according to the plan to the intelligent server (200).

[0063] In the integrated intelligence system (10) described above, the user terminal (100) can provide various intelligent services to the user in response to user input. The user input may include, for example, input via a physical button, touch input, or voice input.

[0064] In one embodiment, the user terminal (100) may provide a voice recognition service through an intelligent app (or voice recognition app) stored internally. In this case, for example, the user terminal (100) may recognize a user utterance or voice input received through the microphone and provide the user with a service corresponding to the recognized voice input.

[0065] In one embodiment, the user terminal (100) may perform a designated action based on the received voice input, either alone or in conjunction with the intelligent server and / or service server. For example, the user terminal (100) may execute an app corresponding to the received voice input and perform a designated action through the executed app.

[0066] In one embodiment, when a user terminal (100) provides a service together with an intelligent server (200) and / or a service server (300), the user terminal (100) can detect user speech using the microphone (120) and generate a signal (or voice data) corresponding to the detected user speech. The user terminal (100) can transmit the voice data to the intelligent server (200) using the communication interface (110).

[0067] According to one embodiment, an intelligent server (200) may generate a plan for performing a task corresponding to a voice input received from a user terminal (100), or a result of performing an operation according to the plan, in response to the voice input. The plan may include, for example, a plurality of operations for performing a task corresponding to the user's voice input, and a plurality of concepts related to the plurality of operations. The concept may define parameters input to the execution of the plurality of operations, or result values ​​output by the execution of the plurality of operations. The plan may include association information between the plurality of operations and the plurality of concepts.

[0068] The user terminal (100) of one embodiment can receive the response using the communication interface (110). The user terminal (100) can output a voice signal generated within the user terminal (100) to the outside using the speaker (130), or can output an image generated within the user terminal (100) to the outside using the display (140).

[0069] FIG. 2 is a drawing showing a form in which relationship information between concepts and actions is stored in a database according to various embodiments.

[0070] The capsule database (e.g., capsule database (230)) of the intelligent server (e.g., intelligent server (200) of FIG. 1) may store capsules in the form of a concept action network (CAN) (400). The capsule database may store operations for processing tasks corresponding to a user's voice input and parameters necessary for the operations in the form of a CAN (concept action network) (400).

[0071] The capsule database may store a plurality of capsules (capsule (A) (401), capsule (B) (404)) corresponding to each of a plurality of domains (e.g., applications). According to one embodiment, one capsule (e.g., capsule (A) (401)) may correspond to one domain (e.g., location (geo), application). In addition, one capsule may correspond to at least one service provider (e.g., CP 1 (402), CP 2 (403), CP 3 (406), or CP 4 (405)) for performing a function for a domain related to the capsule. According to one embodiment, one capsule may include at least one operation and at least one concept for performing a specified function.

[0072] The above natural language platform (e.g., the natural language platform (220) of FIG. 1) can generate a plan for performing a task corresponding to a received voice input using capsules stored in a capsule database. For example, the planner module of the natural language platform (e.g., the planner module (225) of FIG. 1) can generate a plan using capsules stored in a capsule database. For example, a plan (407) can be generated using actions (408a, 408b) and concepts (409a, 409b) of capsule A (401) and actions (408c) and concepts (409c) of capsule B (404).

[0073] FIG. 3 is a diagram illustrating a screen for processing voice input received by a user terminal through an intelligent app according to various embodiments.

[0074] The user terminal (100) can execute an intelligent app to process user input through an intelligent server (e.g., the intelligent server (200) of FIG. 1).

[0075] According to one embodiment, on screen 310, when the user terminal (100) recognizes a designated voice input (e.g., wake up!) or receives an input via a hardware key (e.g., a dedicated hardware key), the user terminal (100) may execute an intelligent app for processing the voice input. For example, the user terminal (100) may execute an intelligent app while executing a schedule app. According to one embodiment, the user terminal (100) may display an object (e.g., an icon) (311) corresponding to the intelligent app on a display (e.g., the display (140) of FIG. 1). According to one embodiment, the user terminal (100) may receive a voice input by a user's speech. For example, the user terminal (100) may receive a voice input such as "Tell me my schedule for this week!" According to one embodiment, the user terminal (100) may display a UI (user interface) (313) (e.g., an input window) of the intelligent app, in which text data of the received voice input is displayed, on the display.

[0076] According to one embodiment, on the 320 screen, the user terminal (100) may display a result corresponding to the received voice input on the display (140). For example, the user terminal (100) may receive a plan corresponding to the received user input and display 'this week's schedule' on the display (140) according to the plan.

[0077] FIG. 4 is a block diagram of an electronic device (411) within a network environment (410) according to one embodiment of the present disclosure.

[0078] Referring to FIG. 4, in a network environment (410), an electronic device (411) (e.g., a user terminal (100) of FIGS. 1 to 3) may communicate with an electronic device (412) via a first network (498) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (414) or a server (418) via a second network (499) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (411) may communicate with the electronic device (414) via the server (418). According to one embodiment, the electronic device (411) may include a processor (420), a memory (430), an input module (450), an audio output module (455), a display module (460), an audio module (470), a sensor module (476), an interface (477), a connection terminal (478), a haptic module (479), a camera module (480), a power management module (488), a battery (489), a communication module (490), a subscriber identification module (496), or an antenna module (497). In some embodiments, the electronic device (411) may omit at least one of these components (e.g., the connection terminal (478)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (476), the camera module (480), or the antenna module (497)) may be integrated into one component (e.g., the display module (460)).

[0079] The processor (420) may control at least one other component (e.g., a hardware or software component) of the electronic device (411) connected to the processor (420) by executing, for example, software (e.g., a program (440)), and may perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculation, the processor (420) may store a command or data received from another component (e.g., a sensor module (476) or a communication module (490)) in a volatile memory (432), process the command or data stored in the volatile memory (432), and store the resulting data in a non-volatile memory (434). According to one embodiment, the processor (420) may include a main processor (421) (e.g., a central processing unit or an application processor) or a secondary processor (423) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (411) includes a main processor (421) and a secondary processor (423), the secondary processor (423) may be configured to use less power than the main processor (421) or to be specialized for a given function. The secondary processor (423) may be implemented separately from the main processor (421) or as a part thereof.

[0080] The auxiliary processor (423) may control at least a portion of functions or states associated with at least one component (e.g., a display module (460), a sensor module (476), or a communication module (490)) of the electronic device (411), for example, on behalf of the main processor (421) while the main processor (421) is in an inactive (e.g., sleep) state, or together with the main processor (421) while the main processor (421) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (423) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (480) or a communication module (490)). In one embodiment, the auxiliary processor (423) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (411) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (418)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0081] The memory (430) can store various data used by at least one component (e.g., the processor (420) or the sensor module (476)) of the electronic device (411). The data can include, for example, software (e.g., the program (440)) and input data or output data for commands related thereto. The memory (430) can include a volatile memory (432) or a non-volatile memory (434).

[0082] The program (440) may be stored as software in the memory (430) and may include, for example, an operating system (442), middleware (444), or an application (446).

[0083] The input module (450) can receive commands or data to be used in a component of the electronic device (411) (e.g., a processor (420)) from an external source (e.g., a user) of the electronic device (411). The input module (450) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0084] The audio output module (455) can output audio signals to the outside of the electronic device (411). The audio output module (455) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0085] The display module (460) can visually provide information to an external party (e.g., a user) of the electronic device (411). The display module (460) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (460) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0086] The audio module (470) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (470) can acquire sound through the input module (450), output sound through the sound output module (455), or an external electronic device (e.g., electronic device (412)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (411).

[0087] The sensor module (476) can detect the operating status (e.g., power or temperature) of the electronic device (411) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (476) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0088] The interface (477) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (411) with an external electronic device (e.g., the electronic device (412)). In one embodiment, the interface (477) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0089] The connection terminal (478) may include a connector through which the electronic device (411) may be physically connected to an external electronic device (e.g., the electronic device (412)). According to one embodiment, the connection terminal (478) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0090] The haptic module (479) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (479) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0091] The camera module (480) can capture still images and videos. According to one embodiment, the camera module (480) may include one or more lenses, image sensors, image signal processors, or flashes.

[0092] The power management module (488) can manage power supplied to the electronic device (411). According to one embodiment, the power management module (488) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).

[0093] A battery (489) may power at least one component of the electronic device (411). In one embodiment, the battery (489) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0094] The communication module (490) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (411) and an external electronic device (e.g., electronic device (412), electronic device (414), or server (418)), and the performance of communication through the established communication channel. The communication module (490) may operate independently from the processor (420) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (490) may include a wireless communication module (492) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (494) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (414) via a first network (498) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (499) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (492) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (496) to identify or authenticate the electronic device (411) within a communication network such as the first network (498) or the second network (499).

[0095] The wireless communication module (492) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (492) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (492) may support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (492) may support various requirements specified in the electronic device (411), an external electronic device (e.g., the electronic device (414)), or a network system (e.g., the second network (499)). According to one embodiment, the wireless communication module (492) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0096] The antenna module (497) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (497) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (497) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (498) or the second network (499), may be selected from the plurality of antennas by, for example, the communication module (490). A signal or power may be transmitted or received between the communication module (490) and an external electronic device via the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (497).

[0097] In one embodiment, the antenna module (497) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0098] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0099] According to one embodiment, commands or data may be transmitted or received between the electronic device (411) and an external electronic device (414) via a server (418) connected to a second network (499). Each of the external electronic devices (412 or 414) may be the same or a different type of device as the electronic device (414). According to one embodiment, all or part of the operations executed in the electronic device (411) may be executed in one or more of the external electronic devices (412, 414, or 418). For example, when the electronic device (411) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (411) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (411). The electronic device (411) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (411) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (414) may include an Internet of Things (IoT) device. The server (418) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (414) or the server (418) may be included in the second network (499).The electronic device (411) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0100] FIG. 5 is a diagram illustrating an example of a configuration of an electronic device within a network environment according to one embodiment of the present disclosure.

[0101] Referring to FIG. 5, in one embodiment, an electronic device (501) (e.g., the user terminal (100) of FIGS. 1 to 3 or the electronic device (411) of FIG. 4) may include a microphone (511), a communication circuit (513), a display (515), a speaker (517), a memory (520), and a processor (530).

[0102] In one embodiment, the communication circuit (513) may be included in the communication interface (110) of FIG. 1 or the communication module (490) of FIG. 4. In one embodiment, the communication circuit (513) may communicate with an intelligent server (550) (e.g., the intelligent server (200) of FIG. 1) via a network (541) (e.g., the second network (499) of FIG. 4). In one embodiment, the electronic device (501) (e.g., the communication circuit (513)) may transmit data to the intelligent server (550). For example, the electronic device (501) may transmit input information corresponding to a user query to the intelligent server (550). The electronic device (501) can obtain (or receive) a response corresponding to an utterance analyzed by the intelligent server (550) (e.g., the natural language platform (220) or execution engine (240) of FIG. 1) from the intelligent server (550) via the network (541).

[0103] In one embodiment, an intelligent server (550) (e.g., a prompt handler (551)) may receive input information corresponding to a user query and / or an input prompt included in the user query from an electronic device (501). The intelligent server (550) may transmit the input information to a cloud server (560) or a third-party server (570). The intelligent server (550) may receive (or obtain) output information (or an answer corresponding to the user query) output by at least one of a first third-party model (571) or a second third-party model (573) included in the third-party server (570) in response to the input information from the third-party server (570). The intelligent server (550) can receive (or acquire) output information (or an answer corresponding to a user query) output in response to input information from the cloud server (560) by at least one of the first cloud model (561), the second cloud model (563), or the third cloud model (565) included in the cloud server (560). The intelligent server (550) can transmit the output information to the electronic device (501) via the network (541). In one embodiment, the first third-party model (571), the second third-party model (573), the first cloud model (561), the second cloud model (563), or the third cloud model (565) can include a generative AI model trained to output an answer corresponding to a user utterance based on receiving a user utterance or an input prompt corresponding to the user utterance. Generative AI models may include large language models (LLMs) trained to output text information, image generation models trained to output image information, or retrieval-augmented generation (LLM / RAG) models trained to generate output information based on a search database.A large-scale language model (LLM) can be referred to as a language model comprised of an artificial neural network pre-trained on a massive amount of text data. A large-scale language model (LLM) can contain more than ten times as many parameters (e.g., more than 100 billion parameters) as a conventional language model. A large-scale language model (LLM) can utilize a transformer artificial neural network structure based on an attention mechanism. The attention mechanism is a technique that helps an artificial intelligence model focus on important parts of input data. The attention mechanism can be used to predict output data by predicting the degree to which at least a portion of time-series input data (e.g., input data such as voice or video, or input data of some layers of a neural network) contributes to the intermediate or final output of the neural network. The recurrent neural network (RNN) structure, which sequentially processes each element of a sequence, has poor prediction performance when there is information dependency between long time series distances, but the attention mechanism can consider information dependency between long time series distances by controlling the degree of weight concentration within the overall (or partial) context of the input data.

[0104] For example, a large-scale language model (LLM) or transformer may include an encoder-decoder structure. The encoder may process input data and output compressed information (e.g., an attention mechanism), and the decoder may process the compressed information and output token-based data. Each encoder and decoder may include an independent attention network, and a cross-attention network connecting the encoder and decoder may be included.

[0105] For example, a large-scale language model (LLM) can be trained in two stages: pre-training and fine-tuning. Pre-training involves training a large-scale language model to process massive amounts of text data and acquire general linguistic knowledge. For example, this could involve self-supervised learning, such as predicting the next word using a sequence of previous words in a text sequence. Fine-tuning involves training a large-scale language model to fit a specific domain (e.g., chatbot, translation, summarization, Q&A) or task. This can be further supervised (or adaptive) based on a pre-trained model using a dataset tailored to the domain's objectives. A large-scale language model can perform tasks with text inputs containing natural language, called prompts. For example, a large-scale language model (LLM) could include BERT (bidirectional encoder representations from transformer) and GPT (generative pre-trained transformer). The term "LLM" can refer to the neural network model itself, but can also refer to the model of an LLM-based application (e.g., chatbot, translation, summarization, text classification, sentence generation). For example, an LLM-based chatbot such as chatGPT can also be referred to as an LLM. "LLM" can also include an inference engine that utilizes the LLM neural network model. For example, "entering an input prompt into an LLM" can refer to "entering an input prompt into an LLM-based inference engine."

[0106] In one embodiment, the memory (520) may be included in the memory (150) of FIG. 1 or the memory (430) of FIG. 4. In one embodiment, the memory (520) may include a first application (521), a first on-device model (523), a second on-device model (525), and a third on-device model (527). In one embodiment, the first application (521) may be included in the apps (155) of FIG. 1 or the applications of FIG. 4. The electronic device (501) may execute the first application (521) based on receiving a user utterance. The first application (521) may provide a response corresponding to the user utterance based on obtaining the user utterance. In one embodiment, the electronic device (e.g., the processor (530)) may perform a function of a natural language platform included in the intelligent server (550) by executing the first application (521). For example, the natural language platform may include an automatic speech recognition module (e.g., an automatic speech recognition module (221) of FIG. 1), a natural language understanding module (223) of FIG. 1, a planner module (225) of FIG. 1, a natural language generation module (227) of FIG. 1, or a text-to-speech conversion module (229) of FIG. 1, and the functions of the natural language platform performed by the intelligent server (550) may be performed by the electronic device (501).

[0107] In one embodiment, the first on-device model (523), the second on-device model (525), or the third on-device model (527) may include an AI model trained to output an answer corresponding to a user utterance based on receiving a user utterance or an input prompt corresponding to the user utterance. In one embodiment, the models (523, 525, 527) included in the electronic device (501) may include a deep-learning AI model or a generative AI model. In one embodiment, the time required for the electronic device (501) to obtain output information from the models (523, 525, 527) included in the electronic device (501) may be relatively shorter than the time required for the electronic device (501) to obtain output information from the models (561, 563, 565, 571, 573) included in at least one external electronic device (cloud server (560) or third-party server (570)).

[0108] In one embodiment, the processor (530) may be included in the processor (160) of FIG. 1 or the processor (420) of FIG. 4. In one embodiment, the processor (530) may control the overall operation for providing a response corresponding to a user query. In one embodiment, the processor (530) may include one or more processors for processing user utterances. The operations performed by the processor (530) to process user utterances will be described below with reference to FIGS. 6 to 21A, and 21B.

[0109] Although the electronic device (501) in FIG. 5 is illustrated as including a microphone (511), a communication circuit (513), a display (515), a speaker (517), a memory (520), and / or a processor (530), the present invention is not limited thereto. For example, the electronic device (501) may further include at least one configuration illustrated in FIG. 4.

[0110] FIG. 6 is a diagram illustrating an example of a configuration of an electronic device within a network environment according to one embodiment of the present disclosure.

[0111] In one embodiment, an electronic device (501) (e.g., the user terminal (100) of FIGS. 1 to 3 or the electronic device (411) of FIG. 4) may include a conversation system (610). The conversation system (610) may include a threshold time manager (611) and a sentence generator (613). The threshold time manager (611) may determine whether an answer corresponding to a user query is obtained from a generative AI model (e.g., at least one of a first on-device model (523), a second on-device model (525), a third on-device model (527), a first cloud model (561), a second cloud model (563), a third cloud model (565), a first third-party model (571), or a second third-party model (573)) before a timer corresponding to a threshold time elapses. A time interval (e.g., a threshold time) corresponding to a timer may be changed according to a setting. The time interval corresponding to the timer can be set, for example, between 2 seconds and 3 seconds, but is not limited thereto. The sentence generator (613) can check the similarity between output information (or answers) obtained from multiple models. The similarity between output information or the similarity between answers can include, for example, a probability (or score) indicating the similarity between words included in the answers. The similarity between output information or the similarity between answers can also include a probability (or score) indicating the similarity between contexts of the answers. The sentence generator (613) can generate (or obtain) additional answers based on checking the similarity between output information (or answers) obtained from multiple models.

[0112] In one embodiment, the modules (e.g., the critical time manager (611) and the sentence generator (613)) implemented (or stored) in the electronic device (501) may be implemented in the form of an application, a program, computer code, instructions, a routine, a process, software, firmware, or a combination of at least two or more thereof, executable by a processor (e.g., the processor (530) of FIG. 5). For example, when the modules are executed, the processor may perform an operation corresponding to each of them. Therefore, the description below that a specific module performs an operation may be understood as the processor performing an operation corresponding to the specific module as the specific module is executed. In one embodiment, at least some of the modules may include multiple programs, but are not limited thereto. Alternatively, at least some of the modules may be implemented in hardware form (e.g., a processing circuit (not shown)).

[0113] FIG. 7 is a flowchart illustrating an operation of an electronic device according to one embodiment of the present disclosure to provide an answer corresponding to a user query.

[0114] In one embodiment, the operations illustrated in FIG. 7 may be performed in various orders, not limited to the order illustrated. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. In one embodiment, more operations may be performed than those illustrated in FIG. 7, or at least one operation may be performed less than those illustrated in FIG.

[0115] Referring to FIG. 7, in operation 701, in one embodiment, an electronic device (501) (e.g., processor (530) of FIG. 5) may obtain a first user query.

[0116] In one embodiment, the electronic device (501) may obtain a user query through at least one component included in the electronic device (501). The user query may also be referred to as a "user utterance." For example, the electronic device (501) may obtain a first user query through a microphone (e.g., the microphone (511) of FIG. 5 ).

[0117] In operation 703, in one embodiment, the electronic device (501) may input input information corresponding to a first user query into each of a plurality of models. In one embodiment, the electronic device (501) may obtain input information (or an input prompt) corresponding to the first user query. The input information corresponding to the first user query may include information of a type that can be input into the AI ​​model (e.g., text information, image information, or audio information). The operation of obtaining the input information corresponding to the first user query may be referred to as a "preprocessing operation." The electronic device (501) may input input information corresponding to the first user utterance into each of a plurality of models (e.g., at least one of the first on-device model (523), the second on-device model (525), the third on-device model (527), the first cloud model (561), the second cloud model (563), the third cloud model (565), the first third-party model (571), or the second third-party model (573)) so that answers corresponding to the user query are generated in parallel by the plurality of models. The electronic device may, for example, provide input information corresponding to the first user query as an input to each of the on-device models (e.g., at least one of the first on-device model (523), the second on-device model (525), or the third on-device model (527)). The electronic device may transmit input information corresponding to the first user query to an intelligent server (e.g., the intelligent server (550) of FIG. 5) via a network (e.g., the network (541) of FIG. 5) so that each of the external models (e.g., at least one of the first cloud model (561), the second cloud model (563), the third cloud model (565), the first third-party model (571), or the second third-party model (573)) is provided with the input information.

[0118] In operation 705, in one embodiment, the electronic device (501) may provide a first answer based on first output information output from at least one first model.

[0119] In one embodiment, the electronic device (501) may provide a first answer corresponding to the first user query based on first output information output from at least one first model among a plurality of models. In one embodiment, the first model may be an AI model that first provided an answer corresponding to the first user query. The first model may be, for example, an AI model having a relatively fast computation speed based on personalization. In one embodiment, the electronic device (501) may provide a first answer corresponding to the first user query through a display (e.g., the display (515) of FIG. 5). In one embodiment, the at least one first model may be an AI model that outputs output information within a threshold time. For example, the electronic device (501) may provide the first answer based on output information output from an AI model set according to a priority based on reliability and / or accuracy among the output information output from the at least one first model. The electronic device (501) can provide a first answer based on output information output from the first model, taking into account that the average computational speed of the generative AI model may be relatively slow.

[0120] In operation 707, in one embodiment, the electronic device (501) may obtain second output information. The electronic device (501) may obtain second output information output from at least one second model among a plurality of models. The second model may be, for example, a generative AI model that provides relatively rich answers. The electronic device (501) may obtain the second output information from one second model or may obtain the second output information from a plurality of second models.

[0121] In operation 709, in one embodiment, the electronic device (501) can determine the similarity between the first output information and the second output information. The similarity can be determined based on a method of obtaining, for example, cosine similarity, Euclidean similarity, Manhattan similarity, or Jaccard similarity, but is not limited thereto. The similarity can include, for example, a probability (or score) indicating the degree of similarity between the first output information and the second output information. The similarity can also include, for example, a class corresponding to a category into which the first output information and the second output information are included.

[0122] In operation 711, in one embodiment, the electronic device (501) may provide a second answer obtained based on the first output information and the second output information.

[0123] In one embodiment, the electronic device (501) may provide a second answer based on the verified similarity. The second answer may be obtained based on the first output information and the second output information. If the similarity between the first output information and the second output information is verified to be high, the electronic device (501) may provide an answer similar to the first answer. In one embodiment, if the similarity between the first output information and the second output information is verified to be relatively low, the electronic device (501) may also provide a supplementary answer to the first answer. By providing the first answer before providing the second answer, the electronic device (501) may provide an answer relatively quickly in response to a user query. By providing the second answer based on the output information of the second model, which has a relatively slow computation speed, after providing the first answer, the electronic device (501) may provide a relatively richer answer than the first answer.

[0124] FIG. 8 is a flowchart illustrating an operation of an electronic device according to an embodiment of the present disclosure to provide an answer corresponding to a user query based on whether a first time interval has elapsed. The embodiment of FIG. 8 will be described with reference to FIG. 9. FIG. 9 is an exemplary diagram illustrating an operation of an electronic device according to an embodiment of the present disclosure to provide an answer corresponding to a user query based on whether a first time interval has elapsed.

[0125] Referring to FIG. 8, in operation 801, in one embodiment, the electronic device (501) (e.g., the processor (530) of FIG. 5) may determine whether output information corresponding to input information has been obtained from at least one generative AI model before a first time interval has elapsed. In one embodiment, the electronic device (501) (e.g., the threshold time manager (611) of FIG. 6) may determine whether output information corresponding to input information has been obtained from at least one generative AI model (e.g., at least one of the first on-device model (523), the second on-device model (525), the third on-device model (527), the first cloud model (561), the second cloud model (563), the third cloud model (565), the first third-party model (571), or the second third-party model (573)) before the first time interval has elapsed.

[0126] Referring to FIG. 9, in one embodiment, an electronic device (501) may obtain a first user query (901). The electronic device (501) may input input information corresponding to the first user query into each of a plurality of models (903). The electronic device (101) may determine whether output information corresponding to the input information has been obtained from at least one generative AI model before a first time interval (910) elapses. The operation of the electronic device (501) obtaining the first user query and the operation of the electronic device (501) inputting input information corresponding to the first user query into each of the plurality of models have been described in detail in operations 701 and 703 of FIG. 7, and therefore, any overlapping description may not be repeated herein. In one embodiment, the generative AI model may include at least one first generative AI model (e.g., at least one of a first on-device model (523), a second on-device model (525), or a third on-device model (527)) included in the electronic device (501) and at least one second generative AI model (e.g., at least one of a first cloud model (561), a second cloud model (563), a third cloud model (565), a first third-party model (571), or a second third-party model (573)) included in at least one external electronic device (560, 570).

[0127] In operation 803, in one embodiment, the electronic device (501) may provide an answer based on confirming that output information corresponding to the input information has been obtained from at least one generative AI model before the first time interval elapses (operation 801: Yes). In one embodiment, the electronic device (501) may confirm a generative AI model that has output information within the first time interval (or, threshold time). The electronic device (501) may provide an answer including a relatively rich answer based on confirming that output information output from the generative AI model has been obtained before the first time interval elapses.

[0128] In operation 805, in one embodiment, the electronic device (501) may provide a first answer based on the first output information based on determining that no output information corresponding to the input information has been obtained from at least one generative AI model before the first time interval elapses (operation 801: No).

[0129] In one embodiment, referring to FIG. 9, the electronic device (101) may obtain (905) first output information from at least one first model. The electronic device (101) may provide (907) a first answer based on the first output information. Considering that the average computational speed of a generative AI model may be relatively slow, the electronic device (501) may quickly provide the first answer based on the output information output from the first model.

[0130] In operation 807, in one embodiment, the electronic device (501) may provide a second answer based on second output information obtained from the generative AI model after a first time interval has elapsed. Referring to FIG. 9, the electronic device (501) may obtain (909) second output information including a relatively rich answer from a generative AI model having a relatively slow computation speed. The electronic device (501) may provide a second answer (911) based on obtaining the second output information. The electronic device (501) may provide a relatively rich answer by additionally providing the second answer after providing the first answer relatively quickly. In one embodiment, in the second answer, overlapping content with the first answer may be deleted. The second answer may include additional information regarding the first answer.

[0131] FIG. 10 is a flowchart illustrating an operation of an electronic device according to an embodiment of the present disclosure to provide an answer corresponding to a user query based on whether the user query corresponds to a multi-intent query. The embodiment of FIG. 10 will be described with reference to FIG. 11. FIG. 11 is an exemplary diagram illustrating an operation of an electronic device according to an embodiment of the present disclosure to provide an answer corresponding to a user query based on a multi-intent query.

[0132] Referring to FIG. 10, in operation 1001, in one embodiment, the electronic device (501) (e.g., the processor (530) of FIG. 5) may determine whether a first user query corresponds to a multi-intent query. In one embodiment, the user query may include at least one intent. An intent may, for example, indicate a purpose or topic of the user query. A user query including one intent may be referred to as a “single-intent query.” A user query including multiple intents may be referred to as a “multi-intent query.” In operation 1003, in one embodiment, the electronic device (501) may provide an answer based on determining that the first user query does not correspond to a multi-intent query (operation 1001: No). In one embodiment, the electronic device (501) may determine that the first user query is a single-intent query. In one embodiment, the electronic device (501) may provide a first answer and a second answer corresponding to the first user query based on performing operations 701 to 711 of FIG. 7.

[0133] In operation 1005, in one embodiment, the electronic device (501) may input input information corresponding to each of the first intent, the second intent, and the multi-intent into each of the plurality of models based on determining that the first user query corresponds to a multi-intent query (operation 1001: Yes). For example, the multi-intent query may include intents as shown in Table 1.

[0134] Query "I have a question. What's the weather tomorrow? What are the tourist attractions in New York?" Intent 1: "Tomorrow's weather." Intent 2: "Tourist attractions in New York." Multi-intent: "Tomorrow's weather, New York tourist attractions."

[0135] Referring to FIG. 11, in one embodiment, the electronic device (501) may provide input information corresponding to each of the first intent, the second intent, and the multi-intent as inputs to the first on-device model (523). The electronic device (501) may provide input information corresponding to each of the first intent, the second intent, and the multi-intent as inputs to the second on-device model (525). The electronic device (501) may provide input information corresponding to each of the first intent, the second intent, and the multi-intent as inputs to the third on-device model (527). The electronic device (501) can transmit input information to an intelligent server (e.g., an intelligent server (550) of FIG. 5) via a network (e.g., a network (541) of FIG. 5) so that input information corresponding to each of the first intent, the second intent, and the multi-intent is provided as input to the first cloud model (561). The electronic device (501) can transmit input information to an intelligent server via a network so that input information corresponding to each of the first intent, the second intent, and the multi-intent is provided as input to the second cloud model (563). The electronic device (501) can transmit input information to an intelligent server via a network so that input information corresponding to each of the first intent, the second intent, and the multi-intent is provided as input to the third cloud model (565). The electronic device (501) can transmit input information to an intelligent server via a network so that input information corresponding to each of the first intent, the second intent, and the multi-intent is provided as input to a first third-party model (571). The electronic device (501) can transmit input information to an intelligent server via a network so that input information corresponding to each of the first intent, the second intent, and the multi-intent is provided as input to a second third-party model (573).In operation 1007, in one embodiment, the electronic device (501) may provide a first answer corresponding to the first intent and the second intent based on obtaining the first output information from the first model before the first time interval elapses. In one embodiment, the electronic device (501) may provide the first answer corresponding to the first intent and the second intent based on the first output information of the first model that provides the output information relatively quickly. The electronic device (501) may quickly provide the first answer including, for example, an answer corresponding to “the weather tomorrow” and an answer corresponding to “tourist attractions in New York.”

[0136] In operation 1009, in one embodiment, the electronic device (501) may provide a second answer corresponding to the multi-intent based on obtaining second output information from the second model after the first time interval has elapsed. In one embodiment, the electronic device (501) may provide the second answer corresponding to the multi-intent based on the second output information of the second model that outputs a relatively richer answer. For example, the electronic device (501) may provide an additional answer that is relatively richer than the first answer by providing a second answer that includes an answer corresponding to "the weather in New York tomorrow" and an answer corresponding to "tourist attractions in New York suitable for tomorrow's weather."

[0137] FIG. 12 is a flowchart illustrating an operation of providing an answer corresponding to a user query based on an electronic device executing a first application according to an embodiment of the present disclosure. The embodiment of FIG. 12 will be described with reference to FIG. 13. FIG. 13 is an exemplary diagram illustrating an operation of providing an answer corresponding to a user query based on an electronic device executing a first application according to an embodiment of the present disclosure.

[0138] Referring to FIG. 12, in operation 1201, in one embodiment, an electronic device (501) (e.g., a processor (530) of FIG. 5) may execute a first application (e.g., a first application (521) of FIG. 5). The electronic device (501) may execute the first application associated with providing an answer corresponding to a query. The first application may provide an answer corresponding to the user utterance based on obtaining the user utterance. In one embodiment, the electronic device may provide at least a portion of the functions of a natural language platform included in an intelligent server by executing the first application.

[0139] Referring to reference numeral 1310 of FIG. 13, in one embodiment, the electronic device (501) may, based on confirming a user input (e.g., a voice command or a key input) for executing a first application, display an object (1311) indicating that the first application is executing through the display (515). In one embodiment, the electronic device (501) may, based on obtaining a user query, display a window (1313) including text information corresponding to the user query through the display (515).

[0140] In operation 1203, in one embodiment, the electronic device (501) may provide (1321) the first answer based on obtaining the first user query while the first application is running.

[0141] Referring to reference numeral 1330 of FIG. 13, the electronic device (501) may display a window (1331) including a first answer through the display (515) while the first application is running. For example, the electronic device (501) may display objects (1333) including a list of applications that can provide answers corresponding to the query “summary of books written by Hillary Clinton” while displaying an object (1311) indicating that the first application is running. The electronic device (501) may display an object (1335a) corresponding to an e-book application, an object (1335b) corresponding to a streaming application, and an object (1335c) corresponding to an Internet application.

[0142] FIG. 14 is a flowchart illustrating an operation of providing an answer corresponding to a user query based on an electronic device executing a first application according to an embodiment of the present disclosure. The embodiment of FIG. 14 will be described with reference to FIGS. 15A, 15B, and 16. FIGS. 15A and 15B are diagrams illustrating an example of a configuration of a conversation system included in an electronic device according to an embodiment of the present disclosure. FIG. 16 is an exemplary diagram illustrating an operation of providing an answer corresponding to a user query based on an electronic device executing a first application according to an embodiment of the present disclosure.

[0143] Referring to FIG. 14, in operation 1401, in one embodiment, the electronic device (501) (e.g., the processor (530) of FIG. 5) may provide a second answer based on second output information being obtained while the first application is running.

[0144] Referring to FIG. 15A, in one embodiment, a conversation system (610) of an electronic device (501) may include a threshold time manager (611), a sentence generator (613), a similarity checker (1501), and an information analyzer (1503). In one embodiment, the threshold time manager (611) and the sentence generator (613) may provide at least some of the same functions as the threshold time manager (611) and the sentence generator (613) of FIG. 6. The information analyzer (1503) may output information associated with a similar context between first output information output from a first model and second output information output from a second model.

[0145] Referring to FIG. 15B, in one embodiment, the similarity checker (1501) can verify (or obtain) the similarity between the first output information and the second output information. The sentence generator (613) can output a similar answer or a supplementary answer to the first output information based on the similarity verified by the similarity checker (1501). The conversation system (610) can output a second answer including the similar answer or a second answer including the supplementary answer.

[0146] Referring to FIG. 16, in one embodiment, the electronic device (501) may provide a second answer based on the acquisition of second output information (1601). In one embodiment, the electronic device (501) may determine that the second answer (1601) has a relatively low similarity to the context of the first answer (e.g., the window (1333) including the list of FIG. 13). Based on the determination that the second answer (1601) is different from the first answer, the electronic device (501) may provide the second answer while the first application is running.

[0147] Referring to reference numeral 1610 of FIG. 16, in one embodiment, the electronic device (501) may display a window (1613) associated with a second answer including a secondary answer, such as “I’ll explain it to you without selecting a service,” while displaying an object (1611) indicating that a first application is running through the display (515). The electronic device (501) may provide the second answer including the secondary answer based on determining that the similarity between the first answer and the second answer (1601) is low.

[0148] In operation 1403, in one embodiment, the electronic device (501) may provide at least one answer corresponding to the second user query based on obtaining a second user query corresponding to the second answer. The electronic device (501) may sequentially provide answers corresponding to the second user query based on performing operations (e.g., operations 701 to 709 of FIG. 7) based on obtaining a second user query different from the first user query, thereby providing a rich answer after providing a quick answer.

[0149] FIG. 17 is a flowchart illustrating an operation of an electronic device according to an embodiment of the present disclosure to provide an answer corresponding to a user query based on whether a similarity between first output information and second output information exceeds a threshold value. The embodiment of FIG. 17 will be described with reference to FIGS. 18A, 18B, 19A, and 19B. FIGS. 18A and 18B are exemplary diagrams illustrating an operation of an electronic device according to an embodiment of the present disclosure to provide an answer corresponding to a user query based on whether a similarity between first output information and second output information exceeds a threshold value. FIGS. 19A and 19B are exemplary diagrams illustrating an operation of an electronic device according to an embodiment of the present disclosure to provide an answer corresponding to a user query based on whether a similarity between first output information and second output information exceeds a threshold value.

[0150] Referring to FIG. 17, in operation 1701, in one embodiment, the electronic device (501) (e.g., the processor (530) of FIG. 5) may determine whether the similarity between the first output information and the second output information exceeds a first threshold value.

[0151] In one embodiment, the electronic device (501) (e.g., the sentence generator (613) of FIG. 6 or the similarity checker (1501) of FIG. 15A) may determine whether the similarity between the first output information and the second output information exceeds a first threshold. For example, the electronic device (501) may determine the similarity between a sentence corresponding to the first output information and a sentence corresponding to the second output information based on text vectorization. The electronic device (501) may perform text vectorization using a weight such as, for example, TF-IDF (term frequency - inverse document frequency). The similarity may be determined as a probability (or score). The similarity may also be determined as a class corresponding to a category. The specific value of the first threshold for determining whether the similarity between the first output information and the second output information is relatively large may vary depending on the embodiment.

[0152] In operation 1703, in one embodiment, the electronic device (501) may provide a second answer that includes a similar answer to the first answer based on determining that the similarity between the first output information and the second output information exceeds a first threshold (operation 1701: Yes).

[0153] Referring to reference numeral 1810 of FIG. 18A, in one embodiment, the electronic device (501) may, based on obtaining a first user query, display an object (1811) indicating that a first application is running through the display (515), while displaying a window (1813) including text information corresponding to the first user query. The electronic device (501) may, based on obtaining the first output information, provide a first answer (1821). Referring to reference numeral 1830 of FIG. 18A, in one embodiment, the electronic device (501) may display a window (1831) including text information corresponding to the first answer through the display (515).

[0154] Referring to FIG. 18B, in one embodiment, the electronic device (501) may obtain second output information (1841). The electronic device (501) may determine the similarity between the first output information and the second output information. Referring to reference numeral 1850, based on determining that the similarity exceeds a first threshold, the electronic device (501) may display a window (1851) including a similar answer through the display (515). The electronic device (501) may provide, for example, a phrase such as “And let me explain in more detail” as a similar answer, indicating that additional information is provided. In the second answer, a sentence that overlaps with the first answer, such as “Michelle Obama married Barack Obama on October 3, 1992,” may be deleted.

[0155] In operation 1705, in one embodiment, the electronic device (501) can determine whether the similarity between the first output information and the second output information exceeds a second threshold based on determining that the similarity between the first output information and the second output information does not exceed a first threshold (operation 1701: No). The electronic device (501) can determine whether the second answer is worthy as a supplementary answer to the first answer based on determining whether the similarity between the first output information and the second output information exceeds the second threshold. In one embodiment, the electronic device (501) can also determine whether the second answer is worthy as a supplementary answer to the first answer based on determining the similarity between a chunk included in the first output information and a chunk included in the second output information. In operation 1709, in one embodiment, the electronic device (501) may discard the second output information based on determining that the similarity between the first output information and the second output information does not exceed a second threshold (operation 1705: No). The electronic device (501) may not provide the second answer based on determining that the second output information is not valuable as a secondary answer.

[0156] In operation 1707, in one embodiment, the electronic device (501) may provide a third answer including a secondary answer to the first answer based on determining that the similarity between the first output information and the second output information exceeds a second threshold (operation 1705: Yes).

[0157] Referring to reference numeral 1910 of FIG. 19A, in one embodiment, the electronic device (501) may, based on obtaining a first user query, display an object (1911) indicating that a first application is running through the display (515), while displaying a window (1913) including text information corresponding to the first user query. The electronic device (501) may, based on obtaining the first output information, provide a first answer (1921). Referring to reference numeral 1930 of FIG. 19A, in one embodiment, the electronic device (501) may display a window (1931) including text information corresponding to the first answer through the display (515).

[0158] Referring to FIG. 19B, in one embodiment, the electronic device (501) may obtain second output information (1941). The electronic device (501) may determine the similarity between the first output information and the second output information. Referring to reference numeral 1950, based on determining that the similarity is less than a first threshold and greater than a second threshold, the electronic device (501) may display a window (1951) including a secondary answer through the display (515). The electronic device (501) may provide a phrase such as "LLM also" as the secondary answer, indicating that it provides information in a different category, for example. The second answer may include an answer that includes examples of how the term LLM is used in fields other than law.

[0159] FIG. 20 is a flowchart illustrating an operation of providing an answer corresponding to a user query based on providing first output information as input to a second model by an electronic device according to an embodiment of the present disclosure. The embodiment of FIG. 20 will be described with reference to FIGS. 21A and 21B. FIGS. 21A and 21B are exemplary diagrams illustrating an operation of providing an answer corresponding to a user query based on providing first output information as input to a second model by an electronic device according to an embodiment of the present disclosure.

[0160] Referring to FIG. 20, at operation 2001, in one embodiment, an electronic device (501) (e.g., processor (530) of FIG. 5) may provide first output information as input to at least one second model.

[0161] Referring to FIG. 21A, in one embodiment, the electronic device (501) may provide first output information output from a first model (2110) as input to at least one second model (2120). A structure in which first output information is provided as input to at least one second model may be referred to as a “hybrid structure.”

[0162] In operation 2003, in one embodiment, the electronic device (501) may provide a second answer based on second output information being output from at least one second model.

[0163] Referring to FIG. 21A, in one embodiment, at least one second model (2120) may output a second answer (or second output information) based on being provided with first output information and a first user query as input.

[0164] Referring to FIG. 21B, in one embodiment, the first output information output from the first model (2110) may be input to the dialogue system (610). The dialogue system (610) may output a second answer based on receiving the first output information output from the first model (2110) and the second output information output from the second model (2120) as inputs. The electronic device (501) may provide a similar answer or a supplementary answer to the first output information as the second answer, thereby providing a relatively rich second answer after quickly providing the first answer.

[0165] According to one embodiment of the present disclosure, an electronic device (e.g., the electronic device 501 of FIG. 5) may include a memory (e.g., the memory 520 of FIG. 5) that stores instructions and a processor (e.g., the processor 530 of FIG. 5). The instructions, when executed by the processor 530, may cause the electronic device 501 to obtain a first user query. The instructions, when executed by the processor 530, may cause the electronic device 501 to input input information corresponding to the first user query into each of a plurality of models. The instructions, when executed by the processor 530, may cause the electronic device 501 to provide a first answer corresponding to the first user query based on first output information output from at least one first model among the plurality of models. The instructions, when executed by the processor (530), may cause the electronic device (501) to obtain second output information output from at least one second model among the plurality of models. The instructions, when executed by the processor (530), may cause the electronic device (501) to determine a similarity between the first output information and the second output information. The instructions, when executed by the processor (530), may cause the electronic device (501) to provide a second answer obtained based on the first output information and the second output information based on the determined similarity.

[0166] In one embodiment, the instructions, when executed by the processor (530), may cause the electronic device (501) to determine whether output information corresponding to the input information is obtained from at least one generative AI model among the plurality of models before a first time interval elapses. The instructions, when executed by the processor (530), may cause the electronic device (501) to provide the first answer based on the first output information based on determining that the output information is not obtained from the at least one generative AI model before the first time interval elapses. The instructions, when executed by the processor (530), may cause the electronic device (501) to provide the second answer based on that the second output information is obtained from the generative AI model after the first time interval elapses.

[0167] In one embodiment, the instructions, when executed by the processor (530), may cause the electronic device (501) to determine whether the first user query corresponds to a multi-intent query. The instructions, when executed by the processor (530), may cause the electronic device (501) to input input information corresponding to each of a first intent, a second intent, and a multi-intent included in the first user query into each of the plurality of models based on determining that the first user query corresponds to the multi-intent query. The instructions, when executed by the processor (530), may cause the electronic device (501) to provide a first answer corresponding to the first intent and the second intent based on obtaining the first output information from the first model before a first time interval elapses. The above instructions, when executed by the processor (530), may cause the electronic device (501) to provide a second answer corresponding to the multi-intent based on obtaining the second output information from the second model after the first time interval has elapsed.

[0168] In one embodiment, the generative AI model may include at least one first generative AI model included in the electronic device (501) and at least one second generative AI model included in at least one external electronic device (560, 570).

[0169] In one embodiment, the instructions, when executed by the processor (530), may cause the electronic device (501) to execute a first application associated with providing an answer corresponding to a query. The instructions, when executed by the processor (530), may cause the electronic device (501) to provide the first answer based on obtaining the first user query while the first application is executing.

[0170] In one embodiment, the instructions, when executed by the processor (530), may cause the electronic device (501) to provide the second answer based on the second output information being acquired while the first application is running. The instructions, when executed by the processor (530), may cause the electronic device (501) to provide at least one answer corresponding to a second user query based on acquiring a second user query corresponding to the second answer.

[0171] In one embodiment, the instructions, when executed by the processor (530), may cause the electronic device (501) to determine whether a similarity between the first output information and the second output information exceeds a first threshold. The instructions, when executed by the processor (530), may cause the electronic device (501) to provide the second answer including a similar answer to the first answer based on determining that the similarity exceeds the first threshold.

[0172] In one embodiment, the instructions, when executed by the processor (530), may cause the electronic device (501) to determine whether the similarity exceeds a second threshold value less than the first threshold value, based on determining that the similarity does not exceed the first threshold value. The instructions, when executed by the processor (530), may cause the electronic device (501) to provide a third answer that includes a secondary answer to the first answer, based on determining that the similarity exceeds the second threshold value.

[0173] In one embodiment, the instructions, when executed by the processor (530), may cause the electronic device (501) to provide the first output information as input to at least one second model. The instructions, when executed by the processor (530), may cause the electronic device (501) to provide the second answer based on the second output information being output from the at least one second model.

[0174] In one embodiment, the electronic device (501) may further include a microphone (511) and a display (515). The instructions, when executed by the processor (530), may cause the electronic device (501) to obtain the first user query through the microphone (511). The instructions, when executed by the processor (530), may cause the electronic device (501) to provide a first answer corresponding to the first user query through the display (515).

[0175] A method according to one embodiment may include an operation of obtaining a first user query. The method may include an operation of inputting input information corresponding to the first user query into each of a plurality of models. The method may include an operation of providing a first answer corresponding to the first user query based on first output information output from at least one first model among the plurality of models. The method may include an operation of obtaining second output information output from at least one second model among the plurality of models. The method may include an operation of checking a similarity between the first output information and the second output information. The method may include an operation of providing a second answer obtained based on the first output information and the second output information based on the checked similarity.

[0176] In one embodiment, the method may further include an operation of determining whether output information corresponding to the input information is obtained from at least one generative AI model among the plurality of models before the first time interval elapses. The method may further include an operation of providing the first answer based on the first output information based on determining that the output information is not obtained from the at least one generative AI model before the first time interval elapses. The method may further include an operation of providing the second answer based on whether the second output information is obtained from the generative AI model after the first time interval elapses.

[0177] In one embodiment, the method may further include an operation of determining whether the first user query corresponds to a multi-intent query. The method may further include an operation of inputting input information corresponding to each of the first intent, the second intent, and the multi-intent included in the first user query into each of the plurality of models based on determining that the first user query corresponds to the multi-intent query. The method may further include an operation of providing a first answer corresponding to the first intent and the second intent based on obtaining the first output information from the first model before a first time interval elapses. The method may further include an operation of providing a second answer corresponding to the multi-intent based on obtaining the second output information from the second model after a first time interval elapses.

[0178] In one embodiment, the generative AI model may include at least one first generative AI model included in the electronic device (501) and at least one second generative AI model included in at least one external electronic device (560, 570).

[0179] In one embodiment, the method may further include executing a first application associated with providing an answer corresponding to the query. The method may further include providing the first answer based on obtaining the first user query while the first application is running.

[0180] In one embodiment, the method may further include providing the second answer based on the acquisition of the second output information while the first application is running. The method may further include providing at least one answer corresponding to the second user query based on the acquisition of the second user query corresponding to the second answer.

[0181] In one embodiment, the method may further include an operation of determining whether the similarity between the first output information and the second output information exceeds a first threshold. The method may further include an operation of providing the second answer, which includes a similar answer to the first answer, based on determining that the similarity exceeds the first threshold.

[0182] In one embodiment, the method may further include, based on determining that the similarity does not exceed the first threshold, determining whether the similarity exceeds a second threshold that is less than the first threshold. The method may further include, based on determining that the similarity exceeds the second threshold, providing a third answer that includes a supplementary answer to the first answer.

[0183] In one embodiment, the method may further include providing the first output information as input to at least one second model. The method may further include providing the second answer based on the second output information being output from the at least one second model.

[0184] In one embodiment, a storage medium storing computer-executable instructions, when executed by a processor (530) of an electronic device (501), may cause the electronic device (501) to perform at least one operation. The at least one operation may include obtaining a first user query. The at least one operation may include inputting input information corresponding to the first user query into each of a plurality of models. The at least one operation may include providing a first answer corresponding to the first user query based on first output information output from at least one first model among the plurality of models. The at least one operation may include obtaining second output information output from at least one second model among the plurality of models. The at least one operation may include checking a similarity between the first output information and the second output information. The at least one operation may include an operation of providing a second answer obtained based on the first output information and the second output information, based on the confirmed similarity.

[0185] An electronic device according to an embodiment disclosed in this document may take various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a home appliance. The electronic device according to an embodiment of this document is not limited to the aforementioned devices.

[0186] It should be understood that the embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to specific embodiments, but include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0187] The term "module" used in one embodiment of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0188] An embodiment of the present document may be implemented as software (e.g., a program (440)) including one or more instructions stored in a storage medium (e.g., an internal memory (436) or an external memory (438)) readable by a machine (e.g., an electronic device (411)). For example, a processor (e.g., a processor (420)) of the machine (e.g., an electronic device (411)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0189] According to one embodiment, the method according to one embodiment disclosed in the present document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0190] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0191] Additionally, the structure of the data used in the embodiments of the present invention described above can be recorded on a computer-readable recording medium through various means. The computer-readable recording medium includes storage media such as magnetic storage media (e.g., ROM, floppy disk, hard disk, etc.) and optical reading media (e.g., CD-ROM, DVD, etc.).

[0192] The present invention has been described above, focusing on preferred embodiments thereof. Those skilled in the art will appreciate that the present invention can be implemented in modified forms without departing from its essential characteristics. Therefore, the disclosed embodiments should be considered illustrative rather than limiting. The scope of the present invention is set forth in the claims, not the foregoing description, and all differences within the scope equivalent thereto should be construed as being encompassed by the present invention.

Claims

1. In the electronic device (501), Memory (520) for storing instructions; and A processor (530) is included, and the instructions, when executed by the processor (530), cause the electronic device (501) to: Obtain the first user query, Input information corresponding to the first user query is input into each of the plurality of models, Provide a first answer corresponding to the first user query based on first output information output from at least one first model among the plurality of models; Obtain second output information output from at least one second model among the plurality of models, Check the similarity between the first output information and the second output information, An electronic device (501) that causes a second answer to be provided based on the first output information and the second output information based on the confirmed similarity.

2. In paragraph 1, The above instructions, when executed by the processor (530), cause the electronic device (501) to: Before the first time interval elapses, it is confirmed whether output information corresponding to the input information is obtained from at least one generative AI model among the plurality of models, Providing the first answer based on the first output information based on confirming that the output information is not obtained from the at least one generative AI model before the first time interval elapses, An electronic device (501) that causes the second answer to be provided based on the second output information being obtained from the generative AI model after the first time interval has elapsed.

3. In paragraph 1 or 2, The above instructions, when executed by the processor (530), cause the electronic device (501) to: Check whether the above first user query corresponds to a multi-intent query, Based on confirming that the first user query corresponds to the multi-intent query, input information corresponding to each of the first intent, the second intent, and the multi-intent included in the first user query is input to each of the plurality of models, Provide a first answer corresponding to the first intent and the second intent based on obtaining the first output information from the first model before the first time interval elapses, An electronic device (501) that causes a second answer corresponding to the multi-intent to be provided based on obtaining the second output information from the second model after the first time interval has elapsed.

4. In any one of paragraphs 1 to 3, An electronic device (501), wherein the generative AI model comprises at least one first generative AI model included in the electronic device (501) and at least one second generative AI model included in at least one external electronic device (560, 570).

5. In any one of paragraphs 1 to 4, The above instructions, when executed by the processor (530), cause the electronic device (501) to: Executes a first application associated with providing an answer corresponding to the query, An electronic device (501) that causes the first answer to be provided based on obtaining the first user query while the first application is running.

6. In any one of paragraphs 1 to 5, The above instructions, when executed by the processor (530), cause the electronic device (501) to: While the first application is running, the second answer is provided based on the second output information being obtained, An electronic device (501) that causes at least one answer corresponding to the second user question to be provided based on obtaining a second user question corresponding to the second answer.

7. In any one of paragraphs 1 to 6, The above instructions, when executed by the processor (530), cause the electronic device (501) to: Check whether the similarity between the first output information and the second output information exceeds the first threshold value, An electronic device (501) that causes a second answer to be provided that includes a similar answer to the first answer based on determining that the similarity exceeds the first threshold.

8. In any one of paragraphs 1 to 7, The above instructions, when executed by the processor (530), cause the electronic device (501) to: Based on confirming that the similarity does not exceed the first threshold, it is confirmed whether the similarity exceeds a second threshold that is less than the first threshold, An electronic device (501) that causes a third answer including a secondary answer to the first answer to be provided based on determining that the similarity exceeds the second threshold.

9. In any one of paragraphs 1 to 8, The above instructions, when executed by the processor (530), cause the electronic device (501) to: Providing the above first output information as input to at least one second model, An electronic device (501) that causes the second answer to be provided based on the second output information being output from the at least one second model.

10. In any one of paragraphs 1 to 9, Mike (511); and Including further display (515), The above instructions, when executed by the processor (530), cause the electronic device (501) to: Obtain the first user query through the above microphone (511), An electronic device (501) that causes a first answer corresponding to the first user query to be provided through the display (515).

11. In the method, The action of obtaining the first user query; An operation of inputting input information corresponding to the first user query into each of a plurality of models; An operation of providing a first answer corresponding to the first user query based on first output information output from at least one first model among the plurality of models; An operation of obtaining second output information output from at least one second model among the plurality of models; An operation of checking the similarity between the first output information and the second output information; and A method comprising an operation of providing a second answer obtained based on the first output information and the second output information based on the confirmed similarity.

12. In paragraph 11, An operation of checking whether output information corresponding to the input information is obtained from at least one generative AI model among the plurality of models before the first time interval elapses; An operation of providing the first answer based on the first output information, based on confirming that the output information is not obtained from the at least one generative AI model before the first time interval elapses; and A method further comprising an action of providing the second answer based on the second output information being obtained from the generative AI model after the first time interval has elapsed.

13. In paragraph 11 or 12, An action to check whether the above first user query corresponds to a multi-intent query; An operation of inputting input information corresponding to each of the first intent, the second intent, and the multi-intent included in the first user query into each of the plurality of models based on confirming that the first user query corresponds to the multi-intent query; An operation of providing a first answer corresponding to the first intent and the second intent based on obtaining the first output information from the first model before the first time interval elapses; and A method further comprising an action of providing a second answer corresponding to the multi-intent based on obtaining the second output information from the second model after the first time interval has elapsed.

14. In any one of paragraphs 11 to 13, A method wherein the generative AI model comprises at least one first generative AI model included in an electronic device (501) and at least one second generative AI model included in at least one external electronic device (560, 570).

15. In any one of paragraphs 11 to 14, The act of executing a first application associated with providing an answer corresponding to the query; and A method further comprising an action of providing the first answer based on obtaining the first user query while the first application is running.

Citation Information

Patent Citations

  • Question information processing method and device, equipment, storage medium and product

    CN117453885A

  • Question answering system and question answering method

    JP7088270B2

  • State determination device, method, and program

    JP7481995B2

  • Electronic fence system

    KR102427548B1

  • Search, question answering, and classifier construction

    US20170278011A1