Electronic device and utterance processing method of electronic device
The electronic device dynamically selects and changes LLMs based on utterance complexity and intent to optimize processing, addressing inefficiencies in existing technologies and enhancing accuracy and resource utilization in user speech processing.
Patent Information
- Application Number
- PCT/KR2024/096970
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-25
- Filing Date
- 2024-12-13
- Publication Date
- 2025-07-10
AI Technical Summary
Existing technologies face challenges in efficiently processing user utterances using large language models (LLMs) due to varying complexity and intent in user speech, leading to inaccuracies and resource inefficiencies.
An electronic device dynamically selects and changes LLMs based on the complexity of user utterances, intent recognition, and context information, optimizing processing levels and resource utilization through a system that includes a communication circuit, memory storing LLM databases, and processors to manage and process user utterances effectively.
This approach enhances the accuracy and efficiency of processing user utterances by selecting appropriate LLMs, optimizing resource usage, and improving the quality and speed of responses, thereby increasing user convenience and reducing costs.
Smart Images

Figure KR2024096970_10072025_PF_FP_ABST
Abstract
Description
Electronic devices and methods for handling ignition of electronic devices
[0001] The embodiments disclosed in this document relate to a technology for processing user speech based on artificial intelligence.
[0002] Electronic devices can perform actions in response to user voice commands through speech recognition and natural language processing. Recently, language models based on artificial intelligence (AI) that process user speech (or text corresponding to user speech) have become widespread. For example, a large language model (LLM) is a language model comprised of an artificial neural network with numerous parameters, and various large language models are being developed. For example, when multiple language models are available, a user can directly select a language model to process the speech from among them. For example, an electronic device can process a user's speech using the language model selected by the user.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0004] An electronic device according to one embodiment disclosed in the present document may include a communication circuit, a memory for storing a large language model (LLM) database including a plurality of LLMs and instructions, and at least one processor. According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to receive information corresponding to at least one user utterance from an external electronic device, recognize at least one intent included in the at least one user utterance and the number of intents included in the at least one user utterance, determine a processing level required to process the at least one user utterance based at least in part on the at least one intent and the number of intents, select a first LLM having a performance corresponding to the determined processing level from among the plurality of LLMs, and process the at least one user utterance using the first LLM.
[0005] In addition, a method according to an embodiment disclosed in the present document may include an operation of receiving information corresponding to at least one user utterance from an external electronic device, an operation of recognizing at least one intent included in the at least one user utterance and a number of intents included in the at least one user utterance, an operation of determining a processing level required to process the at least one user utterance based at least in part on the at least one intent and the number of intents, an operation of selecting a first LLM having a performance corresponding to the determined processing level from among a plurality of LLMs included in an LLM database, and an operation of processing the at least one user utterance using the first LLM.
[0006] In addition, a storage medium according to an embodiment disclosed in the present document may store instructions that, when executed by at least one processor of an electronic device, cause the electronic device to receive information corresponding to at least one user utterance from an external electronic device, recognize at least one intent included in the at least one user utterance and the number of intents included in the at least one user utterance, determine a processing level required to process the at least one user utterance based at least in part on the at least one intent and the number of intents, select a first LLM having a performance corresponding to the determined processing level from among a plurality of LLMs included in an LLM database, and process the at least one user utterance using the first LLM.
[0007] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.
[0008] FIG. 1 is a block diagram of an electronic device according to one embodiment.
[0009] Figure 2 is a block diagram of a system according to one embodiment.
[0010] Figures 3a to 3c illustrate an interface for setting selection criteria for a language model according to one embodiment.
[0011] FIG. 4 is a diagram for explaining an operation of a user terminal processing a user utterance according to one embodiment.
[0012] FIG. 5 is a diagram for explaining an operation of a user terminal processing a user utterance according to one embodiment.
[0013] FIG. 6 is a diagram for explaining an operation of a user terminal processing a user utterance according to one embodiment.
[0014] Fig. 7 is a flowchart of a method for processing ignition of an electronic device according to one embodiment.
[0015] Fig. 8 is a flowchart of a method for processing ignition of an electronic device according to one embodiment.
[0016] Fig. 9 is a flowchart of a method for processing ignition of an electronic device according to one embodiment.
[0017] Fig. 10 is a flowchart of a method for processing ignition of an electronic device according to one embodiment.
[0018] FIG. 11 illustrates an electronic device within a network environment according to various embodiments.
[0019] FIG. 12 is a block diagram illustrating an integrated intelligence system according to one embodiment.
[0020] FIG. 13 is a diagram showing a form in which relationship information between concepts and actions is stored in a database according to one embodiment.
[0021] FIG. 14 is a diagram illustrating a user terminal displaying a screen for processing voice input received through an intelligent app, according to one embodiment.
[0022] FIG. 15 illustrates a generative artificial intelligence system according to one embodiment.
[0023] In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0024] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.
[0025] FIG. 1 is a block diagram of an electronic device according to one embodiment.
[0026] According to one embodiment, an electronic device (100) (e.g., an electronic device (200) of FIG. 2, an electronic device (1101) of FIG. 11, a user terminal (1201) of FIG. 12, an intelligent server (1300), or a generative artificial intelligence system (1500)) includes a communication circuit (110) (e.g., a communication module (1190) of FIG. 11, a communication interface (1290) of FIG. 12, or a front end (1310)), at least one processor (130) (e.g., a processor (1120) of FIG. 11, an LLM management module (210), a state management module (250) of FIG. 2, a processor (1220) or a natural language platform (1320) of FIG. 12, or an artificial intelligence framework (1520) of FIG. 15)) or a memory (120) (e.g., an LLM database (240), a memory (260) of FIG. 2, a memory (1130) of FIG. 11, or a It may include a memory (1230) of 12, or a database (1530) of FIG. 15.
[0027] According to one embodiment, the communication circuit (110) can transmit and receive information and / or data with an external device and / or an external server (e.g., the server (1108) of FIG. 11). For example, the communication circuit (110) can receive information related to at least one user utterance and / or context information related to the external electronic device from an external electronic device (e.g., the external device (201) of FIG. 2, the electronic devices 1102 and 1104 of FIG. 11, or the user terminal (1201) of FIG. 12). For example, the communication circuit (110) can transmit a result of processing a user utterance (e.g., a response of a large language model (LLM) corresponding to the user utterance) to the external electronic device. In the present disclosure, processing of a 'user utterance' is described, but is not limited thereto, and the target of processing using the LLM can include a user utterance, information corresponding to the user utterance (text information corresponding to the user utterance), and / or a user's text input.
[0028] According to one embodiment, the memory (120) may store instructions that control the operation of the electronic device (100) when executed by the processor (130). According to one embodiment, the instructions may be stored in one memory (120) or multiple memories (120). According to one embodiment, the memory (120) may include an LLM database including multiple LLMs. According to one embodiment, the LLM database is described as being included in the memory (120) of the electronic device (100), but is not limited thereto, and the LLM database may be located outside the electronic device (100), or multiple LLM databases may be located outside the electronic device (100) and the electronic device (100).
[0029] According to one embodiment, the processor (130) may receive information corresponding to at least one user utterance from an external electronic device via the communication circuit (110). According to one embodiment, when the electronic device (100) includes a microphone (not shown), the processor (130) may obtain information corresponding to the user utterance received via the microphone.
[0030] According to one embodiment, the processor (130) can recognize at least one intent and the number of intents included in at least one user utterance. For example, the user utterance may include at least one intent. For example, the at least one user utterance may include a single user utterance, a plurality of consecutive user utterances, and / or a plurality of related user utterances received at a predetermined time interval.
[0031] According to one embodiment, the processor (130) may determine a level of processing required to process at least one user utterance based at least in part on at least one intent and / or the number of intents. For example, the processor (130) may determine the complexity of the user utterance based on at least one intent and / or the number of intents. For example, the complexity of the user utterance may be related to the number of processing steps required to process the user utterance. For example, the level of processing may be determined based on at least one of the number of processing steps required to process the user utterance, the processing speed, the amount of data that can be processed, the amount of computation, the type of input data, the learning inference capability, the number of plug-ins available to the LLM, context information related to external electronic devices, and / or a combination thereof.
[0032] According to one embodiment, the processor (130) may receive contextual information related to the external electronic device from the external electronic device. The processor (130) may determine the processing level based at least in part on the contextual information. For example, the contextual information may include a user profile, the number of available plug-ins, personal information stored in the external electronic device (e.g., schedule information), and / or information related to the status of the external electronic device.
[0033] According to one embodiment, the processor (130) may obtain input data related to at least one user utterance. The processor (130) may determine the processing level based at least in part on the type of the input data. For example, the input data may include data provided to the LLM (e.g., document files, image files, audio files, video files, and / or combinations thereof). For example, the processor (130) may determine the processing level based at least in part on whether the type of the input data is text, audio, image, video, and / or a combination thereof.
[0034] According to one embodiment, the processor (130) may use the trained artificial intelligence model to predict at least one of the number of processing steps, processing time, processing speed, computational amount, or a combination thereof required to process at least one user utterance. The processor (130) may determine a processing level based at least in part on the prediction result. For example, the artificial intelligence model may be trained using at least one of a language model, a deep learning model, an artificial neural network, a deep neural network, and / or a combination thereof. For example, the artificial intelligence model may be trained based on previously processed user utterances, a processing history of user utterances, and / or summary information about an operation that processed the user utterance.
[0035] According to one embodiment, the processor (130) may receive user-preference values related to selection criteria of an LLM from an external electronic device (e.g., a device that receives a user utterance). For example, the user-preference values may be set based on user input received through an interface provided by the external electronic device. For example, the selection criteria of the LLM may include performance of the LLM (e.g., maximum capacity, size, processing speed, number of parameters, number of plug-ins supported, and / or specialized field of the LLM) and / or whether an automatic change function of the LLM is activated. The processor (130) may select a first LLM from among a plurality of LLMs based at least in part on the user-preference values.
[0036] In one embodiment, the processor (130) may select a first large language model (LLM) having performance corresponding to a determined processing level from among a plurality of LLMs. For example, the processor (130) may select the first LLM from an LLM database containing a plurality of LLMs.
[0037] According to one embodiment, the processor (130) may select at least one category among a plurality of LLM categories based on at least one intent included in a user utterance. The processor (130) may select a first LLM among the LLMs belonging to at least one selected category based on the number of intents included in the user utterance.
[0038] According to one embodiment, the processor (130) may select a first LLM based on a result of comparing the determined processing level with a preset threshold. For example, there may be one or more thresholds. The thresholds may be set to divide the processing level into a plurality of ranges. For example, each of the plurality of ranges may correspond to a performance of the LLM. For example, the processor (130) may select one of the LLMs having a performance corresponding to the processing level from among the plurality of LLMs based on a result of comparing the processing level with the threshold.
[0039] In one embodiment, the processor (130) may process at least one user utterance using the first LLM. For example, the processor (130) may query the first LLM for the user utterance. For example, the processor (130) may provide the first LLM with a prompt corresponding to the user utterance and obtain a response to the prompt from the first LLM.
[0040] According to one embodiment, the processor (130) may select a second LLM corresponding to the processing state among the plurality of LLMs based on recognizing that the processing state of at least one user utterance does not correspond to a determined processing level while processing at least one user utterance using the first LLM. For example, the processor (130) may recognize that the current processing state of the user utterance (e.g., the number of processing steps, the type of input data, the processing speed, context information related to an external electronic device, the intent of the user utterance, and / or the number of intents included in the user utterance) does not correspond to the determined processing level. For example, the processor (130) may recognize that a new user utterance is additionally input (the intents included in the user utterance may be changed / added or the number of intents may be changed depending on the additional user utterance), or that a processing step while processing the user utterance exceeds a pre-predicted processing level, or that more input data is required / provided, or that an LLM with higher performance or an LLM with different performance is required to process the user utterance (or to increase the accuracy of the user utterance processing result). For example, the processor (130) may select a second LLM that is more suitable for the processing state of the user utterance if the processing state of the user utterance does not correspond to the determined processing level. The processor (130) may change the LLM used for processing the user utterance from the first LLM to the second LLM.
[0041] According to one embodiment, the processor (130) may generate summary information summarizing the content of processing at least one user utterance using the first LLM. For example, the processor (130) may generate the summary information by extracting a portion of information corresponding to each processing step for processing the user utterance. For example, the summary information may include at least a portion of information corresponding to the user utterance and / or information obtained using the first LLM. The summary information may include a smaller number of tokens than the number of tokens included in the information input to the first LLM and / or information obtained (output) from the first LLM. For example, the processor (130) may tokenize and manage the information input to the first LLM and / or information obtained (output) from the first LLM. For example, the processor (130) may generate the summary information including a portion of the tokens obtained from the information input to the LLM and / or information obtained (output) from the first LLM.
[0042] In one embodiment, the processor (130) may provide summary information to the second LLM. For example, the processor (130) may provide the summary information as input to the second LLM (e.g., as a prompt for the second LLM).
[0043] In one embodiment, the processor (130) may process at least one user utterance using a second LLM. For example, the processor (130) may generate a prompt related to an intent contained in the at least one user utterance based on information corresponding to the at least one user utterance and provide the prompt to the second LLM. The processor (130) may obtain a response to the prompt from the second LLM.
[0044] According to one embodiment, the processor (130) may transmit the result of processing the user speech using the LLM (e.g., the first LLM and / or the second LLM) to an external electronic device (e.g., a device that received the user speech) via the communication circuit (110).
[0045] According to one embodiment, the operations described as being performed by the processor (130) may be performed by at least one processor (130). For example, each of the operations described above may be performed by the same or different processor (130), and / or may be performed by at least some of a plurality of processors (130). For example, at least some of the operations described above may be performed by a first processor (130), and at least some of the remaining operations may be performed by a second processor (130). According to one embodiment, the processor (130) may include a circuit such as a central processing unit (CPU), a microprocessor unit (MPU), an application processor (AP), a communication processor (CP), a system on chip (SoC), and / or an integrated circuit (IC).
[0046] According to various embodiments, the configuration of the electronic device (100) is not limited to that described in FIG. 1, and at least some components may be omitted or at least some components (for example, at least one of the components of the electronic device (200) of FIG. 2, the components of the electronic device (1101) of FIG. 11, or the components of the user terminal (1201) and / or the intelligent server (1300) of FIG. 12) may be added. According to one embodiment, although the electronic device (100) is described as receiving information related to a user's utterance and / or context information related to the external electronic device from an external electronic device, it is not limited thereto, and the electronic device (100) may also obtain information related to a user's utterance and / or context information related to the electronic device. For example, if the electronic device (100) is implemented as an on-device rather than a server device, the electronic device (100) can receive user speech through a microphone and / or recognize context information of the electronic device (100).
[0047] According to one embodiment, the electronic device (100) determines the processing level required to process a user utterance, and based on this, selects an LLM having performance suitable for processing the user utterance from among a plurality of LLMs, thereby increasing the accuracy of utterance processing, optimizing service costs, and increasing the efficiency of utterance processing. For example, the electronic device (100) selects an LLM having a corresponding scale based on the complexity of the user utterance to process the utterance, thereby increasing the accuracy and quality of the response to the utterance, and optimizing the resources and costs consumed for utterance processing. For example, the electronic device (100) can increase the efficiency (e.g., token efficiency) of user utterance processing using LLMs by storing and utilizing summary information that summarizes information on the operation of processing the user utterance.
[0048]
[0049] Figure 2 is a block diagram of a system according to one embodiment.
[0050] According to one embodiment, the system may include an electronic device (200) (e.g., the electronic device (100) of FIG. 1, the electronic device (200) of FIG. 2, the electronic device (1101) of FIG. 11, the user terminal (1201) or intelligent server (1300) of FIG. 12, or the generative artificial intelligence system (1500) of FIG. 15) and an external device (201) (e.g., the external device (201) of FIG. 2, the electronic devices (1102, 1104) of FIG. 11, or the user terminal (1201) of FIG. 12). For example, although the system is described as including a plurality of devices, it is not limited thereto, and a system for processing user speech may be implemented as a single device (e.g., the electronic device (200)).
[0051] In one embodiment, the external device (201) may include a voice assistant (2011). In one embodiment, the external device (201) may receive a user utterance. For example, the voice assistant (2011) may analyze the user utterance and perform an action corresponding to the user utterance. In one embodiment, the voice assistant (2011) may include LLM settings (2015) and context information (2017). For example, the LLM settings (2015) may include user-preference values related to selection criteria of the LLM. The voice assistant (2011) may receive the user-preference values related to the selection criteria of the LLM from the user through an interface. For example, the selection criteria of the LLM may include performance of the LLM (e.g., maximum capacity, size, processing speed, number of parameters, number of plug-ins supported, and / or specialized fields of the LLM) and / or whether an automatic change function of the LLM is enabled. For example, context information (2017) may include information related to a user profile, the number of available plug-ins, personal information stored on the external device (201) (e.g., schedule information), and / or the status of the external device (201).
[0052] An external device (201) can provide LLM settings (2015) and context information (2017) to an electronic device (200) through a voice assistant (2011). The external device (201) can provide information related to a user's speech to the electronic device (200). The external device (201) can receive the result of processing the user's speech from the electronic device (200) and provide the result of processing the user's speech through a voice assistant (2011).
[0053] According to one embodiment, the electronic device (200) may include an LLM management module (210), an LLM database (240), a state management module (250), and a memory (260).
[0054] In one embodiment, the LLM management module (210) can dynamically select or change an LLM suitable for processing the user utterance from the LLM database (240) based at least in part on the user utterance. In one embodiment, the LLM management module (210) can include a preprocessing module (220) and an execution module (230).
[0055] According to one embodiment, the preprocessing module (220) may generate and / or extract information used to select an LLM suitable for processing the user utterance based on input information (e.g., information corresponding to the user utterance, information related to the input data, a threshold value (261), an LLM setting (2015) (a user-set value), and / or context information (2017)). According to one embodiment, the preprocessing module (220) may include an intent analysis module (221), a processing level determination module (223), a threshold value (261) management module (225), and an AI training module (227).
[0056] According to one embodiment, the intent analysis module (221) may include an intent counting module (2211) and an intent understanding module (2213). For example, the intent counting module (2211) may recognize the number of intents included in at least one user utterance. For example, the intent understanding module (2213) may analyze at least one user utterance to recognize (and / or understand) the intent included in the user utterance. According to one embodiment, the intent understanding module (2213) and / or the intent counting module (2211) may include at least one language model (e.g., LLM), a deep learning model, a prompt engineering model, and / or a deep neural network model.
[0057] According to one embodiment, the processing level determination module (223) can determine the complexity of a user utterance based on the number of intentions included in the user utterance.
[0058] For example, the processing level determination module (223) may determine the selection criteria for LLM based on the result of comparing the complexity of the user utterance (e.g., the number of intents included in the user utterance) with the first threshold values, if the AI model (2231) (e.g., the generative artificial intelligence model (1550) of FIG. 15) has not been learned (trained). The processing level determination module (223) may provide information on the selection criteria for LLM to the execution module (230) (e.g., the LLM selection module (231)).
[0059] For example, assume that the number of intents contained in a user utterance is x, and the first thresholds are 2 and 4.
[0060] Comparison of the number of intentions and the first threshold Selection criteria for LLM (performance of LLM (e.g. LLM size)) x < 2 Small size (e.g. 1B) 2 ≤ x < 4 Medium size (e.g. 13B) 4 ≤ x Large size (e.g. 70B)
[0061] Referring to Table 1, the processing level decision module (223) may determine to select an LLM with relatively low performance (e.g., relatively small size) when the number of intents is less than 2. The processing level decision module (223) may determine to select an LLM with medium performance (e.g., medium size) when the number of intents is 2 or more but less than 4. The processing level decision module (223) may determine to select an LLM with relatively high performance (e.g., relatively large size) when the number of intents is 4 or more. For example, the contents shown in Table 1 are only examples and are not limited thereto. For example, the performance of the LLM may be a factor other than the size, and the size of the LLM may be set to a specific value (e.g., 1B, 13B, 70B) rather than a relative size.
[0062] According to one embodiment, the processing level determination module (223) may determine a category of LLM to select based on the intent contained in the user's utterance. The processing level determination module (223) may provide the determined category information to the execution module (230) (e.g., the LLM selection module (231)) to select one of the LLMs belonging to the selected category.
[0063] Intent LLM Category LLM Settings (2015) (e.g., whether to enable automatic changes) Coding Fine-tuned LLM (243) No Device Control Third-Party LLM (241) Yes Chat Fine-tuned LLM (243) Yes User Specific PEFT LLM (245) Yes
[0064] Referring to Table 2, the processing level determination module (223) may determine to select an LLM of the fine tuned LLM category if the intent included in the user utterance is related to coding or chatting, to select an LLM of the third-party provided LLM (241) (e.g., an LLM provided by a specific service provider) category if the intent is related to device control, and to select an LLM of the PEFT LLM (245) category if the intent corresponds to a user-specified intent. In Table 2, the LLM setting (2015) (e.g., whether to enable automatic change) may indicate whether it corresponds to a user-set value related to the selection criteria of the LLM. For example, the processing level determination module (223) may determine the category of the LLM to be selected and / or the selection criteria of the LLM according to the user-set value, and / or determine whether the electronic device (200) activates or deactivates a function of dynamically changing the selected LLM. For example, the contents described in Table 2 are examples and are not limited thereto. For example, the categories or classification criteria of LLM may change.
[0065] According to one embodiment, the processing level determination module (223) may determine the processing level (e.g., the number of processing steps) required to process a user utterance. For example, the processing level determination module (223) may include an AI model (2231) used to determine the processing level. For example, the AI model (2231) may include a deep neural network model that predicts the processing level (e.g., the number of processing steps) required to process a user utterance based on information extracted from the user utterance (e.g., the intent and / or the number of intents included in the user utterance), LLM settings (2015) (user setting values) received from an external device (201), and / or context information (2017).
[0066] For example, with regard to Table 3 below, it is assumed that the processing level determination module (223) predicts the number of processing steps required to process a user utterance based on an artificial intelligence model, and sets an LLM selection criterion based on the number of processing steps. For example, the processing level determination module (223) can predict the number of processing steps required to process a user utterance using an artificial intelligence model, and determine an LLM selection criterion based on the result of comparing the number of processing steps with at least one second threshold value.
[0067] Processing Level Comparison with the Second Threshold Selection Criteria for LLM (Performance of LLM (e.g. Size of LLM))Number of Processing Steps (y)y < 3Small size (e.g. 1B)Number of Processing Steps (y)3 ≤ y < 5Medium size (e.g. 13B)Number of Processing Steps (y)5 ≤ yLarge size (e.g. 70B)
[0068] Referring to Table 3, the processing level determination module (223) may determine to select an LLM with relatively low performance (e.g., relatively small size) when the predicted number of processing steps is less than 3. The processing level determination module (223) may determine to select an LLM with medium performance (e.g., medium size) when the predicted number of processing steps is 3 or more but less than 5. The processing level determination module (223) may determine to select an LLM with relatively high performance (e.g., relatively large size) when the predicted number of processing steps is 5 or more. The processing level determination module (223) may change the selection criteria for the LLM based on recognizing that the processing status of the user utterance does not correspond to the predetermined processing level while processing the user utterance. For example, when the predicted number of processing steps based on the artificial intelligence model is 4, the processing level determination module (223) may determine to select an LLM with a medium size. For example, the current processing step while processing the user utterance may exceed the predicted number of processing steps of 4. In this case, the processing level determination module (223) may recognize that the determined processing level does not correspond to the current processing status of the user utterance. For example, the processing level determination module (223) may recognize that a higher performance LLM needs to be used to process the user utterance. The processing level determination module (223) may change the selection criteria of the LLM to select an LLM corresponding to the current processing status based on the second threshold values. For example, the electronic device (200) may change the selection criteria of the LLM to select an LLM having relatively high performance (e.g., relatively large size) and provide information about the changed selection criteria of the LLM to the execution module (230) (e.g., the LLM selection module (231)).For example, the processing level determination module (223) can cause the execution module (230) (e.g., the LLM selection module (231)) to change the previously selected (currently in use) LLM to an LLM suitable for the current processing status of the user's speech.
[0069] Although Table 3 assumes that the processing level is the number of processing steps, this is not limiting. For example, the processing level may be determined based on various factors related to LLM performance, including processing speed and / or computational load, in addition to the number of processing steps, and the second threshold value may be changed to a different value.
[0070] According to one embodiment, the threshold value (261) management module (225) can extract a threshold value (261) necessary for determining a processing level from among the threshold values (261) stored in the memory (260) and provide the extracted threshold value to the processing level determination module (223).
[0071] In one embodiment, the AI training module (227) may train an AI model (2231) used to determine a processing level. For example, the AI training module (227) may update the AI model (2231) by learning information extracted from a user utterance (e.g., intents and / or the number of intents included in the user utterance), LLM settings (2015) (user-set values) received from an external device (201), and / or context information (2017).
[0072] In one embodiment, the execution module (230) may select an LLM suitable for processing a user utterance from the LLM database (240) and process the user utterance using the selected LLM. For example, the execution module (230) may input information corresponding to the user utterance into the selected LLM and obtain a response to the user utterance from the selected LLM. In one embodiment, the execution module (230) may include an LLM selection module (231) and a summary module (233).
[0073] According to one embodiment, the LLM selection module (231) may select an LLM to process the user utterance from the LLM database (240) based on information related to the user utterance (e.g., intent and / or number of intents included in the user utterance) and / or processing level. For example, the LLM selection module (231) may select an LLM having performance (e.g., maximum capacity, size, processing speed, number of parameters, number of plug-ins supported, and / or specialized field of the LLM) corresponding to the information received from the preprocessing module (220) based on information received from the preprocessing module (220) (e.g., intent included in the user utterance, number of intents included in the user utterance, and / or processing level). For example, the LLM selection module (231) may process the user utterance using the selected LLM.
[0074] According to one embodiment, the summary module (233) may generate summary information summarizing the content of at least one user utterance processed. For example, the summary module (233) may generate the summary information by extracting a portion of information corresponding to each processing step of processing the user utterance. For example, the summary information may include at least a portion of the information corresponding to the user utterance and / or the information obtained using the first LLM. The summary information may include fewer tokens than the number of tokens included in the information input to the first LLM and / or the information obtained (output) from the first LLM.
[0075] According to one embodiment, at least some of the components of the LLM management module (210) and the state management module (250) may be implemented as at least one hardware component (e.g., at least one processor (130) of FIG. 1 or processor (1120) of FIG. 11).
[0076] According to one embodiment, the LLM database (240) can store a plurality of LLMs. For example, the LLM database (240) can store a plurality of LLMs by categorized categories. For example, the categories of the LLMs can be classified according to the LLM provider (e.g., a third-party provided LLM), the tuning method (e.g., a fine-tuned LLM (243) or a parameter efficiency fine-tuned (PEFT) LLM), but are not limited thereto. For example, the LLMs can be categorized and stored by specialized fields. According to one embodiment, the LLM database (240) can be located within and / or outside the electronic device (200). For example, at least some of the plurality of LLMs can be stored within the electronic device (200), and the remaining some can be stored outside the electronic device (200). According to one embodiment, the LLM database (240) and memory (260) may be implemented as separate storage devices, or may be implemented in different areas within one storage device (e.g., memory (120) of FIG. 1 or memory (1130) of the eleventh embodiment).
[0077] According to one embodiment, the state management module (250) may include a log recording module (251) and an update module (253).
[0078] According to one embodiment, the log recording module (251) may store the history of processing user utterances using the selected LLM in the execution log (263). For example, the log recording module (251) may store summary information generated by the summary module (233) in the execution log (263).
[0079] In one embodiment, the update module (253) may update at least one threshold value (261) based on the history and / or summary information of processing the user utterance using the selected LLM. For example, the update module (253) may change at least one of the previously stored threshold values (261) to a different value based on the history and / or summary information of processing the user utterance.
[0080] In one embodiment, the memory (260) may store at least one threshold value (261) and an execution log (263). For example, the memory (260) may store at least one threshold value (261) used to determine the level of processing required to process a user utterance. For example, the execution log (263) may include a history and / or summary information of processing a user utterance using LLM.
[0081] According to various embodiments, the configuration of the electronic device (200) is not limited to that described in FIG. 2, and at least some components may be omitted or at least some components (for example, at least one of the components of the electronic device (100) of FIG. 1, the components of the electronic device (1101) of FIG. 11, or the components of the user terminal (1201) and / or the intelligent server (1300) of FIG. 12) may be added. According to one embodiment, although the electronic device (100) is described as receiving information related to user utterance and / or context information related to the external device (201) from the external device (201), it is not limited thereto, and the electronic device (200) may directly obtain information related to user utterance and / or context information related to the electronic device (200). For example, the system may be implemented in the form of an on-device in which the electronic device (200) and the external device (201) are integrated.
[0082] An electronic device (200) according to an embodiment of the present disclosure can improve the accuracy and reliability of the results of processing user speech using LLMs and increase user convenience by dynamically selecting or changing an LLM having functions suitable for processing user speech among a plurality of LLMs. The electronic device (200) can increase the efficiency (e.g., token efficiency) of processing user speech by summarizing, managing, and utilizing the contents of processed user speech.
[0083]
[0084] Figures 3a to 3c illustrate interfaces for setting selection criteria for a language model according to one embodiment. For example, when interfaces (310, 320, 330) are provided by an external electronic device (e.g., an external device (201) of Figure 2, an electronic device (1102, 1104) of Figure 11, or a user terminal (1201) of Figure 12), the electronic device (e.g., an electronic device (100) of Figure 1, an electronic device (200) of Figure 2, an electronic device (1101) of Figure 11, a user terminal (1201) of Figure 12, an intelligent server (1300), or a generative artificial intelligence system (1500)) can receive user setting values set from the external electronic device through the interfaces (310, 320, 330). For example, if an interface (310, 320, 330) is provided in an electronic device, the electronic device can set user-defined values related to selection criteria of the language model based on user input received through the interface (310, 320, 330).
[0085] Referring to FIG. 3A, the first interface (310) may include a first area (311) for setting the maximum capacity of the LLM and a second area (313) for activating the automatic change function of the LLM. For example, the electronic device may set the maximum capacity of the LLM based on a user input (319) received through the first area (311) of the first interface (310). The electronic device may activate or deactivate the automatic change function of the LLM based on a user input (319) received through the second area (313) of the first interface (310). For example, the automatic change function of the LLM may be a function that allows the electronic device to dynamically select and change an LLM suitable for processing a user utterance while processing the user utterance. For example, when the automatic change function of the LLM is activated, the electronic device may select or change an LLM suitable for processing a user utterance from among a plurality of LLMs without a separate user input while processing the user utterance.
[0086] Referring to FIG. 3b, the second interface (320) may include a third area (321) for setting the size condition of the LLM. For example, the electronic device may determine the size of the LLM to be used based on a user input (329) received through the third area (321) of the second interface (320). For example, FIG. 3b illustrates a case where there are three sizes of LLM that can be set (e.g., small LLM, medium LLM, and large LLM), but embodiments of the present disclosure are not limited thereto.
[0087] Referring to FIG. 3c, the third interface (330) may include a fourth area (331) for setting category conditions of the LLM. For example, the categories of the LLM may represent LLMs specialized in specific fields. For example, the categories of the LLM may be related to the intent included in the user utterance (e.g., movies, stocks, coding). For example, the names of the classification criteria displayed in the interface (e.g., “enabled expert model”) and the category names (e.g., “movie expert,” “stock expert,” and “coding expert”) are examples, and the names indicating the classification criteria, categories, and / or LLMs in the interface may be changed or specified by user input (339). For example, FIG. 3c illustrates that the categories of the LLM are classified according to specialized fields (e.g., movies, stocks, coding), but this is not limited thereto, and the number of categories is not limited to that illustrated in FIG. 3c. For example, the category of the LLM may be classified based on the provider of the LLM (e.g., third-party provided LLM) and / or the tuning method (e.g., fine-tuned LLM, parameter efficient fine tuning (PEFT) LLM). For example, the electronic device may determine the category of the LLM to be selected based on user input (339) received through the fourth area (331) of the third interface (330).
[0088] According to various embodiments, the interface for setting selection criteria for a language model is not limited to that illustrated and described in FIGS. 3a to 3c, and the form of the selection criteria and / or interface may be changed.
[0089]
[0090] FIG. 4 is a diagram illustrating an operation of a user terminal processing a user utterance according to one embodiment. For example, FIG. 4 illustrates an example of processing a user utterance through a voice assistant function of a user terminal (e.g., an electronic device (100) of FIG. 1 , an external device (201) of FIG. 2 , an electronic device (1102, 1104) of FIG. 11 , a user terminal (1201) of FIG. 12 , or a generative artificial intelligence system (1500)), and illustrates interface screens (410, 430) related to the voice assistant function.
[0091] According to one embodiment, a user terminal may input a query (411) into an LLM based on information corresponding to a user utterance. For example, the user terminal may generate a prompt (411) corresponding to the user utterance and input the prompt (411) into the LLM. For example, the user terminal may determine a processing level required to process the user utterance and select a first LLM from among a plurality of LLMs based on the determined processing level. The user terminal may process the user utterance using the first LLM. For example, if the user terminal receives the first user utterance “How is the weather today?” from a user (e.g., Jane), the user terminal may select the first LLM that can answer (respond) to the query about today’s weather from among the plurality of LLMs based on information corresponding to the first user utterance. The user terminal may input a query (411) corresponding to the user utterance into the first LLM. The user terminal can obtain an answer (response) (413) to a query (i.e., first user utterance) (411) from the first LLM.
[0092] Thereafter, the user terminal may receive a second user utterance related to schedule management. The user terminal may input a query (415) corresponding to the second user utterance to the first LLM. The user terminal may determine whether the second user utterance can be processed using the selected first LLM. For example, the number of intents included in user utterances (e.g., the first user utterance and the second user utterance) may increase due to the second user utterance. For example, the intent included in the first user utterance may be 'weather', and the intent included in the second user utterance may be 'schedule setting'. If the user terminal cannot process the second user utterance using the first LLM (e.g., if the accuracy of the answer is expected to be low when the second user utterance is processed using the first LLM), the user may prompt the user to change the LLM (431). For example, if the user terminal determines that an LLM with higher performance (e.g., larger size) is required due to an increase in the number of intents included in user utterances, the user terminal may prompt the user to change the LLM (431).
[0093] For example, if a user input (433) agreeing to a change in LLM is received from the user, the user terminal may change the LLM used to process the second user utterance from the first LLM to a second LLM that can answer (respond) to a query (415) about schedule settings. The user terminal may provide a result of processing the second user utterance using the second LLM (e.g., an answer (response) to the second user utterance) (435, 437).
[0094] In FIG. 4, it is described that the user terminal receives user utterances, selects an LLM based on the user utterances, and processes the user utterances using the selected LLM. However, the user terminal may perform at least some operations in conjunction with an external server (e.g., the electronic device (100) of FIG. 1, the electronic device (200) of FIG. 2, the electronic device (1101) of FIG. 11, the user terminal (1201) of FIG. 12, or the intelligent server (1300)). For example, the user terminal may transmit information corresponding to the user utterance to the external server, the external server may select an LLM suitable for the user utterance, and perform an operation of processing the user utterance using the selected LLM. The user terminal may receive the result of processing the user utterance from the external server and provide the result to the user.
[0095] According to one embodiment, a user terminal (and / or an external server) can increase the processing speed, processing performance, and / or accuracy of a user utterance by dynamically selecting or changing an LLM suitable for processing the user utterance among a plurality of LLMs based at least in part on the user utterance.
[0096]
[0097] FIG. 5 is a diagram illustrating an operation of a user terminal processing a user utterance according to one embodiment. For example, FIG. 5 illustrates interface screens (510, 530) related to a voice assistant function as an example of processing a user utterance through a voice assistant function of a user terminal (e.g., an electronic device (100) of FIG. 1 , an external device (201) of FIG. 2 , an electronic device (1102, 1104) of FIG. 11 , a user terminal (1201) of FIG. 12 , or a generative artificial intelligence system (1500)). Hereinafter, any description overlapping with that of FIG. 4 will be omitted or briefly described.
[0098] In one embodiment, a user terminal may input a query (511) to an LLM based on information corresponding to a user utterance. The user terminal may determine a processing level required to process the user utterance and select a first LLM from among a plurality of LLMs based on the determined processing level. The user terminal may process the user utterance using the first LLM. For example, the user terminal may input a query (511) such as "Analyze the domestic bottled water market size and growth rate, and famous brands" to the first LLM using a voice assistant function.
[0099] For example, if the user terminal cannot process the user utterance using the first LLM (e.g., if the accuracy of the response is expected to be low when the user utterance is processed using the first LLM), the user terminal may prompt the user to change the LLM (513). For example, if the user terminal determines that an LLM with higher performance (e.g., larger size) is required based on the intent contained in the user utterance, the user terminal may prompt the user to change the LLM (513).
[0100] For example, if an input (515) agreeing to a change in LLM is received from a user, the electronic device may change the LLM used to process the user utterance from the first LLM to the second LLM. The electronic device may provide a result of processing the user utterance using the second LLM (e.g., a response to the second user utterance) (531).
[0101] In FIG. 5, it is described that the user terminal receives user utterances, selects an LLM based on the user utterances, and processes the user utterances using the selected LLM. However, the user terminal may perform at least some operations in conjunction with an external server (e.g., the electronic device (100) of FIG. 1, the electronic device (200) of FIG. 2, the electronic device (1101) of FIG. 11, the user terminal (1201) of FIG. 12, or the intelligent server (1300)). For example, the user terminal may transmit information corresponding to the user utterance to the external server, the external server may select an LLM suitable for the user utterance, and perform an operation of processing the user utterance using the selected LLM. The user terminal may receive the result of processing the user utterance from the external server and provide the result to the user.
[0102] According to one embodiment, a user terminal (and / or an external server) can increase the processing speed, processing performance, and / or accuracy of a user utterance by dynamically selecting or changing an LLM suitable for processing the user utterance among a plurality of LLMs based at least in part on the user utterance.
[0103]
[0104] FIG. 6 is a diagram illustrating an operation of an electronic device processing a user utterance according to one embodiment. For example, FIG. 6 illustrates an interface screen related to a voice assistant function as an example of processing a user utterance through a voice assistant function of a user terminal (e.g., an electronic device (100) of FIG. 1 , an external device (201) of FIG. 2 , an electronic device (1102, 1104) of FIG. 11 , a user terminal (1201) of FIG. 12 , or a generative artificial intelligence system (1500)). Hereinafter, any description overlapping with that of FIGS. 4 and 5 will be omitted or briefly described.
[0105] According to one embodiment, the user terminal may provide a query (611) and input data (613) related to the user utterance to a pre-selected LLM (e.g., a first LLM) based on information corresponding to the user utterance. The user terminal may determine a processing level required to process the user utterance and recognize one of a plurality of LLMs based on the determined processing level. For example, the user terminal may determine the processing level based on information (611) corresponding to the user utterance and / or information related to the input data (613) (e.g., a volume of the input data or a type of the input data (e.g., text, image, audio, video, and / or a combination thereof). The user terminal may recognize an LLM (e.g., a second LLM) corresponding to the processing level.
[0106] For example, if the user terminal cannot process the user utterance using the first LLM selected previously, the user terminal may prompt the user to change the LLM (615). For example, if the user terminal determines that an LLM with higher performance (e.g., larger size) is required based on the intent contained in the user utterance, the user terminal may prompt the user to change the LLM (615).
[0107] For example, if a user input (617) agreeing to a change in LLM is received from the user, the electronic device may change the LLM used to process the user utterance from the first LLM to the second LLM. The electronic device may provide information (631) indicating that the user utterance is being processed using the second LLM and / or a result of processing the user utterance using the second LLM (e.g., a response to the second user utterance) (633).
[0108] Although FIG. 6 describes that the user terminal receives user utterances, selects an LLM based on the user utterances, and processes the user utterances using the selected LLM, the user terminal may perform at least some operations in conjunction with an external server (e.g., the electronic device (100) of FIG. 1, the electronic device (200) of FIG. 2, the electronic device (1101) of FIG. 11, the user terminal (1201) of FIG. 12, or the intelligent server (1300)). For example, the user terminal may transmit information corresponding to the user utterance to the external server, the external server may select an LLM suitable for the user utterance, and perform an operation of processing the user utterance using the selected LLM, and the user terminal may receive the result of processing the user utterance from the external server and provide the result to the user.
[0109] According to one embodiment, a user terminal (and / or an external server) can increase the processing speed, processing performance, and / or accuracy of a user utterance by dynamically selecting or changing an LLM suitable for processing the user utterance among a plurality of LLMs based at least in part on the user utterance.
[0110]
[0111] For example, if a user directly selects a language model to process a utterance, the response to the user's utterance (i.e., the result of processing the user's utterance) may vary depending on the selected language model. For example, if the selected language model is insufficient to process the user's utterance, or if the appropriate language model is not selected based on the user's utterance, the utterance processing results may be inaccurate, and the user may not obtain the desired information.
[0112] An electronic device according to one embodiment may include a communication circuit, a memory storing instructions and a large language model (LLM) database including a plurality of LLMs, and an LLM database, and at least one processor.
[0113] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to receive information corresponding to at least one user utterance from an external electronic device.
[0114] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to recognize at least one intent included in the at least one user utterance and a number of intents included in the at least one user utterance.
[0115] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to determine a processing level (processing difficulty, processing grade, complexity) required to process the at least one user utterance based at least in part on the at least one intent and the number of intents.
[0116] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to select a first LLM from among the plurality of LLMs having a performance corresponding to the determined processing level.
[0117] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to process the at least one user utterance using the first LLM.
[0118] According to one embodiment, an electronic device determines the level of processing required to process a user utterance and, based on this, selects an LLM with performance suitable for processing the user utterance from among a plurality of LLMs, thereby enhancing the accuracy of utterance processing, optimizing service costs, and increasing the efficiency of utterance processing. For example, the electronic device may select an LLM with a corresponding scale based on the complexity of the user utterance to process the utterance, thereby enhancing the accuracy and quality of responses to the utterance and optimizing the resources and costs consumed in utterance processing.
[0119] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to receive context information related to the external electronic device from the external electronic device.
[0120] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to determine the processing level based at least in part on the context information.
[0121] According to one embodiment, the electronic device can determine the level of processing required to process a user utterance based on contextual information related to the user and / or the user's external electronic device, and determine an LLM suitable for processing the user utterance from among a plurality of LLMs.
[0122] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to obtain input data related to the at least one user utterance.
[0123] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to determine the processing level based at least in part on the type of the input data.
[0124] According to one embodiment, the type of the input data may include at least one of text, image, video, or a combination thereof.
[0125] According to one embodiment, the electronic device can determine a level of processing required to process a user utterance based on a type of input data, and determine an LLM suitable for processing the user utterance from among a plurality of LLMs.
[0126] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to predict, using the trained artificial intelligence model, at least one of a number of processing steps, a processing time, a processing speed, an amount of computation, or a combination thereof, required to process the at least one user utterance using at least one LLM of the plurality of LLMs.
[0127] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to determine the processing level based at least in part on the prediction result.
[0128] In one embodiment, the electronic device can determine an LLM suitable for processing a user utterance without additional input from the user by determining the level of processing required to process the user utterance based on a trained artificial intelligence model.
[0129] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to store a processing history of the at least one user utterance processed using the first LLM.
[0130] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to train and / or update an artificial intelligence model used to determine the processing level based on the processing history.
[0131] According to one embodiment, the plurality of LLMs may be classified into a plurality of designated categories and stored in the LLM database.
[0132] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to select at least one category from the plurality of categories based on the at least one intention.
[0133] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to select the first LLM from among the LLMs belonging to the at least one selected category based on the number of intents.
[0134] In one embodiment, the electronic device may select an LLM of a category suitable for processing a user utterance based on the intent of the user utterance.
[0135] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to select the first LLM based on a result of comparing the processing level with a preset threshold.
[0136] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to select a second LLM corresponding to the processing state among the plurality of LLMs based on recognizing that a processing state of the at least one user utterance does not correspond to the determined processing level while processing the at least one user utterance using the first LLM.
[0137] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to process the at least one user utterance using the second LLM.
[0138] According to one embodiment, the electronic device can increase the accuracy and quality of the user utterance processing result (response to the user utterance) by dynamically selecting and / or changing an LLM suitable for processing the user utterance during the user utterance processing.
[0139] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to generate summary information summarizing the content of the at least one user utterance processed using the first LLM.
[0140] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to provide the summary information to the second LLM.
[0141] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to process the at least one user utterance based on the summary information using the second LLM.
[0142] According to one embodiment, when an LLM used for processing a user utterance is changed, the electronic device generates summary information rather than the entire utterance processing content and provides the summary information to the changed LLM, thereby efficiently managing information and / or data (e.g., increasing the efficiency of tokens associated with the LLM) and efficiently performing processing operations of the user utterance using the LLM. For example, the electronic device may increase the efficiency of utterance processing using the LLM by reducing the amount of resources and / or data required to process the utterance using the LLM.
[0143] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to receive user-defined values related to selection criteria of an LLM from the external electronic device.
[0144] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to select the first LLM based at least in part on the user-configured value.
[0145] According to one embodiment, the electronic device may provide a customized speech processing service to the user by selecting an LLM of a performance and / or category preferred by the user among a plurality of LLMs by selecting the LLM based on user-preferred values.
[0146]
[0147] Fig. 7 is a flowchart of a method for processing ignition of an electronic device according to one embodiment.
[0148] According to one embodiment, in operation 710, an electronic device (e.g., an electronic device (100) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (1101) of FIG. 11, a user terminal (1201) or an intelligent server (1300) of FIG. 12, or a generative artificial intelligence system (1500)) may receive information corresponding to at least one user utterance from an external electronic device (e.g., an external device (201) of FIG. 2, an electronic device (1102, 1104) of FIG. 11, or a user terminal (1201) of FIG. 12). For example, the external electronic device may receive the user utterance through a microphone of the external electronic device and transmit information corresponding to the received user utterance to the electronic device. According to one embodiment, when the electronic device includes a microphone, the electronic device may also receive the user utterance through the microphone of the electronic device. In this case, the electronic device may obtain information corresponding to the user utterance from the received user utterance.
[0149] According to one embodiment, in operation 720, the electronic device may recognize at least one intent and the number of intents included in at least one user utterance. For example, the user utterance may include at least one intent. For example, the at least one user utterance may include a single user utterance, a plurality of consecutive user utterances, and / or a plurality of related user utterances received at a predetermined time interval.
[0150] According to one embodiment, in operation 730, the electronic device may determine a processing level required to process at least one user utterance based at least in part on at least one intent and the number of intents. For example, the electronic device may determine a complexity of the user utterance based on at least one intent and / or the number of intents. For example, the complexity of the user utterance may be related to a number of processing steps required to process the user utterance. For example, the processing level may be determined based on at least one of a number of processing steps required to process the user utterance, a processing speed, an amount of data that can be processed, an amount of computation, a type of input data, a learning inference capability, a number of plug-ins available to the LLM, context information related to an external electronic device, and / or a combination thereof.
[0151] According to one embodiment, an electronic device may receive contextual information related to the external electronic device from an external electronic device. The electronic device may determine the level of processing based at least in part on the contextual information. For example, the contextual information may include a user profile, the number of available plug-ins, personal information stored on the external electronic device (e.g., schedule information), and / or information related to the status of the external electronic device.
[0152] According to one embodiment, an electronic device may obtain input data related to at least one user utterance. The electronic device may determine the processing level based at least in part on the type of the input data. For example, the input data may include data provided to the LLM (e.g., document files, image files, audio files, video files, and / or combinations thereof). For example, the electronic device may determine the processing level based at least in part on whether the type of the input data is text, audio, image, video, and / or a combination thereof.
[0153] According to one embodiment, an electronic device may use a trained artificial intelligence model to predict at least one of the number of processing steps, processing time, processing speed, computational load, or a combination thereof required to process at least one user utterance. The electronic device may determine a processing level based at least in part on the predicted result. For example, the artificial intelligence model may be trained using at least one of a language model, a deep learning model, an artificial neural network, a deep neural network, and / or a combination thereof. For example, the artificial intelligence model may be trained based on previously processed user utterances, a processing history of user utterances, and / or summary information about an operation that processed the user utterance.
[0154] According to one embodiment, the electronic device may receive user-preference values related to selection criteria of an LLM from an external electronic device (e.g., a device that receives a user utterance). For example, the user-preference values may be set based on user input received through an interface provided by the external electronic device. For example, the selection criteria of the LLM may include performance of the LLM (e.g., maximum capacity, size, processing speed, number of parameters, number of plug-ins supported, and / or specialized field of the LLM) and / or whether an automatic change function of the LLM is enabled. The electronic device may select a first LLM from among a plurality of LLMs based at least in part on the user-preference values.
[0155] According to one embodiment, in operation 740, the electronic device may select a first large language model (LLM) having a performance corresponding to the determined processing level from among a plurality of LLMs. For example, the electronic device may include an LLM database including the plurality of LLMs. According to one embodiment, the LLM database may be located external to the electronic device. For example, at least some of the plurality of LLMs may be stored in the electronic device (e.g., the LLM database of the electronic device), and the remaining some may be stored external to the electronic device (e.g., in an external electronic device and / or an LLM database external to the electronic device).
[0156] According to one embodiment, the electronic device may select at least one category from among a plurality of LLM categories based on at least one intent included in a user utterance. The electronic device may select a first LLM from among the LLMs belonging to the at least one selected category based on the number of intents included in the user utterance.
[0157] According to one embodiment, the electronic device may select the first LLM based on a result of comparing the processing level determined in operation 730 with a preset threshold value. For example, there may be one or more threshold values. The threshold values may be set to divide the processing level into a plurality of ranges. For example, each of the plurality of ranges may correspond to a performance of the LLM. For example, the electronic device may select one of the LLMs having a performance corresponding to the processing level from among the plurality of LLMs based on a result of comparing the processing level with the threshold value.
[0158] In one embodiment, in operation 750, the electronic device may process at least one user utterance using the first LLM. For example, the electronic device may query the first LLM for the user utterance. For example, the electronic device may provide the first LLM with a prompt corresponding to the user utterance and obtain a response to the prompt from the first LLM.
[0159] According to one embodiment, the electronic device may store a processing history of at least one user utterance processed using the first LLM. The electronic device may train or update an artificial intelligence model used to determine a processing level based on the processing history.
[0160] According to one embodiment of the present disclosure, an LLM having a function suitable for processing a user utterance is selected from among a plurality of LLMs based on a user utterance, a user setting value, context information related to an electronic device and / or an external electronic device, and / or a type of input data to be provided to the LLM, without a user input for manually selecting an LLM, and the user utterance is processed using the selected LLM, thereby improving the processing performance of the user utterance and / or the accuracy and reliability of the processing result, and increasing user convenience.
[0161] According to various embodiments, at least some of the operations described in FIG. 7 may be performed simultaneously, or the order of operations may be changed. According to various embodiments, at least some of the operations described in FIG. 7 may be omitted, or at least some operations (e.g., at least some of the operations of FIGS. 8 to 10) may be added. For example, the operations of FIG. 7 may be performed as a separate embodiment from the operations of FIGS. 8 to 10, or may be performed as an embodiment linked to each other.
[0162]
[0163] FIG. 8 is a flowchart of a method for processing ignition of an electronic device according to one embodiment. For example, the operations of FIG. 8 may be performed after operation 750 of FIG. 7.
[0164] According to one embodiment, in operation 810, an electronic device (e.g., an electronic device (100) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (1101) of FIG. 11, a user terminal (1201) or an intelligent server (1300) of FIG. 12, or a generative artificial intelligence system (1500)) may select a second LLM corresponding to the processing state among a plurality of LLMs based on recognizing that a processing state of at least one user utterance does not correspond to a determined processing level while processing at least one user utterance using a first LLM. For example, the electronic device may recognize that a current processing state of the user utterance (e.g., a number of processing steps, a type of input data, a processing speed, context information related to an external electronic device, an intent of the user utterance, and / or a number of intents included in the user utterance) does not correspond to a determined processing level. For example, the electronic device may recognize that a new user utterance is additionally input (the intents included in the user utterance may be changed / added or the number of intents may be changed depending on the additional user utterance), or that a processing step while processing the user utterance exceeds a pre-expected processing level, or that more input data is required / provided, or that an LLM with higher performance or an LLM with different performance is required to process the user utterance (or to increase the accuracy of the result of processing the user utterance). For example, the electronic device may select a second LLM that is more suitable for the processing state of the user utterance if the processing state of the user utterance does not correspond to the determined processing level. The electronic device may change the LLM used for processing the user utterance from the first LLM to the second LLM.
[0165] According to one embodiment, in operation 820, the electronic device may generate summary information summarizing the content of processing at least one user utterance using the first LLM. For example, the electronic device may generate the summary information by extracting some of the information corresponding to each processing step of processing the user utterance. For example, the summary information may include at least some of the information corresponding to the user utterance and / or the information obtained using the first LLM. The summary information may include fewer tokens than the number of tokens included in the information input to the first LLM and / or the information obtained (output) from the first LLM. For example, the electronic device may tokenize and manage the information input to the first LLM and / or the information obtained (output) from the first LLM. For example, the electronic device may generate the summary information including some of the tokens obtained from the information input to the LLM and / or the information obtained (output) from the first LLM.
[0166] In one embodiment, in operation 830, the electronic device may provide summary information to the second LLM. For example, the electronic device may provide the summary information as an input to the second LLM (e.g., as a prompt to the second LLM).
[0167] According to one embodiment, in operation 840, the electronic device may process at least one user utterance using a second LLM. For example, the electronic device may generate a prompt related to the intent contained in the at least one user utterance based on information corresponding to the at least one user utterance and provide the prompt to the second LLM. The electronic device may obtain a response to the prompt from the second LLM.
[0168] According to an embodiment of the present disclosure, while processing a user utterance, an LLM used for the user utterance is dynamically changed to an LLM having a function suitable for processing the user utterance among a plurality of LLMs, and the user utterance is processed using the changed LLM, thereby improving the processing speed and / or the accuracy and reliability of the processing result, and increasing user convenience.
[0169] According to various embodiments, at least some of the operations described in FIG. 8 may be performed simultaneously, or the order of operations may be changed. According to various embodiments, at least some of the operations described in FIG. 8 may be omitted, or at least some operations (e.g., at least some of the operations of FIGS. 7, 9, and 10) may be added. For example, the operations of FIG. 8 may be performed as a separate embodiment from the operations of FIGS. 7, 9, and 10, or may be performed as an embodiment linked to each other.
[0170]
[0171] Fig. 9 is a flowchart of a method for processing ignition of an electronic device according to one embodiment. In the following, operations identical or similar to those described in Figs. 7 and 8 are omitted or briefly described.
[0172] According to one embodiment, in operation 910, an electronic device (e.g., an electronic device (100) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (1101) of FIG. 11, a user terminal (1201) or an intelligent server (1300) of FIG. 12, or a generative artificial intelligence system (1500)) may receive information corresponding to a user utterance from an external electronic device (e.g., an external device (201) of FIG. 2, an electronic device (1102, 1104) of FIG. 11, or a user terminal (1201) of FIG. 12). For example, the information corresponding to the user utterance may include voice information corresponding to the user utterance and / or text information corresponding to the user utterance.
[0173] According to one embodiment, in operation 920, the electronic device can check a user-preference value related to an LLM selection criterion. For example, the user-preference value can be set by an external electronic device based on a user input. For example, the external electronic device can set a user-preference value related to the LLM selection criterion through the interface illustrated in FIGS. 3A to 3C. For example, the LLM selection criterion can include a maximum capacity of the LLM, a size of the LLM, whether an automatic change (or selection) function of the LLM is enabled, and / or a category of the LLM. For example, the electronic device can receive a user-preference value related to the LLM selection criterion set by the user from the external electronic device.
[0174] According to one embodiment, in operation 930, the electronic device may determine a processing level required to process the user utterance. For example, the processing level may include the complexity of the user utterance. For example, the complexity of the user utterance may include the number of processing steps required to process the user utterance using LLM. According to one embodiment, the electronic device may determine the processing level based on at least one intent included in the user utterance and / or the number of intents included in the user utterance. For example, the electronic device may determine the processing level based on a result of comparing the number of user intents included in the user utterance with a preset threshold. According to one embodiment, the electronic device may determine the processing level using a trained artificial intelligence model. The artificial intelligence model may be trained and / or updated based on information corresponding to the user utterance, a processing history of the user utterance, and / or summary information. According to one embodiment, the electronic device can determine the level of processing based at least in part on contextual information associated with an external electronic device (e.g., the electronic device that received the user utterance) (e.g., a user profile, the number of available plug-ins, personal information stored on the external electronic device (e.g., schedule information), and / or information associated with the state of the external electronic device). According to one embodiment, when there is input data associated with the user utterance (e.g., input data to be provided to the LLM), the level of processing can be determined at least in part on the type of the input data (e.g., text, audio, video, and / or a combination thereof).
[0175] According to one embodiment, in operation 940, the electronic device may select an LLM corresponding to a processing level from among a plurality of LLMs. For example, the electronic device may select an LLM having performance (e.g., learning inference capability, processing speed, amount of processing data, image processing capability, size, capacity, and / or category) corresponding to a user-set value and a processing level from an LLM database including a plurality of LLMs. For example, if there is no LLM having performance corresponding to the processing level in the LLM database, the electronic device may select an LLM closest to the user-set value and / or the processing level from among the plurality of LLMs. For example, the electronic device may determine a category of the LLM based on an intent included in a user utterance, and select one of the LLMs belonging to the determined category based on a processing level (e.g., number of processing steps or number of intents included in the user utterance).
[0176] According to one embodiment, in operation 940, if the current processing state of the user utterance does not correspond to the determined processing level during user utterance processing, the electronic device may select an LLM having performance corresponding to the processing state. For example, if the processing state does not correspond to the processing level, the electronic device may change the LLM corresponding to the previously selected processing level to an LLM corresponding to the processing state.
[0177] In one embodiment, in operation 950, the electronic device may process the user utterance using the selected LLM. For example, the electronic device may provide a prompt to the selected LLM based on information corresponding to the user utterance. The electronic device may obtain a response to the prompt from the selected LLM.
[0178] According to one embodiment, in operation 960, the electronic device may generate summary information regarding the processing of the user utterance using the selected LLM. For example, the electronic device may generate summary information related to a series of operations or processing steps that processed the user utterance using the selected LLM. The summary information may include at least a portion of the information input into the selected LLM and information obtained from the selected LLM. For example, the summary information may include at least a portion of the tokens included in the information input into the selected LLM and information obtained from the selected LLM.
[0179] In one embodiment, when an LLM is changed, the electronic device can provide summary information to the changed LLM. For example, the electronic device can efficiently manage tokens associated with a user utterance by providing summary information about the user utterance processed using the pre-change LLM to the changed LLM.
[0180] According to an embodiment of the present disclosure, by dynamically selecting or changing an LLM having a function suitable for processing user speech among a plurality of LLMs, the accuracy and reliability of the result of processing user speech using LLMs can be improved, user convenience can be increased, and the efficiency of processing user speech (e.g., token efficiency) can be increased by summarizing and managing and utilizing the content of processed user speech.
[0181] According to various embodiments, at least some of the operations described in FIG. 9 may be performed simultaneously, or the order of operations may be changed. According to various embodiments, at least some of the operations described in FIG. 9 may be omitted, or at least some operations (e.g., at least some of the operations of FIGS. 7, 8, and 10) may be added. For example, the operations of FIG. 9 may be performed as a separate embodiment from the operations of FIGS. 7, 8, and 10, or may be performed as an embodiment linked to each other.
[0182]
[0183] Fig. 10 is a flowchart of a method for processing ignition of an electronic device according to one embodiment. In the following, operations identical or similar to those described in Figs. 7 to 9 are omitted or briefly described.
[0184] According to one embodiment, in operation 1005, an electronic device (e.g., an electronic device (100) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (1101) of FIG. 11, a user terminal (1201) or an intelligent server (1300) of FIG. 12, or a generative artificial intelligence system (1500)) may obtain information corresponding to a user utterance. For example, the information corresponding to the user utterance may include voice information corresponding to the user utterance and / or text information corresponding to the user utterance.
[0185] According to one embodiment, in operation 1010, the electronic device may recognize the number of intents and user settings included in the user utterance. For example, the user utterance may include at least one intent. For example, the electronic device may receive user setting values related to LLM selection criteria from an external electronic device.
[0186] In one embodiment, in operation 1015, the electronic device may determine whether a learned model related to speech processing exists. For example, the learned model may be used to determine the level of processing required to process the user speech. For example, the learned model may be implemented based on at least one language model, a deep learning model, an artificial neural network, a deep neural network, prompt engineering, or a combination thereof.
[0187] For example, if a learned model exists, the electronic device can perform 1030 operations, and if a learned model does not exist, the electronic device can perform 1020 operations.
[0188] According to one embodiment, in operation 1020, the electronic device may determine a first threshold value associated with the user utterance. For example, the electronic device may compare the number of intents included in the user utterance with a preset first threshold value. For example, the threshold value may be one or more. For example, if there are multiple first threshold values, the electronic device may recognize a section in which the number of intents included in the user utterance is divided by the multiple first threshold values. For example, if the first threshold values are 2 and 4, the electronic device may determine whether the number of intents is less than 2, the number of intents is greater than or equal to 2 but less than or equal to 4, or the number of intentions is greater than or equal to 4.
[0189] According to one embodiment, in operation 1025, the electronic device may select an LLM based on a first threshold. For example, the electronic device may select a corresponding LLM based on a result of comparing the number of intents included in the user utterance with the first threshold. For example, assume that the first thresholds are 2 and 4. If the number of intents included in the user utterance is less than 2, the electronic device may select a first LLM having relatively low performance (e.g., relatively small size) from among the plurality of LLMs. If the number of intents included in the user utterance is 2 or more but less than 4, the electronic device may select a second LLM having relatively medium performance (e.g., relatively medium size) from among the plurality of LLMs. If the number of intents included in the user utterance is 4 or more, the electronic device may select a third LLM having relatively high performance (e.g., relatively large size) from among the plurality of LLMs.
[0190] According to one embodiment, in operation 1030, the electronic device may select an LLM using the learned model. According to one embodiment, the electronic device may determine a processing level required to process a user utterance based on the learned model. For example, the electronic device may predict the number of processing steps required to process the user utterance based on the learned model (hereinafter, the number of processing steps may be referred to as the term 'utterance complexity'). For example, the electronic device may select an LLM having a performance corresponding to the number of processing steps from among a plurality of LLMs. For example, the electronic device may recognize a pre-specified selection criterion corresponding to the number of processing steps (e.g., at least one threshold value related to the number of processing steps). For example, assume that the threshold values related to the number of processing steps are 3 and 5. The electronic device may select an LLM having relatively low performance (e.g., relatively small size) among multiple LLMs when the number of processing steps is less than three, select an LLM having relatively high performance (e.g., relatively large size) when the number of processing steps is five or more, and select an LLM having relatively medium performance (e.g., relatively medium size) when the number of processing steps is three or more but less than five. Although the processing level is described as being determined based on the number of utterance processing steps, the present invention is not limited thereto, and according to various embodiments, the processing level may be determined based on at least one of the processing speed, the amount of data that can be processed, the amount of computation, the type of input data, the learning inference capability, the number of plug-ins that the LLM can use, context information related to an external electronic device, and / or a combination thereof in addition to the number of processing steps.
[0191] In one embodiment, in operation 1035, the electronic device may process the user utterance using the selected LLM. For example, the electronic device may provide a prompt corresponding to the user utterance to the selected LLM. The electronic device may provide input data to the selected LLM along with the user utterance. The electronic device may obtain a response to the prompt from the selected LLM. In one embodiment, if there is summary information generated prior to using the selected LLM (e.g., summary information about an operation of processing the user utterance using another LLM), the electronic device may process the user utterance using the selected LLM based at least in part on the summary information. For example, the electronic device may provide the summary information as an input (or a prompt) to the selected LLM.
[0192] According to one embodiment, in operation 1040, the electronic device may determine whether a processing state of a user utterance exceeds a second threshold during utterance processing. For example, the electronic device may determine whether a second threshold related to a user utterance processing level (e.g., number of processing steps) during utterance processing is exceeded. For example, the second threshold may be a predetermined value related to the processing state of the user utterance and corresponding to a selected LLM. For example, if the electronic device predicts the number of processing steps of the user utterance and selects an LLM based on the predicted number of processing steps, an upper limit of the number of processing steps corresponding to the selected LLM may be set as the second threshold. For example, if the electronic device predicts that the number of processing steps is less than 3 and selects a first LLM corresponding to the predicted number of processing steps, the second threshold may be 3. If the electronic device predicts that the number of processing steps is 3 or more and less than 5 and selects a second LLM corresponding to the predicted number of processing steps, the second threshold may be 5.
[0193] For example, if the electronic device is using the LLM with the highest performance, operation 1040 may be omitted and operation 1045 may be performed. For example, if the threshold is exceeded, the electronic device may perform operation 1050, and if the threshold is below the second threshold, the electronic device may perform operation 1045.
[0194] While the case where the processing status of the utterance exceeds the second threshold was previously described, this is not limited thereto. For example, the second threshold may include a lower limit, rather than an upper limit, corresponding to the selected LLM. In this case, the electronic device may determine whether the processing status falls below the second threshold, and if so, may perform operation 1050.
[0195] In one embodiment, in operation 1045, the electronic device may complete processing of the user utterance. For example, the electronic device may obtain a response to the user utterance through the selected LLM and provide the obtained response to the user (e.g., an external electronic device that received the user utterance).
[0196] In one embodiment, in operation 1050, the electronic device may generate summary information regarding the processed operation. For example, the electronic device may generate summary information summarizing the operations that processed the user utterance (e.g., information corresponding to the user utterance and / or a response obtained from the LLM). For example, the summary information may include a smaller number of tokens than the tokens included in the information input to the LLM and / or the information obtained from the LLM.
[0197] According to one embodiment, in operation 1055, the electronic device may select another LLM. For example, the electronic device may select another LLM corresponding to the current processing status of the user utterance determined in operation 1035. For example, if the electronic device determines that a higher-performance LLM is required to process the user utterance, the electronic device may select the higher-performance LLM. For example, if the electronic device determines that a lower-performance LLM is sufficient to process the user utterance, the electronic device may select the lower-performance LLM. For example, if the electronic device determines that an LLM with a different performance (e.g., an LLM of a different category) is more suitable for processing the user utterance, the electronic device may select the LLM with a different performance. According to one embodiment, the electronic device may provide summary information as input to the selected other LLM.
[0198] According to an embodiment of the present disclosure, by dynamically selecting or changing an LLM having a function suitable for processing user speech among a plurality of LLMs, the accuracy and reliability of the result of processing user speech using LLMs can be improved, user convenience can be increased, and the efficiency of processing user speech (e.g., token efficiency) can be increased by summarizing and managing and utilizing the content of processed user speech.
[0199] According to various embodiments, at least some of the operations described in FIG. 10 may be performed simultaneously, or the order of operations may be changed. According to various embodiments, at least some of the operations described in FIG. 10 may be omitted, or at least some operations (e.g., at least some of the operations of FIGS. 7 to 9) may be added. For example, the operations of FIG. 10 may be performed as a separate embodiment from the operations of FIGS. 7 to 9, or may be performed as an embodiment linked to each other.
[0200]
[0201] A method for processing utterances of an electronic device according to one embodiment may include an operation of receiving information corresponding to at least one user utterance from an external electronic device.
[0202] According to one embodiment, the method may include an operation of recognizing at least one intent included in the at least one user utterance and a number of intents included in the at least one user utterance.
[0203] According to one embodiment, the method may include determining a level of processing required to process the at least one user utterance based at least in part on the at least one intent and the number of intents.
[0204] According to one embodiment, the method may include an operation of selecting a first LLM having a performance corresponding to the determined processing level from among a plurality of LLMs included in the LLM database.
[0205] According to one embodiment, the method may include processing the at least one user utterance using the first LLM.
[0206] According to one embodiment, the operation of determining the processing level may include receiving context information related to the external electronic device from the external electronic device.
[0207] According to one embodiment, the operation of determining the processing level may include an operation of determining the processing level based at least in part on the context information.
[0208] According to one embodiment, the operation of determining the processing level may include the operation of obtaining input data related to the at least one user utterance.
[0209] According to one embodiment, the operation of determining the processing level may include an operation of determining the processing level based at least in part on a type of the input data.
[0210] According to one embodiment, the type of the input data may include at least one of text, image, video, or a combination thereof.
[0211] According to one embodiment, the operation of determining the processing level may include an operation of predicting, using a trained artificial intelligence model, at least one of a number of processing steps, a processing time, a processing speed, an amount of computation, or a combination thereof required to process the at least one user utterance using at least one LLM among the plurality of LLMs.
[0212] According to one embodiment, the operation of determining the processing level may include an operation of determining the processing level based at least in part on the prediction result.
[0213] According to one embodiment, the method may include an operation of storing a processing history of the at least one user utterance processed using the first LLM.
[0214] In one embodiment, the method may include training or updating an artificial intelligence model used to determine the processing level based on the processing history.
[0215] According to one embodiment, the plurality of LLMs may be classified into a plurality of designated categories and stored in the LLM database.
[0216] According to one embodiment, the act of selecting the first LLM may include an act of selecting at least one category from among the plurality of categories based on the at least one intention.
[0217] According to one embodiment, the operation of selecting the first LLM may include an operation of selecting the first LLM from among LLMs belonging to the at least one selected category based on the number of intentions.
[0218]
[0219] According to one embodiment, the operation of selecting the first LLM may include an operation of selecting the first LLM based on a result of comparing the processing level with a preset threshold value.
[0220] According to one embodiment, the method may include an operation of selecting a second LLM corresponding to the processing state among the plurality of LLMs based on recognizing that, while processing the at least one user utterance using the first LLM, the processing state of the at least one user utterance does not correspond to the determined processing level.
[0221] According to one embodiment, the method may include processing the at least one user utterance using the second LLM.
[0222] According to one embodiment, the method may include an operation of generating summary information summarizing the content of the at least one user utterance processed using the first LLM.
[0223] In one embodiment, the method may include providing the summary information to the second LLM.
[0224] According to one embodiment, the operation of processing the at least one user utterance using the second LLM may include the operation of processing the at least one user utterance based on the summary information using the second LLM.
[0225] According to one embodiment, the method may include receiving user-defined values related to selection criteria of the LLM from the external electronic device.
[0226] According to one embodiment, the method may include selecting the first LLM based at least in part on the user-defined value.
[0227]
[0228] FIG. 11 is a block diagram of an electronic device (1101) within a network environment (1100) according to various embodiments. Referring to FIG. 11, in the network environment (1100), the electronic device (1101) may communicate with the electronic device (1102) via a first network (1198) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (1104) or the server (1108) via a second network (1199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1101) may communicate with the electronic device (1104) via the server (1108). According to one embodiment, the electronic device (1101) may include a processor (1120), a memory (1130), an input module (1150), an audio output module (1155), a display module (1160), an audio module (1170), a sensor module (1176), an interface (1177), a connection terminal (1178), a haptic module (1179), a camera module (1180), a power management module (1188), a battery (1189), a communication module (1190), a subscriber identification module (1196), or an antenna module (1197). In some embodiments, the electronic device (1101) may omit at least one of these components (e.g., the connection terminal (1178)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1176), camera module (1180), or antenna module (1197)) may be integrated into a single component (e.g., display module (1160)).
[0229] The processor (1120) may, for example, execute software (e.g., a program (1140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (1101) connected to the processor (1120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1120) may store commands or data received from other components (e.g., a sensor module (1176) or a communication module (1190)) in a volatile memory (1132), process the commands or data stored in the volatile memory (1132), and store result data in a non-volatile memory (1134). According to one embodiment, the processor (1120) may include a main processor (1121) (e.g., a central processing unit or an application processor) or an auxiliary processor (1123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1121). For example, when the electronic device (1101) includes the main processor (1121) and the auxiliary processor (1123), the auxiliary processor (1123) may be configured to use less power than the main processor (1121) or to be specialized for a given function. The auxiliary processor (1123) may be implemented separately from the main processor (1121) or as a part thereof.
[0230] The auxiliary processor (1123) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1160), the sensor module (1176), or the communication module (1190)) of the electronic device (1101), for example, on behalf of the main processor (1121) while the main processor (1121) is in an inactive (e.g., sleep) state, or together with the main processor (1121) while the main processor (1121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1180) or a communication module (1190)). In one embodiment, the auxiliary processor (1123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0231] The memory (1130) can store various data used by at least one component (e.g., the processor (1120) or the sensor module (1176)) of the electronic device (1101). The data can include, for example, software (e.g., the program (1140)) and input data or output data for commands related thereto. The memory (1130) can include a volatile memory (1132) or a non-volatile memory (1134).
[0232] The program (1140) may be stored as software in memory (1130) and may include, for example, an operating system (1142), middleware (1144), or an application (1146).
[0233] The input module (1150) can receive commands or data to be used in a component of the electronic device (1101) (e.g., a processor (1120)) from an external source (e.g., a user) of the electronic device (1101). The input module (1150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0234] The audio output module (1155) can output audio signals to the outside of the electronic device (1101). The audio output module (1155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0235] The display module (1160) can visually provide information to an external party (e.g., a user) of the electronic device (1101). The display module (1160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (1160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0236] The audio module (1170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (1170) can acquire sound through the input module (1150), output sound through the sound output module (1155), or an external electronic device (e.g., electronic device (1102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1101).
[0237] The sensor module (1176) can detect the operating status (e.g., power or temperature) of the electronic device (1101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0238] The interface (1177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1101) with an external electronic device (e.g., the electronic device (1102)). In one embodiment, the interface (1177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0239] The connection terminal (1178) may include a connector through which the electronic device (1101) may be physically connected to an external electronic device (e.g., the electronic device (1102)). According to one embodiment, the connection terminal (1178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0240] The haptic module (1179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0241] The camera module (1180) can capture still images and videos. According to one embodiment, the camera module (1180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0242] The power management module (1188) can manage power supplied to the electronic device (1101). According to one embodiment, the power management module (1188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0243] A battery (1189) may power at least one component of the electronic device (1101). In one embodiment, the battery (1189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0244] The communication module (1190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1101) and an external electronic device (e.g., electronic device (1102), electronic device (1104), or server (1108)), and the performance of communication through the established communication channel. The communication module (1190) may operate independently from the processor (1120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1190) may include a wireless communication module (1192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, a corresponding communication module can communicate with an external electronic device (1104) via a first network (1198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1192) can verify or authenticate the electronic device (1101) within a communication network such as the first network (1198) or the second network (1199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1196).
[0245] The wireless communication module (1192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1192) may support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1192) may support various requirements specified in the electronic device (1101), an external electronic device (e.g., the electronic device (1104)), or a network system (e.g., the second network (1199)). According to one embodiment, the wireless communication module (1192) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0246] The antenna module (1197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1198) or the second network (1199), may be selected from the plurality of antennas by, for example, the communication module (1190). A signal or power may be transmitted or received between the communication module (1190) and the external electronic device via the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1197).
[0247] According to various embodiments, the antenna module (1197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0248] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0249] According to one embodiment, commands or data may be transmitted or received between the electronic device (1101) and an external electronic device (1104) via a server (1108) connected to a second network (1199). Each of the external electronic devices (1102 or 104) may be the same or a different type of device as the electronic device (1101). According to one embodiment, all or part of the operations executed in the electronic device (1101) may be executed in one or more of the external electronic devices (1102, 104, or 108). For example, when the electronic device (1101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1101). The electronic device (1101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (1104) may include an Internet of Things (IoT) device. The server (1108) may be an intelligent server utilizing machine learning and / or a neural network.In one embodiment, an external electronic device (1104) or server (1108) may be included in the second network (1199). The electronic device (1101) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.
[0250]
[0251] FIG. 12 is a block diagram illustrating an integrated intelligence system according to one embodiment.
[0252] Referring to FIG. 12, an integrated intelligence system of one embodiment may include a user terminal (1201), an intelligent server (1300), and a service server (1400).
[0253] A user terminal (1201) of one embodiment (e.g., electronic device (1101) of FIG. 11) may be a terminal device (or electronic device) that can connect to the Internet, and may be, for example, a mobile phone, a smart phone, a personal digital assistant (PDA), a laptop computer, a television, white goods, a wearable device, an HMD (head mounted device), or a smart speaker.
[0254] According to the illustrated embodiment, the user terminal (1201) may include a communication interface (1290), a microphone (1270), a speaker (1255), a display (1260), a memory (1230), and / or a processor (1220). The above-listed components may be operatively or electrically connected to each other.
[0255] A communication interface (1290) (e.g., a communication module (1190) of FIG. 11) may be configured to be connected to an external device and transmit and receive data. A microphone (1270) (e.g., an audio module (1170) of FIG. 11) may receive sound (e.g., a user's speech) and convert it into an electrical signal. A speaker (1255) (e.g., an audio output module (1155) of FIG. 11) may output an electrical signal as sound (e.g., a voice). A display (1260) (e.g., a display module (1160) of FIG. 11) may be configured to display an image or video. In one embodiment, the display (1260) may also display a graphical user interface (GUI) of an app (or application program) being executed.
[0256] The memory (1230) of one embodiment (e.g., the memory (1130) of FIG. 11) may store a client module (1231), a software development kit (SDK) (1233), and multiple applications. The client module (1231) and the SDK (1233) may constitute a framework (or solution program) for performing general-purpose functions. In addition, the client module (1231) or the SDK (1233) may constitute a framework for processing voice input.
[0257] The above-described multiple applications (e.g., 1255a, 1255b) may be programs for performing a designated function. According to one embodiment, the multiple applications may include a first app (1235a) and / or a second app (1235b). According to one embodiment, each of the multiple applications may include a plurality of operations for performing a designated function. For example, the applications may include an alarm app, a message app, and / or a schedule app. According to one embodiment, the multiple applications may be executed by the processor (1220) to sequentially execute at least some of the multiple operations.
[0258] The processor (1220) of one embodiment can control the overall operation of the user terminal (1201). For example, the processor (1220) can be electrically connected to a communication interface (1290), a microphone (1270), a speaker (1255), and a display (1260) to perform a specified operation. For example, the processor (1220) can include at least one processor.
[0259] The processor (1220) of one embodiment may also execute a program stored in the memory (1230) to perform a designated function. For example, the processor (1220) may execute at least one of the client module (1231) or the SDK (1233) to perform the following operations for processing voice input. The processor (1220) may control the operations of multiple applications, for example, through the SDK (1233). The following operations described as operations of the client module (1231) or the SDK (1233) may be operations performed by the execution of the processor (1220).
[0260] The client module (1231) of one embodiment can receive a voice input. For example, the client module (1231) can receive a voice signal corresponding to a user utterance detected through a microphone (1270). The client module (1231) can transmit the received voice input (e.g., a voice signal) to the intelligent server (1300). The client module (1231) can transmit status information of the user terminal (1201) to the intelligent server (1300) together with the received voice input. The status information can be, for example, execution status information of an app.
[0261] The client module (1231) of one embodiment can receive a result corresponding to the received voice input from the intelligent server (1300). For example, the client module (1231) can receive a result corresponding to the received voice input if the intelligent server (1300) can produce a result corresponding to the received voice input. The client module (1231) can display the received result on the display (1260).
[0262] In one embodiment, the client module (1231) may receive a plan corresponding to the received voice input. The client module (1231) may display the results of executing multiple operations of the app according to the plan on the display (1260). For example, the client module (1231) may sequentially display the results of executing multiple operations on the display. The user terminal (1201) may, for example, display only some of the results of executing multiple operations (e.g., the result of the last operation) on the display.
[0263] According to one embodiment, the client module (1231) may receive a request from the intelligent server (1300) to obtain information necessary to produce a result corresponding to a voice input. According to one embodiment, the client module (1231) may transmit the necessary information to the intelligent server (1300) in response to the request.
[0264] The client module (1231) of one embodiment can transmit result information of executing multiple operations according to a plan to the intelligent server (1300). The intelligent server (1300) can use the result information to confirm that the received voice input has been processed correctly.
[0265] The client module (1231) of one embodiment may include a voice recognition module. According to one embodiment, the client module (1231) may recognize voice inputs that perform limited functions through the voice recognition module. For example, the client module (1231) may execute an intelligent app for processing voice inputs by performing organic actions in response to a specified voice input (e.g., "Wake up!").
[0266] An intelligent server (1300) of one embodiment may receive information related to a user voice input from a user terminal (1201) via a network (1299) (e.g., the first network (1198) and / or the second network (1199) of FIG. 11). According to one embodiment, the intelligent server (1300) may convert data related to the received voice input into text data. According to one embodiment, the intelligent server (1300) may generate at least one plan for performing a task corresponding to the user voice input based on the text data.
[0267] In one embodiment, the plan may be generated by an artificial intelligence (AI) system. The AI system may be a rule-based system, a neural network-based system (e.g., a feedforward neural network (FNN) and / or a recurrent neural network (RNN)), or a combination of the above or another AI system. In one embodiment, the plan may be selected from a set of predefined plans or may be generated in real time in response to a user request. For example, the AI system may select at least one plan from a plurality of predefined plans.
[0268] An intelligent server (1300) of one embodiment may transmit results according to a generated plan to a user terminal (1201), or transmit the generated plan to the user terminal (1201). According to one embodiment, the user terminal (1201) may display results according to the plan on a display. According to one embodiment, the user terminal (1201) may display results of executing an operation according to the plan on a display.
[0269] An intelligent server (1300) of one embodiment may include a front end (1310), a natural language platform (1320), a capsule database (1330), an execution engine (1340), an end user interface (1350), a management platform (1360), a big data platform (1370), or an analytic platform (1380).
[0270] The front end (1310) of one embodiment can receive a voice input received by the user terminal (1201) from the user terminal (1201). The front end (1310) can transmit a response corresponding to the voice input to the user terminal (1201).
[0271] According to one embodiment, the natural language platform (1320) may include an automatic speech recognition module (ASR module) (1321), a natural language understanding module (NLU module) (1323), a planner module (1325), a natural language generator module (NLG module) (1327), and / or a text to speech module (TTS module) (1329).
[0272] An automatic speech recognition module (1321) of one embodiment can convert a voice input received from a user terminal (1201) into text data. A natural language understanding module (1323) of one embodiment can use the text data of the voice input to determine the user's intention. For example, the natural language understanding module (1323) can perform syntactic analysis and / or semantic analysis to determine the user's intention. The natural language understanding module (1323) of one embodiment can use linguistic features (e.g., grammatical elements) of morphemes or phrases to determine the meaning of words extracted from the voice input, and can match the meaning of the determined words to the intention to determine the user's intention.
[0273] In one embodiment, the planner module (1325) can generate a plan using the intent and parameters determined by the natural language understanding module (1323). According to one embodiment, the planner module (1325) can determine a plurality of domains necessary to perform a task based on the determined intent. The planner module (1325) can determine a plurality of operations included in each of the plurality of domains determined based on the intent. According to one embodiment, the planner module (1325) can determine parameters necessary to execute the determined plurality of operations or result values output by the execution of the plurality of operations. The parameters and the result values can be defined as concepts of a specified format (or class). Accordingly, the plan can include a plurality of operations and / or a plurality of concepts determined by the user's intent. The planner module (1325) can determine the relationships between the plurality of operations and the plurality of concepts in a stepwise (or hierarchical) manner. For example, the planner module (1325) can determine the execution order of a plurality of actions based on the user's intention based on a plurality of concepts. In other words, the planner module (1325) can determine the execution order of a plurality of actions based on parameters required for the execution of the plurality of actions and results output by the execution of the plurality of actions. Accordingly, the planner module (1325) can generate a plan including association information (e.g., ontology) between the plurality of actions and the plurality of concepts. The planner module (1325) can generate the plan using information stored in a capsule database (1330) in which a set of relationships between concepts and actions is stored.
[0274] The natural language generation module (1327) of one embodiment can convert specified information into text format. The information converted into text format may be in the form of natural language speech. The text-to-speech conversion module (1329) of one embodiment can convert information in text format into information in speech format.
[0275] According to one embodiment, some or all of the functions of the natural language platform (1320) may also be implemented in the user terminal (1201). For example, the user terminal (1201) may include an automatic speech recognition module and / or a natural language understanding module. After the user terminal (1201) recognizes a user voice command, it may transmit text information corresponding to the recognized voice command to the intelligent server (1300). For example, the user terminal (1201) may include a text-to-speech conversion module. The user terminal (1201) may receive text information from the intelligent server (1300) and output the received text information as voice.
[0276] The capsule database (1330) may store information on the relationships between multiple concepts and actions corresponding to multiple domains. According to one embodiment, a capsule may include multiple action objects (or action information) and / or concept objects (or concept information) included in a plan. According to one embodiment, the capsule database (1330) may store multiple capsules in the form of a concept action network (CAN). According to one embodiment, the multiple capsules may be stored in a function registry included in the capsule database (1330).
[0277] The capsule database (1330) may include a strategy registry that stores strategy information required when determining a plan corresponding to a voice input. The strategy information may include reference information for determining a single plan when there are multiple plans corresponding to a voice input. According to one embodiment, the capsule database (1330) may include a follow-up registry that stores information on follow-up actions for suggesting follow-up actions to a user in a specified situation. The follow-up actions may include, for example, follow-up utterances. According to one embodiment, the capsule database (1330) may include a layout registry that stores layout information of information output through the user terminal (1201). According to one embodiment, the capsule database (1330) may include a vocabulary registry that stores vocabulary information included in capsule information. According to one embodiment, the capsule database (1330) may include a dialog registry that stores information on dialogue (or interaction) with a user. The capsule database (1330) may update stored objects through a developer tool. The developer tool may include, for example, a function editor for updating action objects or concept objects. The developer tool may include a vocabulary editor for updating vocabulary. The developer tool may include a strategy editor for creating and registering strategies that determine plans.The developer tool may include a dialog editor that creates a dialogue with the user. The developer tool may also include a follow-up editor that activates follow-up goals and allows editing of follow-up utterances that provide hints. The follow-up goals may be determined based on the currently set goals, user preferences, or environmental conditions. In one embodiment, the capsule database (1330) may also be implemented within the user terminal (1201).
[0278] The execution engine (1340) of one embodiment can produce a result using the generated plan. The end user interface (1350) can transmit the produced result to the user terminal (1201). Accordingly, the user terminal (1201) can receive the result and provide the received result to the user. The management platform (1360) of one embodiment can manage information used in the intelligent server (1300). The big data platform (1370) of one embodiment can collect user data. The analysis platform (1380) of one embodiment can manage the quality of service (QoS) of the intelligent server (1300). For example, the analysis platform (1380) can manage the components and processing speed (or efficiency) of the intelligent server (1300).
[0279] In one embodiment, a service server (1400) may provide a designated service (e.g., food ordering or hotel reservation) to a user terminal (1201). According to one embodiment, the service server (1400) may be a server operated by a third party. In one embodiment, the service server (1400) may provide information for generating a plan corresponding to a received voice input to an intelligent server (1300). The provided information may be stored in a capsule database (1330). In addition, the service server (1400) may provide result information according to the plan to the intelligent server (1300). The service server (1400) may communicate with the intelligent server (1300) and / or the user terminal (1201) via a network (1299). The service server (1400) may communicate with the intelligent server (1300) via a separate connection. Although the service server (1400) is depicted as a single server in FIG. 12, the embodiments of this document are not limited thereto. At least one of the services (1401, 1402, and 1403) of the service server (1400) may be implemented as a separate server.
[0280] In the integrated intelligence system described above, the user terminal (1201) can provide various intelligent services to the user in response to user input. The user input may include, for example, input via a physical button, touch input, or voice input.
[0281] In one embodiment, the user terminal (1201) may provide a voice recognition service through an intelligent app (or voice recognition app) stored internally. In this case, for example, the user terminal (1201) may recognize a user utterance or voice input received through the microphone and provide the user with a service corresponding to the recognized voice input.
[0282] In one embodiment, the user terminal (1201) may perform a designated action based on the received voice input, either alone or in conjunction with the intelligent server and / or service server. For example, the user terminal (1201) may execute an app corresponding to the received voice input and perform a designated action through the executed app.
[0283] In one embodiment, when a user terminal (1201) provides a service together with an intelligent server (1300) and / or a service server, the user terminal may detect user speech using the microphone (1270) and generate a signal (or voice data) corresponding to the detected user speech. The user terminal may transmit the voice data to the intelligent server (1300) using a communication interface (1290).
[0284] According to one embodiment, an intelligent server (1300) may generate a plan for performing a task corresponding to a voice input received from a user terminal (1201), or a result of performing an operation according to the plan, in response to a voice input. The plan may include, for example, a plurality of operations for performing a task corresponding to the user's voice input and / or a plurality of concepts related to the plurality of operations. The concept may define parameters input to the execution of the plurality of operations or result values output by the execution of the plurality of operations. The plan may include association information between the plurality of operations and / or the plurality of concepts.
[0285] In one embodiment, the user terminal (1201) can receive the response using the communication interface (1290). The user terminal (1201) can output a voice signal generated within the user terminal (1201) to the outside using the speaker (1255), or can output an image generated within the user terminal (1201) to the outside using the display (1260).
[0286] FIG. 13 is a diagram showing a form in which relationship information between concepts and actions is stored in a database according to one embodiment.
[0287] The capsule database (e.g., capsule database (1330)) of the above intelligent server (1300) can store capsules in the form of a CAN (concept action network). The capsule database can store operations for processing tasks corresponding to a user's voice input and parameters necessary for the operations in the form of a CAN (concept action network).
[0288] The capsule database may store a plurality of capsules (e.g., Capsule A (1331), Capsule B (1334)) corresponding to each of a plurality of domains (e.g., applications). According to one embodiment, one capsule (e.g., Capsule A (1331)) may correspond to one domain (e.g., location (geo), application). In addition, one capsule may correspond to at least one service provider's capsule (e.g., CP 1 (1332), CP 2 (1333), CP3 (1335), and / or CP4 (1336)) for performing a function for a domain related to the capsule. According to one embodiment, one capsule may include at least one operation (1330a) and at least one concept (1330b) for performing a specified function.
[0289] The natural language platform (1320) can generate a plan for performing a task corresponding to a received voice input using capsules stored in the capsule database (1330). For example, the planner module (1325) of the natural language platform can generate a plan using capsules stored in the capsule database. For example, a plan (1337) can be generated using the actions (1331a, 1332a) and concepts (1331b, 1332b) of capsule A (13310) and the action (1334a) and concept (1334b) of capsule B (1334).
[0290]
[0291] FIG. 14 is a diagram showing a screen in which a user terminal processes voice input received through an intelligent app according to one embodiment.
[0292] The user terminal (1201) can execute an intelligent app to process user input through an intelligent server (1300).
[0293]
[0294] According to one embodiment, in the first screen (1410), when the user terminal (1201) recognizes a designated voice input (e.g., wake up!) or receives an input via a hardware key (e.g., a dedicated hardware key), the user terminal (1201) may execute an intelligent app for processing the voice input. For example, the user terminal (1201) may execute the intelligent app while the schedule app is running. According to one embodiment, the user terminal (1201) may display an object (e.g., an icon) (1411) corresponding to the intelligent app on the display (1260). According to one embodiment, the user terminal (1201) may receive a voice input by a user's speech. For example, the user terminal (1201) may receive a voice input such as "Tell me my schedule for this week!" According to one embodiment, the user terminal (1201) can display a user interface (UI) (1213) (e.g., an input window) of an intelligent app on which text data of a received voice input is displayed.
[0295] According to one embodiment, in the second screen (1415), the user terminal (1201) may display a result corresponding to the received voice input on the display. For example, the user terminal (1201) may receive a plan corresponding to the received user input and display "This Week's Schedule" on the display according to the plan.
[0296]
[0297] FIG. 15 illustrates a generative artificial intelligence system according to one embodiment.
[0298] Referring to FIG. 15, in a generative artificial intelligence system (1500), a user question / response interface (1510) can receive user input. The user input may be in the form of natural language, images, and / or videos. Additionally, contextual information may also be transmitted when the user input is transmitted. The contextual information may include various additional information at the time of the user input. For example, information on the application currently being used by the user or information on the user's location. Furthermore, the user input may also be in a form that combines the aforementioned natural language, images, sounds, and contextual information. Furthermore, the user input may also be in a form other than natural language, such as selecting a menu.
[0299] According to one embodiment, the user question / response interface (1510) may output the results of a generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user. The user question / response interface (1510) may output the results of a generative artificial intelligence system (1500) to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user.
[0300] According to one embodiment, the artificial intelligence framework (1520) can receive user input and coordinate and control each component or module necessary to perform the user's intention based on the user's query.
[0301] In one embodiment, user input received from the user question / response interface (1510) may be transmitted to a prompt design module (1521). The prompt design module (1521) may be used to generate prompts suitable for inputting the user input into a large language model (LMM) or a large multi-modal model (LMM). The prompt design module (1521) may be an artificial intelligence component that uses a machine learning algorithm or a neural network to develop better prompts over time. The prompt design module (1521) may access a database (1530) containing user preference data, a prompt library, and prompt examples based on the user input to generate prompts, and may transmit the generated prompts to the LLM or LMM.
[0302] According to one embodiment, the API / plug-in management module (1522) may communicate with external information when there is a request for additional information when passing user input as input to the generative artificial intelligence model (1550). The API / plug-in management module (1522) may establish a channel for communicating with the outside of the artificial intelligence interface through an application programming interface (API) and may enable access to various data sources (e.g., a database (1530)) through the established channel. In addition, the API / plug-in management module (1522) may request the application / service module (930) to perform an action that ultimately performs the user input, rather than an intermediate result, through the API when the application / service module (1540) needs to perform the action. Information obtained from the outside may be used to generate a prompt in the prompt design module (1521) together with the user input or may be passed as an input to the generative artificial intelligence model (1550).
[0303] In one embodiment, the transformation module (1523) can fine-tune the output from the generative artificial intelligence model (1550). For example, the transformation module (1523) can verify whether the content generated through the LLM and / or LMM is irrelevant, biased, or harmful. In addition, the transformation module (1523) can determine to what extent the content matches the user's desired result and, if necessary, perform additional processing. The transformation module (1523) can additionally configure and provide the user with hints to avoid undesired output.
[0304] According to one embodiment, a generative artificial intelligence model (1550) may generally refer to an artificial intelligence neural network that creates new types of data based on user input information. The generative artificial intelligence model (1550) may include a model that generates images and / or a model that generates language. Representative models that generate images include a generative adversarial network (GAN) and a variational autoencoder (VAE), and examples include a diffusion-based generative model that uses a VAE and a transformer structure. A model that generates language is a model that is trained to statistically output the most appropriate output value based on an input value, and representative examples include models such as CHAT-GPT 3 and CHAT-GPT 4. In addition, there is also an LMM that can recognize various types of data input such as text, images, and voice and generate new data corresponding thereto.
[0305]
[0306] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0307] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0308] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0309] Various embodiments of the present document may be implemented as software (e.g., a program (1140)) including one or more instructions stored in a storage medium (e.g., an internal memory (1136) or an external memory (1138)) readable by a machine (e.g., an electronic device (1101)). For example, a processor (e.g., a processor (1120)) of the machine (e.g., an electronic device (1101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0310] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0311] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In electronic devices, communication circuit; A memory including a large language model (LLM) database and instructions including multiple large language models (LLMs); and comprising at least one processor, The above instructions, when executed by the at least one processor, cause the electronic device to: Receive information corresponding to at least one user utterance from an external electronic device, Recognize at least one intent included in said at least one user utterance and the number of intents included in said at least one user utterance, Based at least in part on the at least one intent and the number of intents, determine a level of processing (processing difficulty, processing grade, complexity) required to process the at least one user utterance, Among the above multiple LLMs, a first LLM having a performance corresponding to the determined processing level is selected, An electronic device for processing at least one user utterance using the first LLM.
2. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: Receive context information related to the external electronic device from the external electronic device, An electronic device that determines the level of processing based at least in part on said contextual information.
3. In claim 1 or 2, The above instructions, when executed by the at least one processor, cause the electronic device to: Obtaining input data related to at least one user utterance, To determine the level of processing based at least in part on the type of the input data; An electronic device wherein the type of the input data includes at least one of text, image, video, or a combination thereof.
4. In claims 1 to 3, The above instructions, when executed by the at least one processor, cause the electronic device to: Using a trained artificial intelligence model, predicting at least one of the number of processing steps, processing time, processing speed, computational amount, or a combination thereof required to process at least one user utterance using at least one LLM among the plurality of LLMs, An electronic device that determines the processing level based at least in part on the predicted results.
5. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: Store the processing history of at least one user utterance processed using the first LLM, An electronic device for training or updating an artificial intelligence model used to determine the processing level based on the processing history.
6. In claim 1, The above multiple LLMs are classified into multiple designated categories and stored in the LLM database, The above instructions, when executed by the at least one processor, cause the electronic device to: Selecting at least one category from among the plurality of categories based on at least one of the above intentions, An electronic device that selects the first LLM among the LLMs belonging to the at least one selected category based on the number of the intentions.
7. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: An electronic device that selects the first LLM based on a result of comparing the above processing level with a preset threshold value.
8. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: While processing at least one user utterance using the first LLM, based on recognizing that the processing status of the at least one user utterance does not correspond to the determined processing level, selecting a second LLM corresponding to the processing status among the plurality of LLMs, An electronic device for processing at least one user utterance using the second LLM.
9. In claim 8, The above instructions, when executed by the at least one processor, cause the electronic device to: Generate summary information summarizing the content of processing at least one user utterance using the first LLM, Provide the above summary information to the second LLM, An electronic device for processing at least one user utterance based on the summary information using the second LLM.
10. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: Receive user-defined values related to selection criteria of LLM from said external electronic device. An electronic device that selects said first LLM based at least in part on said user-defined values.
11. In a method for handling ignition of an electronic device, An act of receiving information corresponding to at least one user utterance from an external electronic device; An operation of recognizing at least one intent included in said at least one user utterance and a number of intents included in said at least one user utterance; An operation of determining a level of processing required to process said at least one user utterance based at least in part on said at least one intent and a number of said intents; An operation of selecting a first LLM having a performance corresponding to the determined processing level among a plurality of LLMs included in the LLM database; and A method comprising the action of processing at least one user utterance using the first LLM.
12. In claim 11, The action for determining the above processing level is: An operation of receiving context information related to an external electronic device from said external electronic device; and A method comprising determining the level of processing based at least in part on said context information.
13. In claim 11 or 12, The action for determining the above processing level is: An operation of obtaining input data related to at least one user utterance; and comprising an operation for determining the level of processing based at least in part on the type of said input data; A method wherein the type of the input data includes at least one of text, image, video, or a combination thereof.
14. In claims 11 to 13, The action for determining the above processing level is: An operation of predicting at least one of the number of processing steps, processing time, processing speed, computational amount, or a combination thereof required to process at least one user utterance using at least one LLM among the plurality of LLMs, using a trained artificial intelligence model; and A method comprising determining the level of processing based at least in part on the prediction results.
15. In claim 11, The above multiple LLMs are classified into multiple designated categories and stored in the LLM database, The action of selecting the above first LLM is: An operation of selecting at least one category from among the plurality of categories based on at least one intention; and A method comprising an action of selecting the first LLM among the LLMs belonging to the at least one selected category based on the number of the intentions.
Citation Information
Patent Citations
System and device for selecting a speech recognition model
KR102426717B1
Apparatus for providing query response based on query difficulty and method there of
KR102433072B1