Electronic device and its control method
The electronic device automatically selects the most suitable voice assistant by analyzing user keywords and adjusting associations based on historical usage and satisfaction, addressing the inconvenience of trigger words and improving command recognition.
Patent Information
- Application Number
- CN202011333239.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-26
- Filing Date
- 2020-11-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-11-24
AI Technical Summary
In the prior art, a user needs to speak the trigger word when using the voice assistant, which leads to inconvenience, and may directly say the command without saying the trigger word, resulting in the inappropriate selection of the voice assistant.
By setting an interface circuit and processor in the electronic device, the keywords in the user's vocabulary are identified, and based on predefined information and usage history, the voice assistant with the highest correlation degree is automatically selected for speech recognition, and the correlation degree is dynamically adjusted to optimize the selection of the voice assistant.
It enables automatic selection of the most suitable voice assistant without the user inputting additional trigger words, improving the accuracy of user experience and voice recognition.
Smart Images

Figure CN112951222B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an electronic device capable of capturing a user's speech to perform an operation according to an instruction of the speech, and a control method thereof. For example, it relates to an electronic device and a control method thereof related to executing a voice assistant for processing a user's speech. Background Art
[0002] In order to calculate and process predetermined information according to a specific process, electronic devices including electronic components such as a CPU, a chipset, and a memory for calculation can be classified into various types according to what information is to be processed or what its purpose is. For example, an electronic device may include an information processing device such as a PC or a server that processes general information, an image processing device that processes image data, an audio device that processes audio, a home appliance that performs housework, etc. The image processing device may be implemented as a display device that displays the processed image data as an image on a display panel provided therein.
[0003] An electronic device can receive a user input and perform a predetermined operation according to a command instruction corresponding to the received user input. The user input method varies according to the type of interface. For example, an electronic device can receive a control signal from a button or a remote controller according to a user's manipulation, detect a user's gesture through a camera, perform eye tracking through a camera, or receive a voice signal according to a user's speech through a microphone. For example, an electronic device can obtain a command according to a result of speech recognition processing of a user's speech and perform an operation indicated by the obtained command. Among speech processing methods that an electronic device can use, a voice assistant can be used. The voice assistant is set to convert a user's speech into text, analyze the converted text based on deep learning, and perform an operation according to the analyzed content. In this way, the voice assistant is an artificial intelligence-based assistant service that provides a more advanced and improved interpretation and analysis result of the speech than simply recognizing a user's speech to recognize a command.
[0004] According to a design method, an electronic device may include a plurality of voice assistants and may be configured to use any one of these voice assistants to process a user's speech. For example, when a user's speech is input through a microphone, the electronic device checks whether the user's speech includes a trigger word. The trigger word is a predefined word for identifying a voice assistant and is usually guided by a user to be spoken before an instruction. When a voice assistant is identified according to the trigger word, the electronic device causes the identified voice assistant to process the user's speech.
[0005] However, since it is inconvenient for the user to additionally speak a trigger word, there may be a situation where the user does not speak the trigger word but only speaks an instruction. According to the design method, each voice assistant can be optimized for a specific instruction, and there may be a voice assistant that is more suitable for processing the corresponding spoken instruction than the voice assistant specified by the user through the trigger word. Considering these problems, there may be a need for an electronic device that can provide a voice assistant suitable for processing user utterances without additional input from the user. SUMMARY OF THE INVENTION
[0006] An electronic device according to an exemplary embodiment includes: a voice input interface including an interface circuit, configured to receive an utterance; a processor configured to: obtain a keyword of the utterance received through the voice input interface, identify, based on predefined information regarding a degree of association between a plurality of voice assistants and a plurality of keywords, a voice assistant among the plurality of voice assistants whose degree of association with the obtained keyword is greater than a threshold degree of association, and perform speech recognition of the user utterance based on the identified voice assistant.
[0007] The processor may identify a degree of association of each voice assistant defined in the predefined information with the obtained keyword, and select the identified voice assistant having the highest degree of association.
[0008] Based on the number of the obtained keywords being plural, the processor may sum degrees of association of each voice assistant among the voice assistants with the plurality of obtained keywords, and compare the summed degrees of association of the plurality of voice assistants.
[0009] The predefined information may be provided based on a usage history of the electronic device.
[0010] The usage history may include: information obtained by counting a processing history of a predetermined keyword of the utterance by each voice assistant among the plurality of voice assistants.
[0011] The processor may adjust a degree of association of the identified voice assistant with the obtained keyword.
[0012] The processor may identify a satisfaction level with respect to a result of the performed speech recognition, and based on the identified satisfaction level, increase or decrease the degree of association of the identified voice assistant with the obtained keyword.
[0013] The processor may add a first adjustment value to the degree of association based on identifying a relatively high satisfaction level, and add a second adjustment value smaller than the first adjustment value to the degree of association based on identifying a relatively low satisfaction level, to adjust the degree of association.
[0014] The predefined information may include information obtained based on a plurality of other utterances.
[0015] The processor may display a UI configured to guide the recognized voice assistant based on the amount of data identifying the predefined information being no greater than a threshold, and perform speech recognition on the utterance in response to selecting the recognized voice assistant through the UI.
[0016] The processor may display a UI configured to guide the results of speech recognition for each of the plurality of voice assistants based on the amount of data identifying the predefined information being no greater than a threshold, and perform the results on any one of the voice assistants selected through the UI.
[0017] A method of controlling an electronic device according to an exemplary embodiment includes: receiving an utterance; obtaining a keyword of the received utterance; identifying, based on predefined information regarding a degree of association between a plurality of voice assistants and a plurality of keywords, a voice assistant having a degree of association with the obtained keyword greater than a threshold degree of association among the plurality of voice assistants; and performing speech recognition on the utterance based on the identified voice assistant. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] From the following detailed description in conjunction with the accompanying drawings, the above and other aspects, features, and advantages of certain embodiments of the present disclosure will become more apparent.
[0019] Figure 1 is a diagram showing an example environment in which an electronic device includes a plurality of voice assistants according to various embodiments;
[0020] Figure 2 is a block diagram showing an example configuration of an electronic device according to various embodiments;
[0021] Figure 3 is a flowchart showing an example method of controlling an electronic device according to various embodiments;
[0022] Figure 4 is a diagram of a table provided to an electronic device for reference according to various embodiments;
[0023] Figure 5 is a diagram showing reference according to various embodiments Figure 4 of a flowchart of an example operation of identifying a voice assistant based on a table;
[0024] Figure 6 is a diagram showing an example in which an electronic device updates a table according to various embodiments by reflecting a processing result of a user utterance Figure 4 of the updated result;
[0025] Figure 7is a diagram illustrating an example in which an electronic device obtains an initial table according to various embodiments;
[0026] Figure 8 is a diagram illustrating an example of a UI in which an electronic device displays recognition results according to various embodiments;
[0027] Figure 9 is a diagram illustrating an example in which an electronic device adjusts the score of a table in response to a user response according to various embodiments;
[0028] Figure 10 is a diagram illustrating an example in which an electronic device adjusts the score of a table through synonym management according to various embodiments;
[0029] Figure 11 is a flowchart illustrating an example method of controlling an electronic device according to various embodiments;
[0030] Figure 12 is a diagram illustrating an example of a UI in which an electronic device displays information including processing suitability regarding a plurality of voice assistants according to various embodiments;
[0031] Figure 13 is a diagram illustrating an example in which an electronic device adjusts the score of a table through management of related word categories according to various embodiments. DETAILED DESCRIPTION
[0032] Hereinafter, various exemplary embodiments according to the present disclosure will be described in more detail with reference to the accompanying drawings. Unless otherwise specified, the embodiments described with reference to each drawing are not mutually exclusive configurations, and multiple embodiments can be selectively combined and implemented in one device. The combination of multiple embodiments can be arbitrarily selected and applied by those skilled in the art of the present disclosure when implementing the spirit of the present disclosure.
[0033] If there are terms including ordinal numbers such as a first component, a second component, etc. in an embodiment, then these terms are used to describe various components, and the terms are used to distinguish one component from other components, and the components are not limited by these terms. The terms used in the applied embodiments are used to describe the embodiments, and the terms used do not limit the spirit of the present disclosure.
[0034] In addition, when describing "at least one" among a plurality of components in the present disclosure, this description refers not only to the whole of the plurality of components, but also to each component or all combinations of the plurality of components excluding the remaining components.
[0035] Figure 1 is a diagram illustrating an example environment in which an electronic device includes a plurality of voice assistants according to various embodiments.
[0036] As Figure 1As shown, the electronic device 100 according to an embodiment of the present disclosure may be implemented as, for example, a display device capable of displaying images. When the electronic device 100 is implemented as a display device, the electronic device 100 may include, for example, but not limited to, a television, a computer, a tablet computer, a portable media player, a wearable device, a video wall, an electronic photo frame, etc. However, in practice, the electronic device 100 may be implemented as various types of devices, such as, for example, but not limited to, a display device, an image processing device without a display such as a set-top box, a household appliance such as a refrigerator or a washing machine, an information processing device such as a computer body, etc. In addition, the electronic device 100 may be a device installed and used at a fixed position or a mobile device that a user can carry and use while moving.
[0037] The electronic device 100 may receive speech, for example, the user's voice. Then, the user utters a predetermined command, and the electronic device 100 obtains a voice signal according to the speech. To obtain a voice signal according to the speech, the electronic device 100 may include a microphone that collects the speech, or may receive the voice signal from a remote controller 140 having a microphone or a separate external device.
[0038] The electronic device 100 may include a plurality of voice assistants 110, 120, and 130. The voice assistants 110, 120, and 130 may include, for example, application services that determine the intention of the speech by analyzing the content and context of the user's speech based on, for example, artificial intelligence, and perform operations corresponding to the determination result. For example, the voice assistants 110, 120, and 130 may obtain text data by performing speech-to-text (STT) processing on the voice signal of the user's speech input to the electronic device 100, determine the meaning of the obtained text data by performing semantic analysis on the obtained text data based on deep learning or machine learning, and provide a service suitable for the determined meaning.
[0039] For example, the voice assistants 110, 120, and 130 may be device-level voice assistants 110 in which most operations are performed inside the electronic device 100, or may be included on servers 150 and 160 communicating with the electronic device 100, or may be voice assistants 120 and 130 that perform operations related to external devices.
[0040] For example, in the case of operating in combination with servers 150 and 160, the voice assistants 120 and 130 may operate as follows. When the user's speech is input to the electronic device 100, the voice assistants 120 and 130 send the voice signal of the user's speech to the servers 150 and 160. The servers 150 and 160 perform STT processing and semantic analysis on the received voice signal, and send the analysis result to the voice assistants 120 and 130. The voice assistants 120 and 130 perform operations corresponding to the analysis result.
[0041] When a user's utterance is input, the electronic device 100 may select any one of multiple voice assistants 110, 120, and 130 according to, for example, preset conditions. The preset conditions may vary according to the design method of the electronic device 100, and, for example, the voice assistants 110, 120, and 130 may be selected based on a trigger word input together with the user's utterance. The trigger word may include, for example, predefined words for identifying each of the voice assistants 110, 120, and 130. The electronic device 100 may perform STT processing on the voice signal in advance to identify the trigger word in the voice signal of the user's utterance.
[0042] However, the electronic device 100 according to an embodiment may handle a situation where the user's utterance does not include a trigger word and may recommend a more suitable voice assistant 110, 120, and 130 for the user's utterance. For such an operation, the electronic device 100 may automatically select the voice assistants 110, 120, and 130 in response to preset conditions separate from the trigger word, which will be described in more detail below.
[0043] Figure 2 is a block diagram showing an example configuration of an electronic device according to various embodiments.
[0044] As Figure 2 shown, the electronic device 210 may include a communication interface (e.g., including a communication circuit) 211, a signal input / output interface (e.g., including an input / output circuit) 212, a display 213, a user input interface (e.g., including an interface circuit) 214, a memory 215, a microphone 216, and a processor (e.g., including a processing circuit) 217.
[0045] Hereinafter, the configuration of the electronic device 210 will be described. Although this embodiment describes an example in which the electronic device 210 is a television, the electronic device 210 may be implemented as various types of devices, and thus this embodiment does not limit the configuration of the electronic device 210. In an example where the electronic device 210 cannot be implemented as a display device, the electronic device 210 may not include components for displaying an image, such as the display 213. For example, when the electronic device 210 is implemented as a set-top box, the electronic device 210 may output an image signal to an external television through the signal input / output interface 212.
[0046] The communication interface 211 may include various communication circuits, including, for example, a bidirectional communication circuit that includes at least one of components such as a communication module and a communication chip corresponding to various types of wired and wireless communication protocols. For example, the communication interface 211 may include: a wireless communication module that performs wireless communication with an AP according to the Wi-Fi system, a wireless communication module that performs one-to-one direct wireless communication such as Bluetooth, an IR module for IR communication, a LAN card that is wired to a router or gateway, and the like. The communication interface 211 may communicate with external devices such as a server 220 and a mobile device 240 on a network. The communication interface 211 may communicate with a remote controller 230 separated from the main body of the electronic device 210 to receive a signal transmitted from the remote controller 230.
[0047] The signal input / output interface 212 may include various input / output circuits and may be wired to an external device (e.g., a set-top box or an optical media player) in a 1:1 or 1:N (N is a natural number) manner, for example, to receive data from the external device or output data to the external device. The signal input / output interface 212 includes connectors, ports, etc. according to a predetermined transmission standard, for example, an HDMI port, a DisplayPort, a DVI port, a thunderbolt, and a USB port.
[0048] The display 213 may include a display panel capable of displaying an image on a screen. The display panel is set to, for example but not limited to, a light-receiving structure such as a liquid crystal type or a self-luminous structure such as an OLED type. Depending on the structure of the display panel, the display 213 may further include additional components. For example, if the display panel is of the liquid crystal type, the display 213 includes a liquid crystal display panel, a backlight unit that provides light, and a panel driving substrate that drives the liquid crystal of the liquid crystal display panel.
[0049] The user input interface 214 may include various interface circuits and may include various types of circuits related to the input interface, which are set to be manipulated by a user to perform user input. The user input interface 214 may be configured in various forms according to the type of the electronic device 210, and the user input interface 214 includes, for example, a mechanical or electronic button unit of the electronic device 210, a touchpad, a sensor, a camera, a touch screen mounted on the display 213, and the like.
[0050] The memory 215 stores digital data. The memory 215 may include: a non-volatile memory that can store data regardless of whether the memory is powered; and a volatile memory that can be loaded together with the data processed by the processor 217 and cannot store data when the memory is not powered. The memory may include a flash memory, a hard disk drive (HDD), a solid state drive (SSD), a read-only memory (ROM), etc., and the memory includes a buffer, a random access memory (RAM), etc. The memory 215 according to the present embodiment may store a plurality of applications for each of a plurality of voice assistants for processing voice signals. The applications stored in the memory 215 are driven by the processor 217 to execute the voice assistant.
[0051] The microphone 216 or the voice input interface collects sound including the user's speech from the external environment. The microphone 216 sends the voice signal of the collected sound to the processor 217.
[0052] The processor 217 may include various processing circuits, including, for example, one or more hardware processors implemented as, for example but not limited to, a CPU, a chipset, a buffer, a circuit, etc. mounted on a printed circuit board, and according to the design method, the processor 217 may be implemented as a system on a chip (SOC). When the electronic device 210 is implemented as a display device, the processor 270 may include various modules corresponding to various processes, such as a demultiplexer, a decoder, a scaler, an audio digital signal processor (DSP), and an amplifier. Some or all of these modules may be implemented as an SOC. For example, modules related to image processing, such as a demultiplexer, a decoder, and a scaler, may be implemented as an image processing SOC, and the audio DSP may be implemented as a chipset separate from the SOC.
[0053] When the voice signal of the user's speech is received through a predetermined path, the processor 217 according to the present embodiment may execute the selected voice assistant according to a preset method to process the voice signal. When a plurality of voice assistants are operating in the background, the processor may select any one of the voice assistants according to a preset method and cause the selected voice assistant to process the corresponding voice signal. An example method of selecting a voice assistant by the processor 217 will be described in more detail below.
[0054] The electronic device 210 may obtain the voice of the user's speech in various ways.
[0055] For example, the electronic device 210 may include a microphone 216 that collects sound. The voice signal of the user's speech collected by the microphone 216 is converted into a digital signal and sent to the processor 217.
[0056] When the remote controller 230 includes a microphone 236, the electronic device 210 may transmit, through the communication interface 211, a voice signal of a user's speech collected by the remote controller 230 through the microphone 236. The remote controller 230 may convert the voice signal of the user's speech collected through the microphone 236 into a digital signal, and transmit the digital signal to the communication interface 211 through the remote control communication interface 231 according to a protocol that the communication interface 211 can receive.
[0057] In the case of a general device such as the mobile device 240, the mobile device 240 may operate similarly to the remote controller 230 by installing and executing an application set for controlling the electronic device 210. The mobile device 240 may convert the voice signal of the user's speech collected through the microphone 246 into a digital signal while the application is being executed, and transmit the digital signal to the communication interface 211 through the mobile communication interface 241.
[0058] Hereinafter, a method in which the processor 217 selects any one of a plurality of voice assistants and causes the selected voice assistant to process a user's speech according to an embodiment of the present disclosure will be described in more detail.
[0059] Figure 3 is a flowchart illustrating an example method of controlling an electronic device according to various embodiments.
[0060] As Figure 3 shown, the following operations may be performed, for example, by a processor of an electronic device. In addition, the electronic device may include a plurality of voice assistants.
[0061] In operation 310, the electronic device receives a user's speech.
[0062] In operation 320, the electronic device obtains text of the received user's speech.
[0063] In operation 330, the electronic device obtains one or more keywords of the text from the obtained text.
[0064] In operation 340, the electronic device obtains predefined information on the degree of association between a plurality of voice assistants and a plurality of keywords. This information may include, for example, a table in which scores for each keyword among the plurality of keywords are recorded for each voice assistant.
[0065] In operation 350, the electronic device identifies, based on the obtained information, a voice assistant having a high degree of association (e.g., a degree of association greater than a threshold) with the keyword among the plurality of voice assistants. As an example of the identification method, the electronic device may select, from the table, a voice assistant having the highest score for the keyword included in the text of the user's speech.
[0066] In operation 360, the electronic device performs speech recognition on the user's utterance based on the recognized voice assistant.
[0067] Therefore, even if the user does not utter a trigger word or does not specify a particular voice assistant through a separate UI, the electronic device can automatically identify a voice assistant suitable for processing the user's utterance. The electronic device can provide a customized service to the user by selecting an appropriate voice assistant to process the user's utterance, regardless of whether the user specifies a voice assistant to process the user's utterance.
[0068] The processor of the electronic device can obtain the keywords of the user's utterance as described above; identify a voice assistant having a high degree of association with the obtained keywords based on predefined information about the degree of association between multiple voice assistants and multiple keywords; and based on the identified voice assistant, use at least one of, for example, but not limited to, machine learning, neural networks, deep learning algorithms (e.g., rule-based algorithms and artificial intelligence algorithms), etc., to perform at least a part of data analysis, processing, and result information generation for the operation of performing speech recognition on the user's utterance.
[0069] For example, the processor of the electronic device can perform the functions of the learning unit and the recognition unit together. The learning unit can perform the function of generating a trained neural network, and the recognition unit can perform the function of recognizing (or inferring, predicting, estimating, and determining) data using the trained neural network. The learning unit can generate or update the neural network. The learning unit can obtain learning data to generate the neural network. For example, the learning unit can obtain learning data from the memory of the electronic device or from the outside. The learning data can be data used for learning the neural network, and the data performing the above operations can be used as learning data to train the neural network.
[0070] Before learning the neural network using the learning data, the learning unit can perform a preprocessing operation on the obtained learning data or select the data to be used for learning from multiple learning data. For example, the learning unit can process or filter the learning data in a predetermined format, or process the data in a form suitable for learning by adding / removing noise. The learning unit can generate a neural network configured to perform the above operations using the preprocessed learning data.
[0071] The trained neural network may include multiple neural networks (or layers). Nodes of the multiple neural networks may have weights, and the multiple neural networks may be connected to each other such that the output value of one neural network is used as the input value of other neural networks. Examples of neural networks may include, for example, but not limited to, the following models: such as convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), deep Q-network, etc.
[0072] To perform the above operations, the recognition unit may obtain target data. The target data may be obtained from the memory of the electronic device or externally. The target data may include data to be recognized by the neural network. Before applying the target data to the trained neural network, the recognition unit may perform a preprocessing operation on the obtained target data, or select data to be used for recognition from among the multiple target data. For example, the recognition unit may process or filter target data in a predetermined format, or process the data in a form suitable for recognition by adding / removing noise. The recognition unit may obtain an output value output from the neural network by applying the preprocessed target data to the neural network. The recognition unit may obtain a probability value or a reliability value as well as the output value.
[0073] Hereinafter, examples of predefined information regarding the degree of association between multiple voice assistants and multiple keywords will be described in more detail.
[0074] Figure 4 FIG. is a diagram showing an example table provided to an electronic device for reference according to various embodiments.
[0075] As Figure 4 shown, the electronic device may obtain Table 400 indicating the degree of association between multiple voice assistants and multiple keywords. Each keyword included in Table 400 may be selected according to various methods. For example, after collecting words frequently used by the user, words of a predetermined rank or higher may be selected from the words according to the frequency of use. If words are selected by the server, the server may obtain the usage history of instructions from multiple clients, and select frequently used words from the obtained individual usage histories. Alternatively, words may be selected based on nouns excluding auxiliary words, adverbs, and pronouns from the attributes of the words.
[0076] Table 400 may include, for example, scores recorded for each of multiple predefined words for each of multiple voice assistants provided in the electronic device. The items shown in Table 400 are only examples for convenience and simplification of description, and the corresponding items do not limit the form, method, and content of the information regarding the degree of association between multiple voice assistants and multiple keywords.
[0077] For example, consider the following scenario: There are words such as "today", "weather", "recommendation", and "broadcast", and two voice assistants, such as a first assistant and a second assistant, are set in the electronic device. Considering this scenario, Table 400 includes the scores of the first assistant for each word and the scores of the second assistant for each word. For the word "today" in Table 400, the first assistant indicates a score of 10, while the second assistant indicates a score of 20.
[0078] The scores in Table 400 can be calculated (determined) according to various methods. For example, the scores can be prepared by reflecting the usage history of the electronic device. For example, the usage history can correspond to the historical count of the number of times a specific voice assistant processes a specific word when the user utters the word. The score can be set to a value equal to the cumulative number of times the word is processed, or can be set to a value obtained by multiplying the cumulative number of times the word is processed by a preset weight. For example, if the cumulative number of times the first assistant processes the word "today" is 10 times, then the score of the first assistant for "today" in Table 400 can be calculated as 10. If the predefined weight is 3 under the same conditions, then the score of the first assistant for "today" in Table 400 can be calculated as 30.
[0079] For example, the scores recorded in Table 400 can indicate the number of times a specific voice assistant processes a specific word. When designing a voice assistant based on AI, the more learning history there is, the more accurately the voice processing result can fit the user's intention. Therefore, it is desirable that a voice assistant with a high score will process the corresponding word more accurately. From this perspective, a voice assistant with a high score for a predetermined word is considered more suitable for processing that word.
[0080] The electronic device can generate Table 400 at the beginning, update and use the Table 400 with default settings in the manufacturing stage according to the usage history, or update and use the Table 400 provided by the server according to the usage history. In any case, the electronic device can achieve a Table 400 optimized for the user by updating the respective scores in Table 400 according to the usage history.
[0081] Hereinafter, a method by which the electronic device uses Table 400 to identify a voice assistant more suitable for the user's speech will be described in more detail.
[0082] Figure 5 is a flowchart showing an example operation of an electronic device identifying a voice assistant with reference to Figure 4 the table according to various embodiments.
[0083] As Figure 5 shown, the following operations can be performed, for example, by a processor of the electronic device. In addition, the electronic device can include multiple voice assistants, and can obtain Table 400 as Figure 4 shown (see Figure 4 ).
[0084] In operation 510, the electronic device recognizes one or more words from the text of the user's utterance. For example, consider the case where the user says "How's the weather today". The electronic device recognizes one or more words from the text of the user's utterance. Various criteria for recognizing words are possible. By way of example, the electronic device recognizes words predefined in Table 400 (see Figure 4 ). In this example case, the electronic device can recognize two words: "today" and "weather".
[0085] In operation 520, the electronic device obtains a table that defines scores of multiple voice assistants for multiple words. In this embodiment, reference is made to the table shown above in Figure 4 .
[0086] In operation 530, the electronic device recognizes the scores of the one or more recognized words from the table. For each of the recognized words "today" and "weather", the electronic device checks the scores of the first assistant and the second assistant. The score of the first assistant for the word "today" is 10, while the score of the second assistant is 20. The score of the first assistant for the word "weather" is 30, while the score of the second assistant is 10.
[0087] In operation 540, the electronic device compares the sums of the scores of each voice assistant in the multiple voice assistants for the recognized words with each other. For example, the sum of the scores of the first assistant is 10 + 30 = 40, and the sum of the scores of the second assistant is 20 + 10 = 30. Thus, the sum of the scores of the first assistant is greater than the sum of the scores of the second assistant.
[0088] In operation 550, the electronic device recognizes the voice assistant with the largest sum. When the user's utterance is "How's the weather today", the sum of the scores of the first assistant is the largest.
[0089] In operation 560, the electronic device processes the user's utterance through the recognized voice assistant. For example, when the first assistant is recognized, the electronic device causes the first assistant to perform speech recognition on the user's utterance. If necessary, the first assistant can send the voice signal of the user's utterance to the server to perform speech recognition.
[0090] In operation 570, the electronic device adjusts the scores of the corresponding voice assistant in the table for the recognized words according to the recognition of the voice assistant. For example, the electronic device can update the table by reflecting the history of the currently executed operations.
[0091] In the above exemplary embodiments, the following example has been described, in which the electronic device identifies, from a table, the scores of a voice assistant for each word recognized from a user's utterance, sums up the recognized scores for each voice assistant, compares the sums of the scores, and finally selects the voice assistant with the largest sum. However, the method by which the electronic device selects any one of the voice assistants based on the scores is not limited to comparing the sums of multiple scores, but various modified methods can be applied.
[0092] For example, in the exemplary embodiments, each score is simply summed, and there is no difference in weight among the multiple recognized words. However, according to the design method, it is feasible to add additional weights to the words considered more important by creating a difference in weight among the recognized words.
[0093] For example, assume that the words "today" and "weather" have been recognized from a user's utterance, and it is pre-set that the value obtained by multiplying the score for the word "today" by 1.2 and the value obtained by multiplying the score for the word "weather" by 1.0. This corresponds to the following example: the word "today" is reflected with higher importance compared to the word "weather". In Figure 4 the case of the table as an example, the result value of the first assistant is 10 * 1.2 + 30 * 1.0 = 42, and the result value of the second assistant is 20 * 1.2 + 10 * 1.0 = 34. By comparing the result values of multiple voice assistants in this way, the voice assistant with the largest result value can be selected.
[0094] Operation 570 in the above exemplary embodiments is a case where the table is updated by reflecting the result of the currently executed operation. Hereinafter, an example of such an operation will be described in more detail.
[0095] Figure 6 is a diagram showing an example of the result of updating the table of Figure 4 by the electronic device by reflecting the processing result of the user's utterance according to each embodiment.
[0096] As Figure 6 shown, when the electronic device identifies a voice assistant suitable for the user's utterance, the table 500 can be updated by reflecting the recognized result. For example, consider the following situation: the user's utterance is "How's the weather today", the recognized words are "today" and "weather", and the first assistant is finally selected from multiple voice assistants.
[0097] The electronic device checks, in the table 500, the scores of the finally selected voice assistant for the words recognized from the user's utterance. Since the recognized words are "today" and "weather", and the selected voice assistant is the first assistant, the electronic device can check in the table 500 that the score for "today" is 10 and the score for "weather" is 30.
[0098] The electronic device adds a score corresponding to a preset value only to the scores of the first assistant for the recognized words "today" and "weather". In this case, the score of the first assistant for "today" increases from 10 to 11, and the score of the first assistant for "weather" increases from 30 to 31. This embodiment has described that the increased value of the score is 1, but the specific value of the increased value is not limited. For the increased value of the score, according to the design method, different values can be applied corresponding to different conditions. In addition, the method of adjusting the score value can be not limited to increasing, but according to the conditions, the method of subtracting the score is feasible. Examples of the design method of the score will be described later.
[0099] Hereinafter, embodiments of a method by which the electronic device obtains an initial table will be described in more detail.
[0100] Figure 7 is a block diagram showing an example configuration in which an electronic device obtains an initial table according to various embodiments.
[0101] As shown in Figure 7 the figure, the electronic device 710 is connected to a network to communicate with the server 720. The server 720 communicates with a plurality of clients 730. The electronic device 710 may also be one of the plurality of clients 730, but different terms are specified only for mutual identification.
[0102] When constructing the table 711 as described in the above example embodiment, the electronic device 710 may generate the table 711 by accumulating usage history from the start. However, in this case, it takes time to accumulate the usage history until a specific amount of data is ensured to guarantee the reliability of the table 711.
[0103] As another method, during the manufacturing stage of the electronic device 710, a table 711 having initial values may be stored in the memory of the electronic device 710, and the electronic device 710 may be released as a product.
[0104] As another method, the electronic device 710 may receive the table 721 from the server 720, and may construct the table 711 by additionally reflecting the usage history to the initial values of the received table 721 and updating the table 721. The server 720 may provide the stored table 721 to the electronic device 710 only unidirectionally. The server 720 may receive feedback information related to the update of the table 711 from the electronic device 710, and update the previously stored table 721.
[0105] The server 720 can collect information about the tables 731 stored by each of the multiple clients 730, where the multiple clients 730 are communicatively connected in the same manner as the electronic device 710. Each client 730 can store the table 731 individually and update the table 731 it owns based on its own usage history. The method of updating each table 731 is the same as or similar to the method described in the above embodiments.
[0106] Each client 730 can periodically or in response to a request from the server 720 provide the current information of the table 731 to the server 720. The current information of the table 731 can include, for example but not limited to: predefined words, identification names of multiple voice assistants, scores of each word by the voice assistants, etc.
[0107] The server 720 can create a new table 721 or update the table 721 based on the information about the table 731 obtained from each client 730. Various design methods can be applied to the method of generating or updating the table 721. For example, the server 720 can obtain the scores of the same voice assistant for a specific word according to the information collected from each client 730, and obtain the value obtained by dividing the sum of the obtained scores by the number of the clients 730 as the score of the voice assistant for the corresponding word. This method is only an example, and the method of obtaining the value of the table 721 is not limited to this.
[0108] In response to a request from the electronic device 710 or in response to detecting that the electronic device 710 is connected to the server 720, the server 720 can provide the table 721 constructed as described above to the electronic device 710. The server 720 can obtain the updated information of the table 711 from the electronic device 710 and reflect the obtained updated information of the table 721 in the same manner as in the case of the client 730 described above.
[0109] When the electronic device identifies a voice assistant based on the table, the amount of data in the table needs to be greater than a predetermined threshold to ensure the reliability of the recognition result. If the amount of data in the table is less than the threshold, the electronic device can display the recognition result as a UI. Embodiments of displaying the UI will be described in more detail below.
[0110] Figure 8 It is a diagram showing an example of a UI in which an electronic device displays a recognition result according to various embodiments.
[0111] As Figure 8As shown, the electronic device 800 can identify a voice assistant suitable for the received user utterance and display a UI 810 indicating the recognition result. The UI 810 can at least include information indicating what the recognized voice assistant is and can also include the text of the received user utterance. For example, if a user utterance of "What's the weather like today" is received and a first assistant is selected from multiple voice assistants for the received user utterance, the electronic device 800 displays the UI 810 notifying the selection result.
[0112] The UI 810 can be displayed regardless of the situation according to whether the electronic device 800 is set, or can be selectively displayed. For example, when identifying a voice assistant suitable for a user utterance based on a previously stored table, the electronic device 800 identifies whether the amount of data in the table is greater than a threshold. When the amount of data in the table is greater than the threshold, the recognition result using the table has sufficient reliability. On the other hand, when the amount of data in the table is not greater than the threshold, for example, it may indicate that the reliability of the recognition result using the table is not high.
[0113] Therefore, when the amount of data in the table is greater than the threshold, the electronic device 800 does not display the UI 810 and causes the recognized voice assistant to process the user utterance. On the other hand, when the amount of data in the table is not greater than the threshold, the electronic device 800 displays the UI 810 to allow the user to select whether the recognition result is affirmative. To this end, the UI 810 can guide the recognition result and provide options for the user to select whether the recognition result is affirmative (i.e., whether the user can accept the recognition result) or negative (i.e., the user may not accept the recognition result and whether the user wants to process it through the voice assistant).
[0114] When an affirmative option for the recognition result is selected in the UI 810, the electronic device 800 causes the voice assistant recognized as in the above-described embodiment to process the user utterance to update the score of the table. On the other hand, when a negative option for the recognition result is selected in the UI 810, the electronic device 800 can proceed to a new process for identifying other voice assistants or can provide a separate UI configured to specify the voice assistant the user wants.
[0115] The electronic device can obtain the user's reaction to the recognition result of the voice assistant (i.e., an affirmative reaction or a negative reaction to the recognition result, or whether the user's satisfaction with the recognition result is high or low) through various methods. For example, the electronic device can identify whether the user is affirmative (i.e., whether the user's satisfaction is high) or negative (i.e., whether the user's satisfaction is low) with the recognition result based on the selection options provided through the UI 810 as in this embodiment.
[0116] The electronic device can recognize: the processing result of the first utterance provided by the voice assistant recognized for the user at the first time point, and receive a second utterance with the same content as the first utterance at a second time point within a preset time from the first time point. For example, this can indicate that the user is not satisfied with the processing result of the first utterance provided at the first time point. Therefore, in the case of recognition as described above, the electronic device can recognize that the user makes a negative reaction to the processing result of the first utterance provided at the first time point.
[0117] On the other hand, after the electronic device provides the processing result of the first utterance at the first time point, a second utterance with the same content as the first utterance may not be received within the preset time from the first time point. In this case, the electronic device can recognize that the user makes a positive reaction to the processing result of the first utterance provided at the first time point.
[0118] After the electronic device provides the processing result of the first utterance at the first time point, if an instruction such as cancel or stop is input within the preset time from the first time point, the electronic device can recognize that the user makes a negative reaction to the processing result of the first utterance provided at the first time point.
[0119] The reaction of the user recognized in this way can be reflected in the table by making the weights of the scores different when updating the scores in the table. Hereinafter, exemplary embodiments in which the electronic device adjusts the scores in the table in response to user actions will be described in more detail.
[0120] Figure 9 FIG. is a diagram illustrating an example in which an electronic device adjusts the scores in the table in response to a user reaction according to various embodiments.
[0121] As Figure 9 shown, when a user utterance is input, the electronic device recognizes predefined words from the user utterance, recognizes the scores of each voice assistant for the corresponding words by referring to Table 910, selects a voice assistant to process the user utterance according to the recognition result, and adjusts the scores in Table 910 according to the selection result. This process is as described in the above embodiments. For example, consider the following situation: two words, "today" and "weather", are recognized for the user utterance "What's the weather like today", and thus, based on Table 910, a first assistant is selected from multiple voice assistants.
[0122] The electronic device recognizes the user's reaction to the processing result of the user utterance "What's the weather like today" by the first assistant. The electronic device can recognize whether the user reaction is positive or negative according to various methods as described above.
[0123] If it is recognized that the user response to the processing result of the user's utterance by the first assistant is affirmative, the electronic device adds a preset first weight to the scores of the words recognized from the user's utterance by the first assistant in Table 910 to update Table 920. For example, assume that the score of the first assistant for "today" in Table 910 is 100 and the score for "weather" is 300. If the first weight value when the user response is affirmative is +5, the score of the first assistant for "today" in the updated Table 920 is 105 and the score for "weather" is 305.
[0124] On the other hand, if it is recognized that the user response to the processing result of the user's utterance by the first assistant is negative, the electronic device may add a second weight smaller than the first weight to the corresponding scores in Table 910 to update Table 930. For example, when the second weight is +1, the score of the first assistant for "today" in the updated Table 930 is 101 and the score for "weather" is 301.
[0125] Alternatively, if it is recognized that the user response to the processing result of the user's utterance by the first assistant is negative, the electronic device may add a third weight with a negative value to the corresponding scores in Table 910 to update Table 940. For example, when the third weight is -3, the score of the first assistant for "today" in the updated Table 940 is 97 and the score for "weather" is 297.
[0126] As described above, according to the user response to the processing result of the user's utterance, weights for the scores in Table 910 are applied differently, so the user's taste and preference can be more accurately reflected in Table 910.
[0127] On the other hand, in the exemplary embodiment, weights for the scores may be applied differently in response to the user response to the processing result, but the weights for the scores are not provided differently only in response to the user response. For example, the electronic device may identify the user who makes a sound, and may apply weights for the scores differently according to the identified user. There are several feasible methods for identifying the user. The electronic device may identify the user based on the currently logged-in account, or may identify the user corresponding to a profile by analyzing the waveform of the voice signal of the user's utterance.
[0128] As one of the methods for differently allocating weights to scores, the electronic device may manage synonyms. Hereinafter, an embodiment in which the electronic device adjusts scores through synonym management will be described in more detail.
[0129] Figure 10 is a diagram illustrating an example in which an electronic device adjusts scores of a table through synonym management according to various embodiments.
[0130] For example, asFigure 10 As shown, consider the following scenario: For the user's utterance "How's the weather today", two words "today" and "weather" are recognized, and thus, based on the scores in Table 1010, the first assistant is selected from multiple voice assistants. For example, if the first weight for the scores of the recognized words is set to +3, the electronic device adjusts the score of 100 of the first assistant for "today" in Table 1010 to 103, and adjusts the score of 300 of the first assistant for "weather" to 303.
[0131] The electronic device can identify whether there are synonyms of the recognized words among the words provided in Table 1010. For example, assume that among the predefined words in Table 1010, there is "climate" which is synonymous with "weather".
[0132] The electronic device can be set to assign a second weight smaller than the first weight for the scores of the recognized words to the synonyms of the recognized words. For example, if the second weight is set to +1 which is less than +3, the electronic device adjusts the score of 60 of the first assistant for "climate" in Table 1010 to 61.
[0133] In this way, in the updated Table 1020, the first weight is reflected in the scores of the recognized words, and the second weight smaller than the first weight is reflected in the scores of the synonyms of the recognized words.
[0134] On the other hand, if it is recognized that the amount of data in the table is less than the threshold or the usage history is small enough not to create a table, the electronic device can provide the user with information about the processing results of multiple voice assistants through the UI.
[0135] Figure 11 is a flowchart showing an example method of controlling an electronic device according to various embodiments.
[0136] As Figure 11 shown, the following operations can be performed by the processor of the electronic device. The electronic device may include multiple voice assistants for processing user utterances respectively.
[0137] In operation 1110, the electronic device receives a user utterance.
[0138] In operation 1120, the electronic device obtains information for identifying voice assistants. This information can be predefined information about the degree of association between multiple voice assistants and multiple keywords, and can represent the table in the above embodiments. Alternatively, this information can be the cumulative usage history for creating a table.
[0139] In operation 1130, the electronic device identifies whether the amount of the obtained information is less than the threshold.
[0140] If it is recognized that the amount of the obtained information is less than the threshold value ("Yes" in operation 1130), then in operation 1140, the electronic device causes each of the plurality of voice assistants to process the user's utterance.
[0141] When the processing results are output from each of the plurality of voice assistants, in operation 1150, the electronic device displays a UI for guiding the processing results of each of the plurality of voice assistants.
[0142] In operation 1160, the electronic device executes the processing result selected through the UI and updates the processing result by reflecting the processing result in the above information.
[0143] On the other hand, if it is recognized that the amount of the obtained information is not less than the threshold value ("No" in operation 1130), then in operation 1170, the electronic device selects any one of the plurality of voice assistants based on the obtained information.
[0144] In operation 1180, the electronic device processes the user's utterance through the selected voice assistant.
[0145] Accordingly, the electronic device can perform selective operations in response to the amount of information for identifying the voice assistant.
[0146] In the above exemplary embodiment, an example is described in which the electronic device causes the best voice assistant selected based on the table score to provide the processing result of the user's utterance. As a separate example, the electronic device may also be configured to provide the user with a rank list of a plurality of voice assistants and allow the user to select a voice assistant.
[0147] Figure 12 FIG. [FIGURE NUMBER] is a diagram illustrating an example in which an electronic device displays a UI including information on the processing suitability of a plurality of voice assistants according to various embodiments.
[0148] As Figure 12 shown, the electronic device 1200 may display a UI 1210 including information indicating the processing suitability of each of the plurality of voice assistants for the received user's utterance.
[0149] When a user's utterance is input, the electronic device 1200 recognizes a predefined word from the user's utterance, calculates the sum of the scores of each voice assistant for the corresponding word by referring to a table, selects the voice assistant with the highest sum of scores to process the user's utterance according to the calculation result, and adjusts the scores of the table according to the selection result. This process is as described in the above embodiment.
[0150] In the above process, when calculating the sum of scores of each voice assistant in multiple voice assistants for the corresponding word by referring to a table, according to the design method, the electronic device 1200 may display the UI 1210 without immediately selecting a voice assistant. The UI 1210 displays the sum of scores of each voice assistant in the multiple voice assistants for the user's utterance in a sorted manner. The user can compare the suitability of the multiple voice assistants set in the electronic device 1200 for the user's utterance through the UI 1210 and can select any one of these voice assistants.
[0151] The electronic device 1200 receives the user's utterance of "What's the weather like today", and when two words "today" and "weather" are recognized, calculates the sum of scores of each voice assistant in the multiple voice assistants for these two words based on the table. The calculation method is as described in the above embodiments.
[0152] The electronic device 1200 may display items related to the multiple voice assistants and the sum of scores of each voice assistant together on the UI 1210. In this case, in the UI 1210, the items of the multiple voice assistants may be arranged and displayed in ascending order of the sum of scores. According to the example of the accompanying drawings, since the sum of scores of the second assistant is the largest, which is 700, the second assistant is displayed at the highest position in the UI 1210.
[0153] The UI 1210 may be displayed through the settings of the electronic device 1200, or may be selectively displayed when it is recognized that the reliability of the table is not high.
[0154] In the above embodiments with reference to Figure 10 As one of the methods of differently assigning weights to scores, the case where the electronic device manages synonyms has been described. However, it may not be limited to only associating synonyms of predefined words prepared for recognition with the corresponding words, but may manage word categories related to the corresponding words according to various criteria.
[0155] Figure 13 is a diagram showing an example in which an electronic device according to various embodiments adjusts the scores of a table through the management of related word categories.
[0156] For example, as Figure 13 shown, consider the following situation: For the user's utterance of "What's the weather like today", two words "today" and "weather" are recognized, and thus, the first assistant is selected from the multiple voice assistants based on the scores in Table 1310. For example, if the first weight for the scores of the recognized words is set to +3, the electronic device adjusts the score of 100 of the first assistant for "today" in Table 1310 to 103, and adjusts the score of 300 of the first assistant for "weather" to 303.
[0157] The electronic device identifies whether there is a word category designated as related to the identified word among the words provided in Table 1310. Such a word category is a group of predefined words that can be synonymous with a reference word (i.e., a predefined word to be identified) (refer to the embodiments related to Figure 10 ), or can be words that are considered related even if they have different meanings, or can include search terms with various correlations to current trends in the SNS, etc. That is, words in a category related to a reference word can be selected by various methods and criteria, and for example, an AI model can be used to specify the word category related to the reference word.
[0158] For example, assume that among the predefined words in Table 1310, there is a "region" which is a word related to "weather".
[0159] The electronic device can be set to assign a second weight smaller than the first weight of the score for the identified word to words in the category related to the identified word. For example, if the second weight is set to +1 which is less than +3, the electronic device adjusts the score of 65 for "region" by the first assistant in Table 1310 to 66.
[0160] In this way, in the updated Table 1320, the first weight is reflected in the score for the identified word, and a second weight smaller than the first weight is reflected in the scores for other words related to the identified word.
[0161] The electronic device can store a predefined DB or list related to word categories synonymous or related to the reference word, and can identify other words related to the identified reference word from the stored DB or list. Since such a DB is provided to the electronic device by the server, the electronic device can use a DB that is periodically updated with words. When the server provides the updated DB to several electronic devices, the server can update the DB based on the usage history of the DB collected from each electronic device.
[0162] The operations of the device as described in the above embodiments can be performed by artificial intelligence installed in the device. Artificial intelligence can be applied to various systems using machine learning algorithms. An artificial intelligence system can include a computer system that realizes intelligence corresponding to or comparable to the human level, and can include a system in which a machine, device, or system autonomously performs learning and determination, and the recognition rate and determination accuracy are improved based on the accumulation of usage experience. Artificial intelligence technology can include machine learning technology that classifies / learns the features of input data using algorithms, elemental technologies that simulate functions such as recognition and determination of the human brain using machine learning algorithms, etc.
[0163] Examples of elemental technologies may include, for example but not limited to, at least one of the following: language understanding technologies for identifying human languages / characters, visual understanding technologies for identifying objects as humans do visually, reasoning / prediction technologies for logically inferring and predicting information by determining information, knowledge representation technologies for processing human experience information using knowledge data, or motion control technologies for controlling the autonomous driving of vehicles and the movement of robots.
[0164] For example, language understanding may refer to technologies for identifying and applying / processing human languages / characters, and includes natural language processing, machine translation, dialogue systems, question and answer, speech recognition / synthesis, etc.
[0165] For example, reasoning / prediction may refer to technologies for determining and logically predicting information, and includes knowledge / probability-based inference, optimization prediction, preference-based planning, recommendation, etc.
[0166] For example, knowledge representation may refer to technologies for automatically processing human experience information into knowledge data, and includes knowledge establishment (data generation / classification), knowledge management (data utilization), etc.
[0167] The method according to an embodiment of the present disclosure may be implemented in the form of program commands executable by various computer devices and may be recorded in a computer-readable recording medium. The computer-readable recording medium may separately include program commands, data files, data structures, etc. or include a combination thereof. For example, the computer-readable recording medium may be stored in a non-volatile memory (e.g., a USB memory device), a memory (e.g., a random access memory (RAM)), a read-only memory (ROM), a flash memory, a memory chip, or an integrated circuit, or a storage medium readable optically or magnetically by a machine (e.g., a computer) (e.g., a compact disc (CD), a digital versatile disc (DVD), a magnetic disk, a magnetic tape, etc.), regardless of whether the data is erasable or rewritable. The memory included in a mobile terminal is an example of a storage medium suitable for storing one or more programs including instructions readable by a machine for implementing an embodiment of the present disclosure. The program instructions recorded in the storage medium may be specifically designed and constructed for the present disclosure, or may be known and available to those skilled in the computer software field. The computer program instructions may also be implemented by a computer program product.
[0168] Although the present disclosure has been shown and described with reference to various exemplary embodiments, it should be understood that the embodiments are intended to be illustrative rather than restrictive. Those of ordinary skill in the art should also understand that various changes in form and detail may be made without departing from the true spirit and full scope of the present disclosure including the appended claims and their equivalents.
Claims
1. An electronic device, comprising: A voice input interface including a circuit, configured to receive a user voice input; And A processor, configured to: Obtain predefined information, in which, for each keyword of the voice input previously received through the voice input interface, the number of times each keyword is processed by a corresponding voice assistant records the score of each voice assistant among a plurality of voice assistants associated with each keyword among the plurality of keywords, Obtain keywords of the voice input received through the voice input interface, Based on the predefined information, obtain the score of the keyword obtained by the voice assistant associated with the obtained keyword among the plurality of voice assistants, identify the voice assistant having the highest score for the obtained keyword among the plurality of voice assistants, and Use the identified voice assistant to perform voice recognition on the voice input.
2. The electronic device according to claim 1, wherein, The processor is further configured to: in response to the number of the obtained keywords being plural, based on the predefined information, obtain the sum of the scores of each voice assistant among the plurality of voice assistants associated with the plurality of obtained keywords for the plurality of obtained keywords, and select the voice assistant with the highest sum of scores as the identified voice assistant.
3. The electronic device according to claim 1, wherein, The predefined information is based on the usage history of the electronic device.
4. The electronic device according to claim 3, wherein, The usage history includes: information obtained by counting the processing history of a predetermined keyword of the user's speech by each voice assistant among the plurality of voice assistants.
5. The electronic device according to claim 1, wherein, The processor is further configured to: adjust the score of the keyword obtained by the identified voice assistant.
6. The electronic device according to claim 5, wherein, The processor is further configured to: Identify the satisfaction with the result of the performed voice recognition, and Based on the identified satisfaction, increase or decrease the score of the keyword obtained by the voice assistant identified in the predefined information.
7. The electronic device according to claim 6, wherein, The processor is further configured to: add a first adjustment value to the score of the keyword obtained by the voice assistant identified in the predefined information based on identifying that the satisfaction is high, and add a second adjustment value less than the first adjustment value to the score of the keyword obtained by the voice assistant identified in the predefined information based on identifying that the satisfaction is low.
8. The electronic device according to claim 1, wherein, The predefined information includes information obtained based on a plurality of other utterances.
9. The electronic device according to claim 1, wherein, The processor is further configured to: based on identifying that the data volume of the predefined information is not greater than a threshold, control the display to display a UI indicating the identified voice assistant, and in response to selecting the identified voice assistant through the UI, perform voice recognition on the voice input.
10. The electronic device according to claim 1, wherein, The processor is further configured to: based on identifying that the data volume of the predefined information is not greater than a threshold, control the display to display a UI indicating the results of voice recognition performed by each voice assistant among the plurality of voice assistants, and perform the results on any one of the voice assistants selected through the UI.
11. A method for controlling an electronic device, comprising: Obtain predefined information in which the number of times each keyword based on a previously received voice input is processed by a corresponding voice assistant records the score of each voice assistant among a plurality of voice assistants associated with each keyword among the plurality of keywords; In response to receiving a voice input, obtain the keyword of the voice input; Identify, among the plurality of voice assistants, the voice assistant having the highest score for the obtained keyword by obtaining, based on the predefined information, the score of the obtained keyword by the voice assistant associated with the obtained keyword among the plurality of voice assistants; and Use the identified voice assistant to perform speech recognition on the voice input.
12. The method according to claim 11, further comprising: In response to the number of the obtained keywords being plural, obtain, based on the predefined information, the sum of the scores of the obtained plurality of keywords by each voice assistant associated with the obtained plurality of keywords among the plurality of voice assistants, and select the voice assistant with the highest sum of scores as the identified voice assistant.
13. The method according to claim 11 further comprises: Provide the predefined information based on the usage history of the electronic device.
Citation Information
Patent Citations
Synthesized voice selection for computational agents
US20180096675A1
Indicating a responding virtual assistant from a plurality of virtual assistants
US20180293484A1
Speaker command and key phrase management for muli -virtual assistant systems
US20190013019A1
Voice information processing method and device, and terminal
WO2019071607A1