Electronic device and Method for controlling the electronic device thereof

KR103015272B1Active Publication Date: 2026-09-04SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
KR1020190156158
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2019-11-28
Publication Date
2026-09-04
Estimated Expiration
2039-11-28

Smart Images

  • Figure 112019123219833-PAT00001_ABST
    Figure 112019123219833-PAT00001_ABST
Patent Text Reader

Abstract

An electronic device and a method for controlling the same are provided. The electronic device comprises a communication unit including a circuit, a microphone, a memory storing at least one instruction, and a processor executing at least one instruction. The processor determines whether to transmit a user voice input through the microphone to a server including a first conversational system by executing at least one instruction. If it is determined to transmit the user voice to the server, the processor controls the communication unit to transmit at least a portion of the user voice and the stored conversational history information to the server. Through the communication unit, the processor receives conversational history information related to the user voice from the server and controls the processor to store the received conversational history information in the memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to an electronic device and a method for controlling the same, and more specifically, to an electronic device and a method for controlling the same for determining a conversational system for providing a response to a user's voice based on a voice recognition result for an input user's voice. Background Technology

[0002] Artificial Intelligence (AI) systems are computer systems that achieve human-level intelligence. Unlike existing rule-based smart systems, they are systems that learn, make judgments, and become smarter on their own. As AI systems improve in recognition accuracy and become capable of understanding user preferences more accurately with continued use, existing rule-based smart systems are gradually being replaced by deep learning-based AI systems.

[0003] Artificial intelligence technology consists of machine learning (e.g., deep learning) and elemental technologies utilizing machine learning.

[0004] Machine learning is an algorithmic technology that autonomously classifies and learns the features of input data, and component technology is a technology that mimics the functions of the human brain, such as cognition and judgment, by utilizing machine learning algorithms such as deep learning, and consists of technology fields such as linguistic understanding, visual understanding, reasoning / prediction, knowledge representation, and motion control. In particular, linguistic understanding is a technology that recognizes, applies, and processes human language / characters, and includes natural language processing, machine translation, dialogue systems, question answering, and speech recognition / synthesis.

[0005] Meanwhile, users can utilize artificial intelligence systems stored on a device or server to perform speech recognition technology using machine learning. Previously, depending on the device's network conditions, if one of the AI ​​systems stored on the device or server was selected, only the selected AI system was used to perform speech recognition technology.

[0006] Previously, when a device's network conditions changed, the AI ​​systems in use were switched, but information regarding previously performed tasks was not shared between the switched AI systems. Consequently, there was a limitation in that the switched AI systems had to repeatedly ask questions to obtain the information necessary to perform the requested tasks. The problem to be solved

[0007] The present disclosure is devised to solve the aforementioned problems, and the purpose of the present disclosure is to provide an electronic device and a method for controlling the same, which determine a conversational system to provide a response to a user's voice based on the voice recognition results of the user's voice, and input user voice and conversation history information into the determined conversational system to provide a response to the user's voice. means of solving the problem

[0008] An electronic device according to one embodiment of the present disclosure for achieving the above objective comprises a communication unit including a circuit, a microphone, a memory storing at least one instruction, and a processor executing said at least one instruction. By executing said at least one instruction, the processor determines whether to transmit a user voice input through said microphone to a server including a first conversational system. If it is determined that the user voice is to be transmitted to said server, the processor controls said communication unit to transmit said user voice and at least a portion of said stored conversational history information to said server. Through said communication unit, the processor receives said conversational history information related to said user voice from said server and controls said received conversational history information to be stored in said memory.

[0009] Meanwhile, a server according to one embodiment of the present disclosure for achieving the above objective comprises a communication unit including a circuit, a first conversational system, at least one memory storing at least one instruction, and a processor executing said at least one instruction. The processor, by executing said at least one instruction, receives text corresponding to a user voice input to the electronic device and conversational history information stored in the electronic device from the electronic device through the communication unit, performs language analysis on said text through the first conversational system based on said conversational history information, and transmits the result of said language analysis to said electronic device. The received text may be characterized as being processed through a second conversational system stored in said electronic device.

[0010] Meanwhile, a control method for an electronic device according to one embodiment of the present disclosure for achieving the above objective may include the steps of determining whether to transmit an input user voice to a server including a first conversation system, transmitting at least a portion of the user voice and stored conversation history information to the server when it is determined to transmit the user voice to the server, receiving conversation history information related to the user voice from the server, and storing the received conversation history information.

[0011] Meanwhile, a control method for a server including a memory storing a first conversational system according to an embodiment of the present disclosure for achieving the above objective comprises the steps of receiving from an electronic device text corresponding to a user voice input to the electronic device and conversational history information stored in the electronic device, performing language analysis on the text through the first conversational system based on the conversational history information, and controlling a communication unit to transmit a result according to the performed language analysis to the electronic device, wherein the received text may be characterized as having been processed through a second conversational system stored in the electronic device. Effects of the invention

[0012] As described above, according to various embodiments of the present disclosure, an electronic device determines a conversation system to input user voice based on input user voice, and provides user voice and conversation history information to the determined conversation system to obtain a response to user voice, thereby enabling the user to utilize voice recognition technology more conveniently.

[0013] In addition, the electronic device obtains a response to the user's voice through at least one of a conversational system stored in the electronic device or a conversational system stored on a server, so that the user can receive a response to the user's voice naturally even if the user does not know which conversational system is being used. Brief explanation of the drawing

[0014] FIG. 1a is a diagram illustrating the process of determining a conversation system to transmit input user voice by an electronic device, according to one embodiment of the present disclosure. FIG. 1b is a diagram illustrating the process of determining a conversation system to transmit input user voice by an electronic device, according to one embodiment of the present disclosure. FIG. 2a is a flowchart for explaining a method for controlling an electronic device according to one embodiment of the present disclosure, FIG. 2b is a flowchart for explaining a method of controlling a server according to one embodiment of the present disclosure, FIGS. 3A, 3B, 3C, 3D, and 3E are flowcharts for explaining a process for determining whether an electronic device will transmit a user voice to a server, according to one embodiment of the present disclosure. FIG. 4a is a diagram briefly illustrating the configuration of an electronic device server according to one embodiment of the present disclosure. FIG. 4b is a diagram briefly illustrating the configuration of a server according to one embodiment of the present disclosure. FIG. 5 is a flowchart for explaining a control method of a server server according to one embodiment of the present disclosure, FIGS. 6a and 6b are drawings for explaining the operation between a software module of a server electronic device and a software module of a server according to one embodiment of the present disclosure. FIGS. 7, 8 and 9 are sequence diagrams for explaining the operation between an electronic device and a server according to one embodiment of the present disclosure, FIG. 10 is a block diagram illustrating in detail the configuration of an electronic device according to one embodiment of the present disclosure. Specific details for implementing the invention

[0015] Hereinafter, various embodiments of the present disclosure will be described with reference to the drawings.

[0016] FIG. 1a is a drawing for explaining a method of controlling an electronic device (100) according to one embodiment of the present disclosure.

[0017] As illustrated in FIG. 1a, according to one embodiment of the present disclosure, an electronic device (100) can determine whether to transmit an input user voice (10) (e.g., 'Give me directions to Seorae Village') to a server (200) including a first conversational system. Specifically, the electronic device (100) inputs the user voice (10) into a second conversational system to obtain a reliability value (e.g., a voice recognition reliability value or a language analysis reliability value) or a domain of the user voice (10), and determines whether to transmit the user voice (10) to the server (200) based on the obtained reliability value or domain. The process of the electronic device (100) determining whether to transmit the user voice (10) to the server (200) will be described in detail with reference to FIG. 2a.

[0018] Meanwhile, a domain is a type of data obtained as a result of semantic analysis of voice or text, referring to a category distinguished according to the type of user intent or control command corresponding to the voice or text. For example, if a user voice saying "Tell me today's weather" is input, the domain of the user voice may be "weather." A domain may be identical to a user intent or control command, and a single domain may include multiple user intents or control commands.

[0019] Meanwhile, the dialogue system may include an artificial intelligence model that recognizes and analyzes input user voice and provides a response to the user voice. The first dialogue system stored in the server (200) may be an artificial intelligence model trained using a larger amount of training data than the second dialogue system stored in the electronic device (100), or an artificial intelligence model having a larger amount of data than the amount of data of the second dialogue system stored in the electronic device (100). An artificial intelligence model trained using a large amount of training data can output recognition results with higher reliability for the same input voice compared to an artificial intelligence model trained using a relatively small amount of training data. Additionally, an artificial intelligence model trained using a large amount of training data can process voices related to more domains compared to an artificial intelligence model trained using a relatively small amount of training data, thereby controlling the electronic device to perform domain-related functions. Similarly, an artificial intelligence model having a large amount of data can output recognition results with higher reliability for the same input voice compared to an artificial intelligence model having a relatively small amount of data. In addition, an artificial intelligence model with a large amount of data can process voices related to more domains compared to an artificial intelligence model with a relatively small amount of data, thereby controlling the electronic device to perform domain-related functions. Generally, an artificial intelligence model with a large amount of data is a model trained using a large amount of training data compared to an artificial intelligence model with a relatively small amount of data. Therefore, the first conversational system can output a voice recognition result or a language analysis result with a high confidence value for the user voice (10) compared to the second conversational system. In addition, the first conversational system can perform domain-related functions that the second conversational system cannot process.

[0020] Meanwhile, in one embodiment, when it is determined that user voice (10) is transmitted to a server (200), the electronic device (100) may transmit at least some of the user voice (10) and stored conversation history information to a server (200) including a first conversation system. Meanwhile, in another embodiment, the electronic device (100) may transmit at least some of the text corresponding to the user voice (10) obtained through the second ASR module of the second conversation system and stored conversation history information to the server (200).

[0021] Conversation history information relates to the user voice (10) and includes information regarding voice recognition results, language analysis results, or responses of the conversation system obtained prior to the input of the user voice (10). Additionally, the conversation history information may further include information regarding operations performed by the electronic device (100) prior to the input of the user voice (10), and information regarding the state of the electronic device (100) at the time the user voice is input.

[0022] Additionally, the server (200) can perform language analysis on the text corresponding to the user voice (10) by inputting at least a portion of the text corresponding to the user voice (10) received from the electronic device (100) and the conversation history information stored in the electronic device (100) into the first conversation system. Meanwhile, in another embodiment, when the user voice is received from the electronic device (100), the server (200) can obtain the text corresponding to the user voice through the first conversation system.

[0023] Additionally, the server (200) can transmit the results of the language analysis of the text to the electronic device (100). Specifically, the server (200) can obtain a first language analysis result and a first language analysis reliability value by analyzing the text based on conversation history information, and obtain a second language analysis result and a second language analysis reliability value by analyzing the text only. Additionally, the server (200) can transmit one of the first language analysis result and the second language analysis result to the electronic device (100) based on the first and second language analysis reliability values. An example of the operation of the server (200) will be described in detail with reference to FIG. 2b.

[0024] Meanwhile, in one embodiment, the electronic device (100) may receive and store conversation history information related to the user voice (10) from the server (200). The conversation history information related to the user voice (10) illustrated in FIG. 1a may be situation information requesting directions to the location of the place name 'Seorae Village'. Specifically, the conversation history information may include information about the situation requesting directions, information about the destination of the directions called 'Seorae Village', application information about whether an application capable of providing directions services is installed on the electronic device (100), and information about whether the electronic device (100) has established a communication connection with the server (200).

[0025] Meanwhile, in one embodiment, information regarding the situation requesting directions and information regarding the destination of the directions, 'Seorae Village', may be information obtained through speech recognition results and language analysis results. The form of the conversation history information may be speech or natural language, but is not limited to.

[0026] And, in one embodiment, the conversation history information may include at least a portion of the user's voice or text regarding the user's voice. Additionally, the conversation history information may further include information regarding a response provided by the conversation system. Specifically, the information regarding a response provided by the conversation system may be at least a portion of a natural language sentence regarding a response message provided by the conversation system or at least a portion of the information used by the conversation system to generate the natural language sentence.

[0027] Meanwhile, the electronic device (100) can receive a response corresponding to the user voice (10) from the server (200) and control it to perform an action corresponding to the response. In one embodiment, the electronic device (100) can receive a response from the server (200) regarding the user voice (10) to "perform directions to the location of the place name Seorae Village." Here, the response refers to information about a response message to be output through the electronic device (100), and the information about the response message may include a natural language sentence to be output by the electronic device (100) or information for generating a natural language sentence. Additionally, the response refers to information about an action to be performed by the electronic device (100), and the information about the action may include information about an application to be executed on the electronic device (100) or information about the detailed functions of the application to be performed.

[0028] For example, the electronic device (100) can output a message in the form of voice or display it in the form of text, such as "Starting directions to Seorae Village," in response to a response received from the server (200). Additionally, the electronic device (100) can execute a directions application in response to the received response and perform a detailed function of starting directions to the location of the place name "Seorae Village."

[0029] Meanwhile, as illustrated in FIG. 1a, when additional user voice (20) (e.g., 'Show me the photo taken there') is input, the electronic device (100) can determine whether to transmit the input additional user voice (20) to the server (200). Specifically, the electronic device (100) can input the additional user voice (20) into the second conversational system and determine whether to transmit it to the server (200). The process of the electronic device (100) determining whether to transmit the additional user voice (20) to the server (200) will be explained in detail with reference to FIG. 2a.

[0030] If it is decided not to transmit additional user voice (20) to the server (200), the electronic device (100) can perform speech recognition or language analysis of the additional user voice (20) using conversation history information related to the user voice (10) among conversation history information stored in the second conversation system, and obtain a response to the additional user voice (20) and conversation history information related to the additional user voice (20). The conversation history information related to the user voice (10) may include information related to conversations that were input or responded to prior to the input of the additional user voice (20). At this time, the electronic device (100) may pre-set a predetermined number of times and obtain information related to conversations that were input or responded to within the pre-set number of times among the previously input or responded conversations. Additionally, the electronic device (100) can perform speech recognition or language analysis on the additional user voice (20) and obtain information about conversations related to speech recognition results or language analysis results among the previously input or responded conversations.

[0031] That is, the electronic device (100) can obtain conversation history information and a response related to the additional user voice (20) by utilizing conversation history information related to the user voice (10) received from the server (200) before the additional user voice (20) is input. In one embodiment, the electronic device (100) can identify that the text corresponding to the voice 'there' included in the additional user voice (20) means the place name 'Seorae Village' through situation information requesting directions to the location of the place name 'Seorae Village' among the conversation history information related to the user voice (10).

[0032] Additionally, conversation history information related to the additional user voice (20) may include information regarding a situation requesting a photo search related to 'Seorae Village', application information indicating that an application capable of searching for photos stored in the electronic device (100) is installed, and information regarding the communication connection status between the electronic device (100) and the server (200) at the time the second user voice (20) is input, but this is merely one embodiment.

[0033] Additionally, the electronic device (100) can perform a function corresponding to a response to an additional user voice (20). In one embodiment, the electronic device (100) can output a voice saying "This is a photo taken in Seorae Village" as a response to the additional user voice (20). Then, the electronic device (100) can control the display of a photo among the stored photos by executing an application capable of displaying a photo corresponding to the response, so that the photo with location information of "Seorae Village" is displayed. That is, the series of processes in which the electronic device (100) inputs the input additional user voice (20) into the first conversational system, outputs a response message in the form of voice, and executes a specific application can operate similarly to the embodiment of providing directions to Seorae Village described above in relation to the user voice (10).

[0035] FIG. 1b is a diagram illustrating the process of determining a conversation system to transmit input user voice by an electronic device (100) according to one embodiment of the present disclosure.

[0036] In one embodiment, as illustrated in (a) of FIG. 1b, when a user voice (30) saying "give me directions" is input, the electronic device (100) can determine whether to transmit the user voice (30) to a server (200). For example, the electronic device (100) can input the user voice (30) into a second ASR module of a second conversational system to obtain text and voice recognition reliability values ​​corresponding to the user voice (30).

[0037] In one embodiment, when the voice recognition reliability value exceeds a first threshold, the electronic device (100) may obtain a response to the user voice (30) and conversation history information related to the user voice (30) through a second conversation system. Additionally, the electronic device (100) may store conversation history information related to the user voice (30). The conversation history information related to the user voice (30) may include, but is not limited to, information regarding situations where the user requests directions, information regarding whether an application capable of performing directions is installed, etc. Meanwhile, the threshold value described in this disclosure may be a preset value, but this is merely one embodiment and it is understood that it may be changed by a user command.

[0038] And, the electronic device (100) can provide a response to the user's voice (30). For example, the electronic device (100) can provide a response message asking for a destination to provide directions, such as “Where are you going?” as shown in (a) of FIG. 1b.

[0039] Meanwhile, as illustrated in (b) of FIG. 1b, when a user voice (40) named 'Seorae Village' corresponding to a response to a previous user voice (30) is input, the electronic device (100) can determine whether to transmit the user voice (40) to the server (200). If the voice recognition reliability value for the user voice (40) is below a first threshold value, the electronic device (100) can transmit conversation history information related to the user voice (40) and the previously stored previous user voice (30) to the server (200).

[0040] In addition, in one embodiment, the server (200) can obtain a response to the user voice (40) and conversation history information related to the user voice (40) through a first conversation system based on conversation history information related to the received user voice (40) and the previous user voice (30). Specifically, the server (200) can determine that the intention of the text corresponding to the user voice (40) is to request directions to a destination called 'Seorae Village' by using conversation history information related to the previous user voice (30) which includes situation information that the current user is requesting directions. In addition, the server (200) can obtain a response to the user voice (40) (e.g., a response message saying “Starting directions to Seorae Village” and information related to an application that provides directions to Seorae Village) and conversation history information related to the user voice (40) (e.g., information about the situation of providing directions to a destination called Seorae Village, and information about an application that provides directions) through the first conversation system. And, the server (200) can transmit a response to the user voice (40) and conversation history information related to the user voice (40) to the electronic device (100). And, as shown in (c) of FIG. 1b, when a user voice (50) corresponding to the response to the previous user voice (40), such as "Oh, what is the weather like there now?" is input, the electronic device (100) can decide whether to transmit the user voice (50) again to the server (200). If the voice recognition reliability value of the user voice (50) exceeds a first threshold, the electronic device (100) can identify, through the conversation history information related to the previous user voice (40), that the meaning of the word "there" included in the text corresponding to the current user voice (50) is a destination called "Seorae Village."And, the electronic device (100) can obtain a response to the current user voice (50) and information related to the user voice (50) through a second conversation system. The response to the user voice (50) may be a message about the current weather in Seorae Village (e.g., “The weather in Seorae Village is currently cloudy and the temperature is 15 degrees”). And, conversation history information related to the user voice (50) may include situation information about the question about the weather in Seorae Village and information about the current weather in Seorae Village.

[0041] That is, as illustrated in FIG. 1b, the electronic device (100) can obtain a response to a user voice by using at least one of a second conversation system or a first conversation system stored in a server (200) that stores conversation history information about the user voice and continuously input user voice.

[0043] FIG. 2a is a drawing for explaining a method of controlling an electronic device (100) according to one embodiment of the present disclosure.

[0044] First, the electronic device (100) can determine whether to transmit the input user voice to a server (200) including a first conversational system (S210-1). Specifically, the electronic device (100) can input the user voice into a second conversational system to obtain a reliability value or domain of the user voice, and determine whether to transmit the user voice to the server (200) based on the obtained reliability value or domain.

[0045] In one embodiment, the electronic device (100) can obtain text corresponding to the user's voice and a voice recognition confidence score of the user's voice through the second ASR module of the second conversational system. The voice recognition confidence score is a numerical value indicating how accurately the user's voice is recognized and converted into text.

[0046] Additionally, the electronic device (100) can determine whether to transmit the user voice to the server (200) based on the voice recognition reliability value of the user voice. For example, if the voice recognition reliability value of the user voice is below a first threshold, the electronic device (100) can decide to transmit the user voice to the server (200). Meanwhile, if the voice recognition reliability value of the user voice exceeds the first threshold, the electronic device (100) can obtain a response to the user voice and conversation history information related to the user voice through the second conversation system. An embodiment related to the second ASR module will be described in detail with reference to FIG. 3a.

[0047] In another embodiment, the electronic device (100) can obtain a language analysis reliability value for text corresponding to the user's voice through the second NLU module of the second conversational system. The language analysis reliability value is a numerical value indicating the degree of reliability with which the meaning of the text corresponding to the user's voice was analyzed and determined.

[0048] Additionally, the electronic device (100) can determine whether to transmit the user voice to the server (200) based on the language analysis reliability value of the user voice (10). For example, if the language analysis reliability value of the user voice is below a second threshold, the electronic device (100) can decide to transmit the user voice to the server (200). Additionally, if the language analysis reliability value of the user voice exceeds the second threshold, the electronic device (100) can obtain a response to the user voice and conversation history information related to the user voice through a second conversation system. An embodiment related to the language analysis reliability value will be described in detail with reference to FIG. 3b.

[0049] In another embodiment, the electronic device (100) can obtain a domain of text corresponding to the user voice and information related to the domain through a second NLU module. Then, the electronic device (100) can determine whether to transmit the user voice (10) to the server (200) based on the information related to the domain obtained. An embodiment related to the domain will be described in detail with reference to FIG. 3c.

[0050] In another embodiment, the electronic device (100) may determine whether to transmit user voice to the server (200) based on the status information of the electronic device (100). According to one embodiment, the electronic device (100) may determine whether to transmit user voice to the server (200) based on the current battery charge level of the electronic device (100), the status of the communication connection with the server, etc. An embodiment related to the status information of the electronic device (100) will be described in detail with reference to FIG. 3d.

[0051] In another embodiment, the electronic device (100) may determine whether to transmit the user voice to the server (200) according to a conversation system selected by the user. In one embodiment, when a conversation system stored in the server (200) is selected, the electronic device (100) may decide to transmit the user voice to the server (200). An embodiment related to user selection will be described in detail with reference to FIG. 3e.

[0052] Meanwhile, if it is decided to transmit the user voice to the server (200), the electronic device (100) can transmit at least some of the user voice (or text corresponding to the user voice) and conversation history information to the server (200) (S220-1). Accordingly, the first conversation system of the server (200) can output a response to the user voice and conversation history information related to the user voice using the received conversation history information.

[0053] The electronic device (100) can receive conversation history information regarding the user's voice from the server (200) (S230-1). Additionally, the electronic device (100) can receive a response regarding the user's voice. Furthermore, the electronic device (100) can store conversation history information related to the received user's voice (S240-1) and can provide a response message regarding the user's voice based on the response received from the server (200).

[0055] FIG. 2b is a drawing for explaining a control method of a server (200) according to one embodiment of the present disclosure.

[0056] First, the server (200) can receive text corresponding to the user voice input into the electronic device (100) and conversation history information stored in the electronic device (100) from the electronic device (100) (S210-2). Meanwhile, in another embodiment, the server (200) can receive the user voice input into the electronic device (100) and conversation history information stored in the electronic device (100) from the electronic device (100). In this case, the server (200) can input the user voice into the first ASR module of the first conversation system to obtain text corresponding to the user voice.

[0057] And, the server (200) can perform language analysis on the text through the first conversation system based on conversation history information (S220-2). Specifically, the server (200) can perform language analysis based on the text and conversation history information to obtain the first language analysis result and the first language analysis reliability value, and can perform language analysis based only on the text to obtain the second language analysis result and the second language analysis reliability value.

[0058] If the text corresponding to the user voice currently received from the electronic device (100) is text corresponding to a voice related to the user voice previously entered into the electronic device (100), the server (200) can accurately identify the user's intent by performing language analysis on the currently received text using conversation history information related to the previously entered user voice, rather than performing language analysis based solely on the text. On the other hand, if the text corresponding to the user voice currently received from the electronic device (100) is a voice independent of the previously entered user voice, the server (200) can accurately identify the user's intent by ignoring conversation history information related to the previously entered user voice and performing language analysis based solely on the text. Accordingly, the server (200) can distinguish whether the text corresponding to the user voice entered into the electronic device (100) is text corresponding to a voice related to the previous user voice or text corresponding to an independent utterance unrelated to the previous user voice by performing language analysis on the text in different ways to obtain first and second language analysis reliability values ​​and comparing them. The specific process of performing language analysis will be explained in detail with reference to Fig. 5.

[0059] And, the server (200) can transmit the results of the performed language analysis to the electronic device (100) (S230-2). Specifically, if the first language analysis reliability value exceeds the second language analysis reliability value, the server (200) can transmit the first language analysis result to the electronic device (100).

[0060] Meanwhile, in another embodiment, if the first language analysis reliability value is less than or equal to the second language analysis reliability value, the server (200) may determine whether to transmit the language analysis result to the electronic device (100) based on information regarding the domain of the text among the second language analysis results. In one embodiment, if the domain of the text can be processed by the electronic device (100), the server (200) may transmit the second language analysis result to the electronic device (100). In another example, if the domain of the text cannot be processed by the electronic device (100), the server (200) may obtain response information regarding the user voice and conversation history information related to the user voice through the first conversation system based on the second language analysis result. Then, the server (200) may transmit the obtained response information regarding the user voice and conversation history information related to the user voice to the electronic device (100).

[0062] FIG. 3a is a diagram illustrating a process for determining whether an electronic device (100) will transmit a user voice to a server (200) based on a voice recognition reliability value, according to one embodiment of the present disclosure.

[0063] First, the electronic device (100) can input the input user voice into the second conversational system (S310). Then, the electronic device (100) can obtain text and voice recognition reliability values ​​corresponding to the user voice through the second ASR module of the second conversational system (S320). Specifically, the electronic device (100) can obtain a value indicating the degree of reliability with which the input user voice was recognized and converted into text through the second ASR module. The voice recognition reliability value may be 0 or 1, and the closer it is to 1, the higher the reliability with which the user voice was recognized and converted into text.

[0064] And, the electronic device (100) can identify whether the voice recognition reliability value of the user's voice exceeds a first threshold (S330). In one embodiment, if some of the text corresponding to the user's voice includes text that has not been learned by the language model of the second ASR module, the electronic device (100) can determine that the voice recognition reliability value of the user's voice is below the first threshold through the second ASR module. For example, when a user command such as "Give me directions to Seorae Village" is input, if the text "Seorae Village" has not been learned by the language model of the second ASR module, the electronic device (100) can determine that the voice recognition reliability value of the user's voice is below the first threshold through the second ASR module.

[0065] In one embodiment, when the voice recognition reliability value of the user's voice exceeds a first threshold, the electronic device (100) can obtain conversation history information and responses related to the user's voice through a second conversation system (S330). When the voice recognition reliability value of the user's voice is below the first threshold, the electronic device (100) can decide to transmit the user's voice to a server (200) (S350).

[0067] FIG. 3b is a diagram illustrating a process for determining whether an electronic device (100) will transmit a user voice to a server (200) based on a language analysis reliability value, according to one embodiment of the present disclosure.

[0068] The electronic device (100) can input the input user voice into the second conversational system (S410). Then, the electronic device (100) can obtain a language analysis reliability value for the text corresponding to the user voice through the second NLU model of the second conversational system (S420). That is, the electronic device (100) can obtain a value indicating the degree of reliability with which the text corresponding to the user voice was analyzed and understood through the second NLU module. The language analysis reliability value may be 0 or 1, and the closer it is to 1, the higher the reliability with which the text corresponding to the user voice was analyzed and understood.

[0069] And, the electronic device (100) can identify whether the language analysis reliability value for text corresponding to the user's voice exceeds a second threshold (S430). In one embodiment, if some of the text corresponding to the user's voice contains unlearned language, the electronic device (100) can determine through the second NLU module that the language analysis reliability value of the text corresponding to the user's voice is below the second threshold.

[0070] In one embodiment, when the language analysis reliability value exceeds a second threshold, the electronic device (100) can obtain response and conversation history information for the user's voice through a second conversation system (S440). If the language analysis reliability value is below the second threshold, the electronic device (100) can decide to transmit the user's voice to a server (200) (S450).

[0071] FIG. 3c is a diagram illustrating a process for determining whether to transmit a user voice to a server (200) based on information related to a domain of text corresponding to the user voice, according to one embodiment of the present disclosure.

[0072] The electronic device (100) can input the input user voice into the second conversational system (S510). Then, the electronic device (100) can obtain the domain of the text corresponding to the user voice and information related to the domain through the second NLU module (S520). The information related to the domain may include information regarding whether the domain is a dedicated domain of the first or second conversational system or a domain that can be processed by both conversational systems, and information regarding the processing capacity for the domain.

[0073] In one embodiment, when a user voice saying "Tell me the ASTC conference address" is input, the electronic device (100) can obtain information through the second NLU module about a domain called "location" and whether "location" is a dedicated domain of the first or second conversational system.

[0074] Meanwhile, the electronic device (100) can determine whether the second conversation system can process the user voice based on information related to the domain (S536). In one embodiment, if information is obtained that the 'address' is a domain exclusive to the first conversation system, the electronic device (100) can determine that the input user voice cannot be processed by the second conversation system. In another embodiment, if information is obtained that the 'address' is a domain exclusive to the second conversation system or a domain that can be processed by both the first conversation system and the second conversation system, the electronic device (100) can determine that the input user voice can be processed by the second conversation system.

[0075] If it is determined based on information related to the domain that the second conversation system cannot process the user voice, the electronic device (100) can transmit the user voice to the server (200) (S540). If it is determined based on information related to the domain that the second conversation system can process the user voice, the electronic device (100) does not transmit the user voice to the server (200) (S550), and can obtain conversation history information and responses related to the user voice through the second conversation system.

[0077] FIG. 3d is a flowchart for explaining the process of determining whether an electronic device (100) will transmit a user voice to a server (200) based on the state information of the electronic device (100), according to one embodiment of the present disclosure.

[0078] First, the electronic device (100) can receive user voice input (S610). Then, the electronic device (100) can decide whether to transmit the user voice to the server (200) based on the status information of the electronic device (100) (S620).

[0079] In one embodiment, the electronic device (100) can determine a conversation system to input user voice based on the state of communication connection with the server (200). For example, if communication connection with the server (200) is not established, the electronic device (100) can input the user voice into a second conversation system without transmitting it to the server (200) to obtain a response to the user voice and conversation history information related to the user voice.

[0080] Meanwhile, when a communication connection is established with the server (200) while additional user voice is being input, the electronic device (100) can determine whether to transmit the additional user voice to the server in the second conversational system. That is, when a communication connection is established with the server (200), the electronic device (100) can determine whether to transmit the user voice to the server (200) based on the reliability value and domain of the user voice obtained in the second conversational system.

[0081] In another embodiment, if at least some of the user voice and conversation history information is transmitted to the server (200) but the conversation history information related to the user voice is not received within a threshold time, the electronic device (100) may identify that the communication connection status with the server (200) is poor. Then, the electronic device (100) may decide not to transmit the user voice to the server (200) and input the user voice into a second conversation system to obtain conversation history information related to the user voice.

[0082] In another embodiment, the electronic device (100) may determine whether to transmit user voice to the server (200) based on the battery charge status of the electronic device (100). If the battery charge level of the electronic device (100) is below a threshold value, the electronic device (100) may decide not to transmit user voice to the server (200) in order to reduce battery consumption. Additionally, the electronic device (100) may input user voice into a second conversation system to obtain conversation history information related to user voice.

[0084] FIG. 3e is a diagram illustrating the process of an electronic device (100) selecting a conversational system to provide a response to a user's voice according to one embodiment of the present disclosure.

[0085] According to one embodiment of the present disclosure, when a conversational system to provide a response to a user's voice is selected, the electronic device (100) can determine the conversational system to provide a response to the user's voice as the selected conversational system. Meanwhile, in FIG. 3e, the electronic device (100) is implemented as a smartphone and the input unit (170) is implemented as a touch screen, but this is merely one embodiment. That is, the electronic device (100) can receive a user command to select a conversational system to provide a response to the user's voice through the input unit (170) implemented in various ways.

[0086] In one embodiment, as illustrated in FIG. 3e, the electronic device (100) may display a UI that allows selecting a conversation system to provide a response to the user's voice. When a UI (710) representing a first conversation system stored in a server (200) is selected via a touch screen, the electronic device (100) may transmit the user's voice to the server (200). Then, when a UI (720) representing a second conversation system built into the electronic device (100) is selected via a touch screen by the user, the electronic device (100) may input the user's voice into the second conversation system to obtain conversation history information related to the user's voice and a response to the user's voice.

[0088] FIG. 4a is a block diagram briefly illustrating the configuration of an electronic device according to one embodiment of the present disclosure. As illustrated in FIG. 4a, the electronic device (100) may include a communication unit (110), a microphone (120), a memory (130), and a processor (140). The configuration illustrated in FIG. 4 is an illustrative diagram for implementing embodiments of the present disclosure, and appropriate hardware and software configurations that are obvious to a person skilled in the art may additionally be included in the electronic device (100).

[0089] The communication unit (110) can communicate with an external device through various communication methods. Communication connection between the communication unit (110) and an external device may include communication via a third device (e.g., a repeater, a hub, an access point, a server, or a gateway).

[0090] Meanwhile, the communication unit (110) may include various communication modules to perform communication with an external device. For example, the communication unit (110) may include a wireless communication module, for example, LTE, LTE-A (LTE Advance), 5G (5 TH It may include a cellular communication module using at least one of Generation), CDMA (code division multiple access), WCDMA (wideband CDMA), UMTS (universal mobile telecommunications system), WiBro (Wireless Broadband), or GSM (Global System for Mobile Communications). As another example, the wireless communication module may include at least one of, for example, WiFi (wireless fidelity), Bluetooth, Bluetooth Low Energy (BLE), and Zigbee.

[0091] The microphone (120) is configured to receive user voice input and may be provided inside the electronic device (100), but this is merely one embodiment and may be provided outside the electronic device (100) and electrically connected to the electronic device (100) or connected via communication through the communication unit (110).

[0092] The memory (130) may store instructions or data related to at least one other component of the electronic device (100). An instruction is an action statement for the electronic device (100) in a programming language and is the smallest unit of a program that the electronic device (100) can directly execute. In one embodiment, the memory (130) may be implemented as non-volatile memory, volatile memory, flash memory, a hard disk drive (HDD), or a solid-state drive (SSD). The memory (130) is accessed by a processor (140), and read / write / modify / delete / update data by the processor (140) may be performed.

[0093] In the present disclosure, the term "memory" may include memory (130), ROM (not shown) or RAM (not shown) within a processor (140), or a memory card (not shown) mounted in an electronic device (100) (e.g., micro SD card, memory stick). Additionally, the memory (130) may store programs and data, etc., for configuring various screens to be displayed in the display area of ​​a display (150).

[0094] In particular, the memory (130) may store a program for executing a second conversation system. At this time, the second conversation system is a personalized program for providing various services to the electronic device (100). Additionally, the memory (130) may store a program for obtaining conversation history information related to the second conversation system and conversation history information related to the second conversation system. Additionally, the memory (130) may store a program for obtaining conversation history information related to the first conversation system from the server (200) and conversation history information related to the first conversation system. Furthermore, in one embodiment, the memory (130) may store conversation history information while the second conversation system is executing, and the conversation history information stored in the memory (130) may be deleted when the execution of the second conversation system ends.

[0095] Meanwhile, various software modules may be stored in the memory (130) as illustrated in FIG. 4a. Each software module may be controlled by the processor (140). Specifically, each software module stored in the memory (130) may be loaded into volatile memory (e.g., DRAM (Dynamic Random-Access Memory) and SRAM (Static RAM), etc.) by the control of the processor (140). Volatile memory may be implemented as a separate component that can be linked with the processor (140), but this is merely one embodiment, and volatile memory may also be implemented as a component of the processor (140) and included in the processor (140). Meanwhile, volatile memory refers to memory that requires a continuous power supply to maintain stored information.

[0096] The Voice Assistant Client module (310) can record and store user voice input through the microphone (120) under the control of the processor (140) in a first storage area of ​​volatile memory so that other software modules can process it. Meanwhile, in describing the present disclosure, the first to fifth storage areas of volatile memory are intended to explain the process of each software module accessing a storage area in which specific data is stored within the volatile memory to process specific data. Each storage area may be a separate storage area, but this is merely one embodiment, and some storage areas may be implemented in a form that includes some storage areas as components of another storage area.

[0097] The Coordinator module (320) can access a second storage area in the volatile memory and determine whether to transmit at least one of the user voice or text corresponding to the user voice to the server (200) based on the voice recognition reliability value and language analysis reliability value of the user voice recorded and stored in the second storage area. As another example, the process of the Coordinator module (320) accessing a third storage area in the volatile memory to determine whether to transmit at least one of the text corresponding to the user voice or the user voice to the server (200) based on the domain and language analysis reliability value of the text corresponding to the user voice will be explained in detail later by referring to the operation of the processor (140).

[0098] Meanwhile, the second ASR module (105-1) can access the first storage area of ​​the volatile memory, perform voice recognition on the user voice recorded and stored in the first storage area, and output text corresponding to the recognized user voice. Additionally, the second ASR module (105-1) can calculate a voice recognition reliability value for each user voice. The voice recognition reliability value is a numerical value that quantifies the degree of reliability with which the second ASR module (105-1) recognized the input user voice and converted it into text. Therefore, a high voice recognition reliability value may mean that the second ASR module (105-1) recognized the user voice more reliably and converted it into text corresponding to the user voice. Furthermore, the text corresponding to the user voice and the voice recognition reliability value output by the second ASR module (105-1) can be recorded and stored in the second storage area of ​​the volatile memory under the control of the processor (140).

[0099] In one embodiment of the present disclosure, if some of the text corresponding to the user voice is not learned by the language model (not shown) of the second ASR module (105-1), the second ASR module (105-1) can calculate a voice recognition reliability value below a first threshold. For example, when a user voice saying "Tell me the phone number of the CDE building in Seorae Village" is input through the microphone (120), if the text corresponding to the user voice includes "CDE building" which is not learned by the language model of the second ASR module (105-1), the second ASR module (105-1) can calculate a voice recognition reliability value for the input user voice below a first threshold.

[0100] Meanwhile, the second NLU (Natural Language Understanding) module (105-2) can access a second storage area of ​​volatile memory and determine the user's intent and parameters by using matching rules divided into a domain, an intent, and parameters (or slots) necessary to identify the intent, based on text corresponding to the user's voice recorded and stored in the second storage area. Specifically, one domain (e.g., alarm) may include multiple intents (e.g., alarm setting, alarm disabling), and one intent may include multiple parameters (e.g., time, number of repetitions, notification sound, etc.). The matching rules may be stored in an NLU Database (not shown). Furthermore, the second NLU module (105-2) can determine the user's intent by using linguistic features (e.g., grammatical elements) such as morphemes and phrases to identify the meaning of words extracted from user input, and matching the identified meaning of the words to the domain and intent. And, the language analysis reliability value output by the second NLU module (105-2) and the domain, intent, parameters, etc. of the text corresponding to the user's voice can be recorded and stored in the third storage area of ​​the volatile memory under the control of the processor (140).

[0101] For example, if the user voice converted into text through the second ASR module (105-2) is ‘give me directions to Seorae Village,’ the second NLU module (105-2) can determine the meaning of words such as ‘Seorae Village’ and ‘directions’ to obtain the intention that the user is requesting route guidance to the location of the place name ‘Seorae Village.’

[0102] Additionally, the second NLU module (105-2) can calculate a language analysis reliability value for the text corresponding to the user's voice obtained through the second ASR module. The language analysis reliability value is a numerical value indicating the degree of reliability with which the second NLU module (105-2) analyzed and understood the text corresponding to the user's voice. Therefore, a high language analysis reliability value may mean that the second NLU module (105-2) more reliably analyzed the text corresponding to the user's voice and grasped the user's intention.

[0103] The second DM (Dialogue Manager) module (105-3) can access the third storage area of ​​the volatile memory and determine whether the information regarding the user's intention recorded and stored in the third storage area is clear. Specifically, the second DM module (105-3) can determine whether the user's intention is clear based on whether the information of the parameters is sufficient. Furthermore, if the second DM module (105-3) can perform an action based on the intention and parameters identified through the second NLU module (105-2), it can generate a result (or response) of performing a task corresponding to the user input. The result and response corresponding to the user input output by the second DM module (105-3) can be stored in the fourth storage area of ​​the volatile memory under the control of the processor (140).

[0104] For example, if the second NLU module (105-2) identifies a user's intent to request route guidance to the location of the place name 'Seorae Village', the second DM module (105-3) can generate a response indicating that it will start route guidance to Seorae Village. Meanwhile, as another example, the second DM module (105-3) can determine whether the user's intent identified by the first NLU module (205-2) of the server (200) is clear, based on the language analysis results of the text corresponding to the user's voice received from the server (200) and stored in volatile memory. Since the process of determination has been described above, a redundant explanation will be omitted.

[0105] The second NLG (Natural Language Generator) module (105-4) can access the fourth storage area of ​​the volatile memory and change the response to the user voice recorded and stored in the fourth storage area into text form. The information changed into text form may be in the form of natural language utterance. For example, the second NLG module (105-4) can output the text "Starting directions to Seorae Village" based on a response meaning "starting directions to Seorae Village" recorded and stored in the volatile memory. The text output by the second NLG module (105-4) can be recorded and stored in the fifth storage area of ​​the volatile memory under the control of the processor (140). Then, the response to the user voice changed into text form can be displayed on the electronic device (100). As another example, a TTS (Text to Speech Synthesis) module (not shown) can access the fifth storage area of ​​the volatile memory and change the text recorded and stored in the fifth storage area into speech form and output it.

[0106] Meanwhile, the second conversational system data (105-7) may contain training data for training a software module included in the second conversational system (105).

[0107] The second Context Understanding module (or, second conversation history information understanding unit) (105-5) can identify, based on the input user voice, information about the work performed by the electronic device (100) before the user voice was input, information about the conversational context included in the user voice, and information about the state of the electronic device (100) at the time the user voice was input. For example, if the user voice is 'Give me directions to Seorae Village,' the second Context Understanding module (105-5) can identify, based on the user voice, information about the situation requesting directions to the location of the place name 'Seorae Village,' and information about whether an application capable of providing directions services is installed on the electronic device (100).

[0108] And, the second Context Generator module (or the second conversation history information generation unit) (105-6) can generate conversation history information based on information recorded and stored in volatile memory and store the generated conversation history information in conversation history data (330). The conversation history data (330) may be a database in which conversation history information is classified according to preset conditions (e.g., order of storage, etc.).

[0110] Meanwhile, the processor (140) is electrically connected to the memory (130) and can control the overall operation and function of the electronic device (100). In particular, the processor (140) can determine whether to transmit user voice input through the microphone (120) to the server (200) by executing instructions for a program to execute a second conversation system stored in the memory (130). Specifically, the processor (140) can input user voice into the second conversation system to obtain a reliability value or domain of the user voice, and determine whether to transmit the user voice to the server (200) based on the obtained reliability value or domain of the user voice.

[0111] In one embodiment, the processor (140) obtains a text corresponding to the user voice and a voice recognition reliability value of the user voice through the second ASR module of the second conversational system, and can determine whether to transmit the user voice to the server (200) based on the voice reliability value. Specifically, the processor (140) can determine whether to transmit the user voice to the server (200) depending on whether the voice recognition reliability value of the user voice obtained through the second ASR module exceeds a first threshold value.

[0112] In another embodiment, the processor (140) may obtain a language analysis reliability value of text corresponding to the user voice through the second NLU module of the second conversational system and determine whether to transmit the user voice to the server (200) based on the language analysis reliability value. Specifically, the processor (140) may determine whether to transmit the user voice to the server (200) depending on whether the language analysis reliability value of the text corresponding to the user voice obtained through the second NLU module exceeds a second threshold value.

[0113] In another embodiment, the processor (140) can obtain the domain of the text corresponding to the user voice and information related to the domain through the second NLU module. Then, the processor (140) can determine whether to transmit the user voice to the server (200) based on the information related to the domain. In one embodiment, if the processor (140) obtains information through the second NLU module that the domain of the text corresponding to the user voice cannot be processed by the second conversational system, the processor (140) can decide to transmit the user voice to the server (200).

[0114] In another embodiment, the processor (140) may determine whether to transmit user voice to the server (200) based on the state of the electronic device (100). In one embodiment, the processor (140) may determine whether to transmit user voice to the server (200) based on the state of communication connection between the electronic device (100) and the server (200) or the battery charge state of the electronic device (100).

[0115] In another embodiment, the processor (140) may determine whether to transmit the user voice to the server (200) according to the conversational system that will provide a response to the user voice selected through the input unit (170). For example, if the first conversational system is selected as the conversational system that will provide a response to the user voice through the input unit (170), the processor (140) may determine to transmit all input user voices to the server (200).

[0116] Meanwhile, in one embodiment, if it is decided to transmit the user voice to the server (200), the processor (140) may control the communication unit (110) to transmit at least a portion of the stored conversation history information to the server (200) including the first conversation system. Then, the processor (140) may receive conversation history information related to the user voice and a response to the user voice from the server (200) through the communication unit (110), and may store the received conversation history information in the memory (130). Then, the processor (140) may perform an action corresponding to the response to the received user voice or output a response message corresponding to the received response. For example, the processor (140) may control the display (150) to display a UI corresponding to the response to the user voice or execute an application corresponding to the response. In addition, the processor (140) may output a response message corresponding to the response to the user voice in the form of voice or display it in the form of text.

[0117] Meanwhile, in one embodiment, when additional user voice is input through the microphone (120), the processor (140) can input the additional user voice into the second conversation system and determine whether to transmit the user voice to the server (200). If it is determined not to transmit the additional user voice to the server (200), the processor (140) can use conversation history information related to the user voice among the conversation history information stored in the second conversation system to perform speech recognition or language analysis of the additional user voice, and obtain a response to the additional user voice and conversation history information related to the additional user voice.

[0118] Meanwhile, the artificial intelligence-related functions according to the present disclosure are operated through the processor (140) and memory (130).

[0119] The processor (140) may be composed of one or more processors. In this case, the one or more processors (140) may be a general-purpose processor such as a CPU (Central Processing Unit) or AP (Application Processor), a graphics-dedicated processor such as a GPU (graphics-processing Unit) or VPU (Visual Processing Unit), or an artificial intelligence-dedicated processor such as an NPU (Neural Processing Unit).

[0120] One or more processors control input data to be processed according to predefined operation rules or artificial intelligence models stored in memory (130). The predefined operation rules or artificial intelligence models are characterized by being created through learning.

[0121] Here, being created through learning means that a predefined rule of operation or an artificial intelligence model of desired characteristics is created by applying a learning algorithm to a number of learning data. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server / system.

[0122] An artificial intelligence model may be composed of multiple neural network layers. Each layer has multiple weight values ​​and performs operations of the layer through the operation of the multiple weights and the operation of the previous layer. Examples of neural networks include CNN (Convolutional Neural Network), DNN (Deep Neural Network), RNN (Recurrent Neural Network), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), BRDNN (Bidirectional Recurrent Deep Neural Network), and Deep Q-Networks; however, the neural networks in this disclosure are not limited to the aforementioned examples except where specified.

[0123] A learning algorithm is a method of training a specific target device (e.g., a robot) using a number of learning data to enable the target device to make decisions or predictions on its own. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, and the learning algorithms in this disclosure are not limited to the aforementioned examples except where specified.

[0125] FIG. 4b is a block diagram briefly illustrating the configuration of a server (200) according to one embodiment of the present disclosure. As illustrated in FIG. 4b, the server (200) may include a communication unit (210), a memory (220), and a processor (230). The configuration illustrated in FIG. 4b is an illustrative diagram for implementing embodiments of the present disclosure, and appropriate hardware and software configurations that are obvious to a person skilled in the art may additionally be included in the server (200).

[0126] The communication unit (210) can communicate with an external device (e.g., electronic device (100)) through various communication methods. Communication of the communication unit (210) with an external device may include communicating through a third device (e.g., a repeater, a hub, an access point, a gateway, etc.).

[0127] Meanwhile, the communication unit (210) may include various communication modules to perform communication with an external device. Since the communication modules have been described with reference to FIG. 4a, redundant descriptions will be omitted.

[0128] The memory (220) may be implemented as non-volatile memory, volatile memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD). The memory (220) is accessed by the processor (140), and data reading / writing / modification / deletion / updating, etc., by the processor (230) may be performed. In this disclosure, the term memory may include memory (220), ROM (not shown), RAM (not shown) within the processor (230), or a memory card (not shown) (e.g., micro SD card, memory stick) mounted on the server (200). Additionally, programs and data, etc., for configuring various screens to be displayed in the display area of ​​the display (150) may be stored in the memory (220).

[0129] And, memory (220) can store the first conversation system and at least one instruction.

[0130] In particular, the memory (220) can store a program for executing a first conversation system. The first conversation system is a personalized program for providing various services to the user. Additionally, the memory (220) can store a program for obtaining conversation history information related to the first conversation system and conversation history information related to the first conversation system. Additionally, the memory (220) can store a program for obtaining conversation history information related to the second conversation system from the electronic device (100) and conversation history information related to the second conversation system. Furthermore, in one embodiment, the memory (220) can store conversation history information while the first conversation system is executing, and the conversation history information stored in the memory (220) can be deleted when the execution of the first conversation system ends.

[0131] And, the memory (220) can store a first conversational system (205) containing various software modules as illustrated in FIG. 4b. Each software module can be controlled by a processor (230).

[0132] Meanwhile, the first ASR module (205-1), first NLU module (205-2), first DM module (205-3), first NLG module (205-4), first Context Understanding (or, first conversation history information understanding unit) (205-5), first Context Generator (or, first conversation history information generation unit) (205-6) and first conversation system data (205-7) of the first conversation system (205) stored in the server (200) can perform the same function as the modules corresponding to the second conversation system (105).

[0133] Meanwhile, the amount of data stored in the language model of the first ASR module (205-1) included in the first dialogue system (205) may be greater than the amount of data stored in the language model included in the second ASR module (105-1). Additionally, the first dialogue system (205) may have a greater amount of data that can be processed compared to the second dialogue system.

[0134] In addition, according to one embodiment of the present disclosure, a first NLU module (205-2) included in a first conversational system (205) can perform language analysis based on conversational history information on text corresponding to user voice input to the electronic device (100) received from the electronic device (100). Specifically, the first NLU module (205-2) can perform language analysis based on conversational history information to output a first language analysis result and a first language analysis reliability value, and can perform language analysis based only on text to output a second language analysis result and a second language analysis reliability value. In addition, the language analysis reliability value and language analysis result output by the first NLU module (205-2) can be recorded and stored in volatile memory under the control of a processor (230).

[0135] In one embodiment, the first NLU module (205-2) can identify the domain and intent of the text corresponding to the user's voice through conversation history information. If the text corresponding to the user's voice input into the electronic device (100) is 'Seorae Village' and the conversation history information received from the electronic device (100) includes situation information where the current user is requesting directions, the first NLU module (205-2) can identify through the conversation history information that the domain of the text is 'Location' related to directions, and the intent is to request directions to a destination called 'Seorae Village'. Therefore, when performing language analysis on the text using conversation history information, the first NLU module (205-2) can omit the process of classifying the domain and identifying the intent of the text. Furthermore, the first NLU module (205-2) can output the first language analysis result and output a first language analysis reliability value, which is a numerical value indicating the degree of reliability with which the text corresponding to the user's voice was analyzed and understood.

[0136] Meanwhile, while the first NLU module (205-2) performs language analysis on the text based on conversation history information, language analysis can be performed using only the text. That is, the first NLU module (205-2) can perform the operation of classifying the domain of the text and identifying the intent without utilizing conversation history information, thereby outputting the second language analysis result and the second language analysis reliability value. In the case of the above embodiment, since the text corresponding to the user's voice is associated with the conversation history information, the first language analysis reliability value may be higher than the second language analysis reliability value.

[0137] In another embodiment of the present disclosure, the first NLU module (205-2) can perform language analysis based solely on text and output a second language analysis result and a second language analysis reliability value. For example, if the text corresponding to the user's voice is "Not that, tell me the weather in Beijing" and the conversation history information is situational information where the user is currently requesting directions, the first NLU module (205-2) can ignore the conversation history information and perform language analysis on the text to determine that the domain of the text is "Weather" related to weather and the intention is to ask for the weather in the region called "Beijing." Then, the first NLU module (205-2) can output a second language analysis reliability value, which is a numerical value indicating the degree of reliability with which the text was analyzed and understood.

[0138] Meanwhile, the first NLU module (205-2) can perform language analysis on text based on conversation history information while performing language analysis based only on text, and output a first language analysis reliability value and a first language analysis result. The first NLU module (205-2) can identify that the domain obtainable through the conversation history information is 'Location'. However, the domain of the text corresponding to the user's voice is 'Weather', and since it is an utterance independent of the user's voice corresponding to the conversation history information, the second language analysis reliability value may be higher than the first language analysis reliability value.

[0139] Meanwhile, the processor (230) is electrically connected to the memory (220) and can control the overall operation and functions of the server (200). In particular, the processor (230) can execute instructions for a program to execute a first conversation system stored in the memory (220). In particular, the processor (230) can receive text corresponding to a user voice input to the electronic device (100) and conversation history information stored in the electronic device (100) from the electronic device (100) through the communication unit (210). However, this is merely one embodiment, and the processor (230) can receive a user voice input to the electronic device (100) through the communication unit (210). At this time, the processor (230) can obtain text corresponding to the user voice through the first conversation system.

[0140] And, the processor (230) can perform language analysis on text through the first conversation system based on conversation history information. Specifically, the processor (230) can perform language analysis based on text and conversation history information to obtain a first language analysis result and a first language analysis reliability value, and can perform language analysis only on text to obtain a second language analysis result and a second language analysis reliability value. By performing language analysis on text in a different way to obtain the first and second language analysis results, the processor (230) can distinguish whether the text corresponding to the user voice input into the electronic device (100) is text corresponding to a voice related to the previous user voice or text corresponding to an independent utterance unrelated to the previous user voice.

[0141] Additionally, the processor (230) can transmit one of the first language analysis result and the second language analysis result to the electronic device (100) based on the first language analysis reliability value and the second language analysis reliability value. In one embodiment, if the first language analysis reliability value is higher than the second language analysis reliability value, the processor (230) can control the communication unit (210) to transmit the first language analysis result to the electronic device (100). In another embodiment, if the second language analysis reliability value is higher than the first language analysis reliability value, the processor (230) can determine whether to transmit the text to the electronic device (100) based on information related to the domain of the text among the second language analysis results.

[0142] In one embodiment, the processor (230) can identify whether the domain of the text is a domain that the electronic device (100) can process. For example, if the domain of the text is a domain that can be processed using data stored in memory (220), the processor (230) can identify that the domain of the text is a domain that the electronic device (100) cannot process. In another example, if the domain of the text is a domain that can be processed only using data stored in the electronic device (100), the processor (230) can identify that the domain of the text is a domain that the electronic device (100) can process. In yet another example, the domain that the electronic device (100) can process or the domain that the server (200) can process can be determined by user input.

[0143] In one embodiment, if it is identified that the electronic device (100) can process the domain of the text, the processor (230) can control the communication unit (210) to transmit the second language analysis result to the electronic device (100). In another example, if it is identified that the electronic device (100) cannot process the domain of the text, the processor (230) can obtain a response to the user's voice and conversation history information related to the user's voice based on the second language analysis result. Then, the processor (230) can control the communication unit (210) to transmit the obtained response and conversation history information related to the user's voice to the electronic device (100).

[0144] Meanwhile, the artificial intelligence-related functions according to the present disclosure may be operated through the processor (230) and memory (220). Since the functions related to artificial intelligence (e.g., the learning process) have been described above, redundant descriptions will be omitted.

[0146] FIG. 5 is a flowchart for explaining a control method of a server (200) according to one embodiment of the present disclosure.

[0147] In one embodiment, the server (200) can receive text corresponding to a user voice input into the electronic device (100) and conversation history information stored in the electronic device (100) from the electronic device (100) (S710). Meanwhile, in one embodiment, when a user voice is input from the electronic device (100), the server (200) can obtain text corresponding to the user voice through the first ASR module of the first conversation system.

[0148] And, the server (200) can obtain a first language analysis result and a first language analysis reliability value by performing language analysis based on text and conversation history information, and obtain a second language analysis result and a second language analysis reliability value by performing language analysis based only on text (S720). Specifically, the server (200) can obtain a first language analysis result and a first language analysis reliability value by performing language analysis on the text using information on the domain and intent of the text corresponding to the user voice included in the conversation history information. That is, when the server (200) performs language analysis on the text, the additional classification process of the text domain and intent can be omitted. And, as another example, the server (200) can obtain a second language analysis result and a second language analysis reliability value by performing a language analysis process that classifies the domain and intent of the text itself without using conversation history information.

[0149] And, the server (200) can identify whether the first language analysis reliability value is higher than the second language analysis reliability value (S730). That is, by comparing the first and second language analysis reliability values, the server (200) can identify whether the text corresponding to the user voice input into the electronic device (100) is text corresponding to the voice related to the user previously input into the electronic device (100) or text corresponding to an independent utterance.

[0150] If the first language analysis reliability value is higher than the second language analysis reliability value, the server (200) can transmit the first language analysis result to the electronic device (S730-Y). That is, if the text corresponding to the user voice currently input into the electronic device (100) is identified as text related to the user voice previously input into the electronic device (100), the server (200) can transmit the first language analysis result to the electronic device (100). If the first language analysis reliability value is lower than the second language analysis reliability value, the server (200) can identify whether the domain of the text is processable by the electronic device (100) (S730-N).

[0151] If it is identified that the domain of the text can be processed by the electronic device (100), the server (200) can transmit the second language analysis result to the electronic device (S740). Then, if it is identified that the domain of the text cannot be processed by the electronic device (100), the server (200) can obtain a response to the user's voice and conversation history information related to the user's voice based on the second language analysis result through the first conversation system (S750). Then, the server (200) can transmit the response to the user's voice and conversation history information to the electronic device (100) (S760).

[0153] FIGS. 6a and 6b are drawings illustrating the operation between software modules of a conversational system included in an electronic device (100) and a server (200) as another embodiment of the present disclosure. That is, unlike what is shown in FIGS. 4a and 4b, each of the electronic device (100) and the server (200) can store software modules in memory (130, 220) as shown in FIGS. 6a and 6b. Meanwhile, descriptions that overlap with the content described with reference to FIGS. 4a and 4b will be omitted.

[0154] The second context sharer module (or second conversation history information sharing module) of the second conversation system (105) can share conversation history information with the first conversation system (205). In one embodiment, the second context sharer module (350) can output a signal requesting the server (200) to transmit a signal requesting conversation history information (or data) (330-2) stored in the server (200). Then, the processor (140) can control the communication unit (110) to transmit a signal requesting conversation history information (330-2) to the server (200). Then, the processor (140) can receive conversation history information (330-2) from the server (200) through the communication unit (110).

[0155] Additionally, the processor (140) may receive a signal requesting the sharing of conversation history information from the server (200) through the communication unit (110). When the processor (140) receives the signal requesting conversation history information through the communication unit (110), the second context sharer module (350) may output a signal requesting the transmission of conversation history information (or data) (330-1) to the server (200). Then, the processor (140) may control the communication unit (110) to transmit the conversation history information to the server (200) based on the output signal.

[0157] In another embodiment, the software module of the conversational system included in the electronic device (100) and the server (200) may be implemented as shown in FIG. 6b. Meanwhile, descriptions that overlap with those described with reference to FIG. 4a and FIG. 4b will be omitted.

[0158] The Execute Manager module (or execution manager module) (340) of the second conversation system (105) can be controlled to perform a function corresponding to a response to a user voice obtained from the first conversation system (105) or the second conversation system, which is recorded and stored in volatile memory. For example, when a response to a user voice saying "Give me directions to Seorae Village" is received, the Execute Manager module (340) can be controlled to execute a navigation application that gives directions to Seorae Village.

[0159] Meanwhile, in one embodiment, when a response to a user voice is received from the first conversation system (205) through the communication unit (110), the Execute Manager module (340) can transmit cache information of the response to the second conversation system (206).

[0161] FIG. 7 is a sequence diagram for explaining the operation of an electronic device and a server according to one embodiment of the present disclosure.

[0162] First, when a user voice is input (S810), the electronic device (100) can determine whether to transmit the user voice to the server (200) (S820). Specifically, the electronic device (100) inputs the user voice into a second conversational system to obtain a voice recognition reliability value, a domain, and a language analysis reliability value of the user voice, and can determine whether to transmit the user voice (10) to the server (200) based on the obtained voice recognition reliability value, language analysis reliability value, and domain of the user voice. Since the process of the electronic device (100) determining whether to transmit the user voice to the server (200) has been explained with reference to FIG. 2, a redundant explanation will be omitted.

[0163] If it is decided not to transmit the user voice to the server (200), the electronic device (100) can obtain a response to the user voice and conversation history information regarding the user voice through the second conversation system (S820-N).

[0164] Meanwhile, if it is decided to transmit the user voice to the server (200), the electronic device (100) can transmit the previously stored conversation history information and text corresponding to the user voice to the server (200) including the first conversation system (S820-Y). Meanwhile, in another embodiment, the electronic device (100) can transmit the user voice to the server (200). Then, the server (200) can perform language analysis on the text based on the conversation history information (S830). Specifically, the server (200) can perform language analysis based on the text and conversation history information to obtain a first language analysis result and a first language analysis reliability value, and perform language analysis based only on the text to obtain a second language analysis result and a second language analysis reliability value.

[0165] And, the server (200) can transmit the results of the language analysis to the electronic device (100) (S840). In one embodiment, if the first language analysis reliability value is higher than the second language analysis reliability value, the server (200) can transmit the first language analysis result to the electronic device (100). In another embodiment, if the second language analysis reliability value is higher than the first language analysis reliability value, the server (200) can identify whether the domain of the text is a domain that can be processed by the electronic device (100) based on information related to the domain of the text among the second language analysis results. For example, if the domain of the text is a domain that can be processed by the electronic device (100), the server (200) can transmit the second language analysis result to the electronic device (100). As another example, if the domain of the text is a domain that cannot be processed by the electronic device (100), the server (200) can obtain a response to the user's voice and conversation history information related to the user's voice based on the second language analysis result.

[0166] Meanwhile, the electronic device (100) can obtain a response to the user's voice and conversation history information related to the user's voice based on the results received from the server (200) (S850). Then, the electronic device (100) can provide voice to the user and store conversation history information (S860). That is, the electronic device (100) and the server (200) share conversation history information so that each conversation system can smoothly output a response to the user's voice.

[0168] Meanwhile, FIG. 8 is a sequence diagram for explaining the operation between an electronic device (100) and a server (200) according to another embodiment of the present disclosure.

[0169] When a user voice is input (S910), the electronic device (100) can determine whether to transmit the user voice to the server (200) (S920). If it is determined to transmit the user voice to the server (200), the electronic device (100) can transmit conversation history information and the user voice to the server (200) (S930). The server (200) can obtain a response to the user voice and conversation history information through the first conversation system (S940). Then, the server (200) transmits a response to the user voice to the electronic device (100) (S950), and the electronic device (100) can provide the received response (S960).

[0170] Then, when additional user voice is input, the electronic device (100) can determine whether to transmit the additional user voice to the server (200) (S965). If it is determined not to transmit the user voice to the server (200), the electronic device (100) can transmit a signal to the server (200) for requesting conversation history information regarding the user voice (S970). Upon receiving the signal, the server (200) can transmit conversation history information regarding the user voice to the electronic device (100) (S975). Upon receiving conversation history information regarding the user voice, the electronic device (100) can obtain a response and conversation history information regarding the additional user voice through the second conversation system (S980). Specifically, the electronic device (100) can obtain conversation history information regarding the additional user voice by inputting the conversation history information related to the user voice received from the server (200) and the additional user voice into the second conversation system. And, the electronic device (100) can provide a response to additional user voice and store conversation history information related to additional user voice (S985).

[0172] FIG. 9 is a drawing for explaining the operation of an electronic device (100) and a server (200) according to one embodiment of the present disclosure. Descriptions that overlap with FIG. 8 will be omitted.

[0173] When a user voice is input (S1010), the electronic device (100) inputs the user voice into a second conversational system (S1020) and can transmit it to a server (200) (S1030). Although FIG. 9 illustrates the electronic device (100) transmitting the user voice to the server (200) after inputting it into the second conversational system, the operations according to each step (S1030, S1040) can be performed simultaneously or within a preset time difference.

[0174] Meanwhile, the server (200) can obtain a response to the user's voice and conversation history information related to the user's voice through the second conversation system (S1040). Then, the server (200) can transmit the response to the user's voice to the electronic device (100) (S1050).

[0175] Meanwhile, the electronic device (100) can obtain cache information of a response to a received user voice (S1050). Specifically, it can obtain cache information that caches data related to a response to a user voice (S1060). Then, the electronic device (100) can provide a response to a user voice and store the cache information (S1070).

[0176] And, when additional user voice is input (S1080) and a second conversation system is determined to provide a response to the user voice from the user, the electronic device (100) can obtain a response to the additional user voice and conversation information related to the additional user voice based on cache information (S1090).

[0177] And, the electronic device (100) can provide a response to additional user voice and store conversation history information related to additional user voice (S1095).

[0179] FIG. 10 is a block diagram illustrating in detail the configuration of an electronic device (100) according to one embodiment of the present disclosure. As shown in FIG. 10, the electronic device (100) may include a communication unit (110), a microphone (120), a memory (130), a processor (140), a display (150), a speaker (160), and an input unit (170). Meanwhile, since the communication unit (110), microphone (120), memory (130), and processor (140) have been described in FIG. 4a, a redundant description will be omitted.

[0180] The display (150) can display various information under the control of the processor (140). In particular, the display (150) can display a UI corresponding to a response to a user's voice under the control of the processor (140).

[0181] Additionally, the display (150) may be implemented as a touch screen along with a touch panel. However, it is not limited to the implementation described above, and the display (150) may be implemented differently depending on the type of electronic device (100).

[0182] The speaker (160) is configured to output various audio data, as well as various notification sounds or voice messages, after various processing operations such as decoding, amplification, and noise filtering have been performed by an audio processing unit (not shown). In particular, the speaker (160) can output a response corresponding to the user's voice in the form of voice. However, the speaker (160) is merely one embodiment and may be implemented with other output terminals capable of outputting audio data.

[0183] The input unit (170) can receive various user inputs and transmit them to the processor (140). In particular, the input unit (170) may include a touch sensor, a (digital) pen sensor, a pressure sensor, a key, or a microphone. The touch sensor may use at least one of, for example, capacitive, pressure-sensitive, infrared, or ultrasonic methods. The (digital) pen sensor may, for example, be part of a touch panel or include a separate recognition sheet. The key may include, for example, a physical button, an optical key, or a keypad. Since the case where the input unit (170) is implemented as a touch sensor has been described with reference to FIG. 7, a redundant description will be omitted.

[0185] Meanwhile, various embodiments of the present disclosure are described with reference to the accompanying drawings. However, this is not intended to limit the technology described in the present disclosure to specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives to the embodiments of the present disclosure. In connection with the description of the drawings, similar reference numerals may be used for similar components.

[0186] In the present disclosure, expressions such as “have,” “may have,” “include,” or “may include” indicate the presence of such features (e.g., numerical values, functions, actions, or components such as parts) and do not exclude the presence of additional features.

[0187] In the present disclosure, expressions such as “A or B,” “at least one of A or / and B,” or “one or more of A or / and B” may include all possible combinations of items listed together. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” may refer to cases including (1) at least one A, (2) at least one B, or (3) both at least one A and at least one B.

[0188] Expressions such as "first," "second," "first," or "second" used in this disclosure may modify various components regardless of order and / or importance, and are used only to distinguish one component from another and do not limit said components.

[0189] When it is stated that a component (e.g., a first component) is "(operatively or communicatively) coupled with" or "connected to" another component (e.g., a second component), it should be understood that the component may be directly connected to the other component or connected through another component (e.g., a third component). On the other hand, when it is stated that a component (e.g., a first component) is "directly connected" or "directly connected" to another component (e.g., a second component), it may be understood that no other component (e.g., a third component) exists between the component and the other component.

[0190] As used in this disclosure, the expression “configured to” may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean only “specifically designed to” in hardware. Instead, in some contexts, the expression “device configured to” may mean that the device is “capable of” in conjunction with other devices or components. For example, the phrase “subprocessor configured to perform A, B, and C” may mean a dedicated processor for performing the said operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or an application processor) capable of performing said operations by executing one or more software programs stored in a memory device.

[0191] An electronic device according to various embodiments of the present disclosure may include, for example, at least one of a smartphone, a tablet PC, a desktop PC, a laptop PC, a netbook computer, a server, a PDA, a medical device, or a wearable device. In some embodiments, the electronic device may include, for example, at least one of a television, a refrigerator, an air conditioner, an air purifier, a set-top box, or a media box (e.g., Samsung HomeSync™, Apple TV™, or Google TV™).

[0192] Meanwhile, the term "user" may refer to a person using an electronic device or a device using an electronic device (e.g., an artificial intelligence electronic device). The present disclosure will be described in more detail below with reference to the drawings.

[0193] Various embodiments of the present disclosure may be implemented as software containing instructions stored on a machine-readable storage medium (e.g., a computer). The machine may include an electronic device (e.g., an electronic device (100)) according to the disclosed embodiments, which is a device capable of calling instructions stored on the storage medium and operating according to the called instructions. When the instructions are executed by a processor, the processor may perform a function corresponding to the instructions directly or using other components under the control of the processor. The instructions may include code generated or executed by a compiler or an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory storage medium" means that it does not contain a signal and is tangible, without distinguishing whether data is stored semi-permanently or temporarily on the storage medium. For example, the "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0194] According to one embodiment, the method according to the various embodiments disclosed herein may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed online in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily created in a storage medium such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0195] Each component (e.g., module or program) according to various embodiments may consist of a singular or multiple entities, and some of the aforementioned sub-components may be omitted, or other sub-components may be further included in various embodiments. Generally or additionally, some components (e.g., module or program) may be integrated into a single entity to perform the same or similar functions as those performed by each of the respective components prior to integration. The operations performed by the module, program, or other components according to various embodiments may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or other operations added. Explanation of the symbols

[0197] 110: Communications Unit 120: Microphone 130: Memory 140: Processor 150: Display 160: Speaker 170: Input section

Claims

Claim 1 An electronic device comprising: a communication unit including a circuit; a microphone; at least one instruction; a second conversational system for providing a response to a user voice input through the microphone; at least one memory for storing conversational history information; and a processor for executing the at least one instruction; wherein the processor, by executing the at least one instruction, determines whether to transmit the user voice to a server storing the first conversational system based on the voice recognition reliability value of the user voice, and if it is determined to transmit the user voice to the server, controls the communication unit to transmit at least a portion of the user voice and the stored conversational history information to the server, receives conversational history information related to the user voice from the server through the communication unit, and controls the received conversational history information to be stored in the memory; wherein the second conversational system includes a second ASR module (Automatic Speech Recognition), and the processor obtains text corresponding to the user voice and the voice recognition reliability value of the user voice through the second ASR module. Claim 2 delete Claim 3 An electronic device according to claim 1, wherein the processor determines to transmit at least one of the user voice or text corresponding to the user voice to the server when the voice recognition reliability value of the user voice is below a first threshold, and obtains a response to the user voice and conversation history information related to the user voice through the second conversation system when the voice recognition reliability value of the user voice exceeds the first threshold. Claim 4 In paragraph 3, the second conversational system includes a second NLU (Natural Language Understanding) module, and the processor, when the voice recognition reliability value of the user voice exceeds a first threshold, obtains a language analysis reliability value and a domain for text corresponding to the user voice through the second NLU module, and determines whether to transmit the text corresponding to the user voice to the server based on at least one of the language analysis reliability value and the domain for text corresponding to the user voice. Claim 5 An electronic device according to claim 4, wherein the processor determines to transmit the text corresponding to the user voice to the server when the language analysis reliability value for the text corresponding to the user voice is below a second threshold, and when the language analysis reliability value for the text corresponding to the user voice exceeds the second threshold, obtains a response to the user voice and conversation history information related to the user voice through the second conversation system. Claim 6 A server comprising: a communication unit including a circuit; at least one memory storing a first conversational system and at least one instruction; and a processor executing the at least one instruction; wherein the processor, by executing the at least one instruction, receives text corresponding to a user voice input to the electronic device and conversational history information stored in the electronic device from the electronic device through the communication unit, performs language analysis on the text through the first conversational system based on the conversational history information, and controls the communication unit to transmit the result of the performed language analysis to the electronic device, wherein the received text is processed through a second conversational system stored in the electronic device. Claim 7 In claim 6, the processor performs language analysis based on the text and conversation history information to obtain a first language analysis result and a first language analysis reliability value, performs language analysis based only on the text to obtain a second language analysis result and a second language analysis reliability value, and controls the communication unit to transmit one of the first language analysis result and the second language analysis result to the electronic device based on the first language analysis reliability value and the second language analysis reliability value. Claim 8 In claim 7, the processor controls the communication unit to transmit the first language analysis result to the electronic device when the first language analysis reliability value is higher than the second language analysis reliability value, and the server determines whether to transmit the text to the electronic device based on information related to the domain of the text among the second language analysis results when the second language analysis reliability value is higher than the first language analysis reliability value. Claim 9 In claim 8, the processor identifies whether the electronic device can process the domain of the text through information related to the domain of the text, and if it is identified that the electronic device can process the domain of the text, the server controls the communication unit to transmit the second language analysis result to the electronic device. Claim 10 In claim 9, the processor is a server that obtains a response to the user voice and conversation history information related to the user voice based on the second language analysis result when it is identified that the electronic device cannot process the domain of the text. Claim 11 A method for controlling an electronic device comprising a memory for storing conversation history information and a second conversation system, comprising: a step of determining whether to transmit the user voice to a server including a first conversation system based on a speech recognition reliability value of the input user voice; a step of transmitting at least a portion of the user voice and the stored conversation history information to the server if it is determined to transmit the user voice to the server; a step of receiving conversation history information related to the user voice from the server; and a step of storing the received conversation history information; wherein the second conversation system includes a second ASR module (Automatic Speech Recognition), and the step of determining comprises a step of obtaining text corresponding to the user voice and a speech recognition reliability value of the user voice through the second ASR module. Claim 12 delete Claim 13 A method for controlling an electronic device according to claim 11, wherein the determining step comprises: determining to transmit at least one of the user voice or text corresponding to the user voice to the server when the voice recognition reliability value of the user voice is below a first threshold value; and when the voice recognition reliability value of the user voice exceeds the first threshold value, obtaining a response to the user voice and conversation history information related to the user voice through the second conversation system. Claim 14 In claim 13, the second conversational system includes a second Natural Language Understanding (NLU) module, and the determining step comprises: a step of obtaining a language analysis reliability value and a domain for text corresponding to the user voice through the second NLU module when the voice recognition reliability value of the user voice exceeds a first threshold value; and a step of determining whether to transmit the text corresponding to the user voice to the server based on at least one of the language analysis reliability value and the domain for text corresponding to the user voice. Claim 15 A method for controlling an electronic device according to claim 14, wherein the determining step comprises: determining to transmit the text corresponding to the user voice to the server when the language analysis reliability value for the text corresponding to the user voice is less than or equal to a second threshold value, and when the language analysis reliability value for the text corresponding to the user voice exceeds the second threshold value, obtaining a response to the user voice and conversation history information related to the user voice through the second conversation system. Claim 16 A method for controlling a server comprising at least one memory for storing a first conversational system, comprising: receiving from an electronic device text corresponding to a user voice input to the electronic device and conversational history information stored in the electronic device; performing language analysis on the text through the first conversational system based on the conversational history information; and transmitting a result according to the performed language analysis to the electronic device; wherein the received text is processed through a second conversational system stored in the electronic device. Claim 17 A method for controlling a server according to claim 16, wherein the steps performed above include: a step of performing language analysis based on the text and conversation history information to obtain a first language analysis result and a first language analysis reliability value, and performing language analysis based only on the text to obtain a second language analysis result and a second language analysis reliability value; and a step of transmitting one of the first language analysis result and the second language analysis result to the electronic device based on the first language analysis reliability value and the second language analysis reliability value. Claim 18 In claim 17, the transmitting step comprises: transmitting the first language analysis result to the electronic device when the first language analysis reliability value is higher than the second language analysis reliability value; and determining whether to transmit the text to the electronic device based on information related to the domain of the text among the second language analysis results when the second language analysis reliability value is higher than the first language analysis reliability value; a control method of a server. Claim 19 A method for controlling a server according to claim 18, wherein the transmitting step comprises: a step of identifying whether the electronic device can process the domain of the text through information related to the domain of the text; and a step of transmitting the second language analysis result to the electronic device if it is identified that the electronic device can process the domain of the text. Claim 20 A method for controlling a server according to claim 19, wherein the transmitting step comprises the step of obtaining a response to the user voice and conversation history information related to the user voice based on the second language analysis result when it is identified that the domain of the text cannot be processed by the electronic device.

Citation Information

Patent Citations

  • Electronic apparatus, controlling method of thereof and non-transitory computer readable recording medium

    KR1020180108400A

  • Device and method for recognizing wake-up word using server recognition result

    KR1020190064384A