Electronic device for managing multiple intelligent agents and operating method thereof
By integrating communication circuits, processors, and memory into electronic devices, the processor receives voice data and generates text data to identify the device, thus solving the problem of interlocking incompatibility between heterogeneous voice intelligent agents and achieving cross-agent device control compatibility.
Patent Information
- Application Number
- CN202010739478.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-05
- Filing Date
- 2020-07-28
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2040-07-28
AI Technical Summary
When multiple voice-based smart agents exist in a user's living environment, the interlocking and incompatibility between these smart agents makes it impossible to effectively control heterogeneous devices.
By integrating communication circuits, processors, and memory into electronic devices, the processor can receive voice data, generate text data, analyze the text data to identify the device, and determine whether to send the voice data to an external server that supports heterogeneous voices based on the stored information, thereby enabling cross-agent control.
It enables effective control of intelligent agent devices that support heterogeneous voices, even when voice commands are input, thus improving the compatibility and consistency of device control.
Smart Images

Figure CN112331196B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an electronic device for managing multiple intelligent agents and an operating method thereof. Background Art
[0002] Services for electronic devices using intelligent agents have become widespread. Intelligent agents can provide integrated functionality to users by controlling several external devices functionally connected to the electronic device. Electronic devices can provide voice-based intelligent agent services, allowing users of the electronic device to perform various functions using their voice.
[0003] With the beginning of the application of the Internet of Things (IoT), in which user devices are connected to each other through wired / wireless networks to share information in the user's living environment, it is possible to use various electronic devices such as televisions, refrigerators, etc. to perform voice recognition for other external devices connected through the network.
[0004] Electronic devices that provide voice-based intelligent agent functionality within a user's living environment are increasing. When multiple voice-based intelligent agents exist, electronic devices supporting each of these agents can coexist in the user's living environment. To control a specific device, a user must use the voice-based intelligent agent supported by that device. In other words, interlocking between dissimilar voice-based intelligent agents may be infeasible. Summary of the Invention
[0005] Embodiments of the present disclosure provide an electronic device and an operating method thereof, wherein the electronic device and the operating method thereof can control a device intended to be controlled by a user by processing a received voice command (e.g., speech) to identify the device and sending the voice data to an intelligent agent, wherein the intelligent agent is capable of controlling the identified device based on information about the identified device. For example, even in the case where a voice command is input to a voice-based intelligent agent, the control command can be transmitted to a device that supports a heterogeneous voice-based intelligent agent.
[0006] According to various exemplary embodiments of the present disclosure, an electronic device configured to support a first voice-based intelligent agent may include: a communication circuit; a processor operably connected to the communication circuit; and a memory operably connected to the processor. According to various embodiments, the memory stores instructions that, when executed, cause the processor to control the electronic device to: receive voice data from a user terminal via the communication circuit, process the voice data to generate text data; identify a device to be controlled by analyzing the text data, receive information about an intelligent agent supported by the identified device from a first external server via the communication circuit, and determine whether to send the voice data to a second external server supporting a second voice-based intelligent agent based on the information about the intelligent agent supported by the identified device.
[0007] According to various exemplary embodiments of the present disclosure, an electronic device configured to support a first voice-based intelligent agent may include: a communication circuit; a processor operably connected to the communication circuit; and a memory operably connected to the processor. According to various exemplary embodiments, the memory stores information about at least one device registered in an account and information about an intelligent agent supported by the at least one device, and the memory stores instructions that, when executed, cause the processor to control the electronic device to: receive voice data from a user terminal via the communication circuit; process the voice data to generate text data; identify a device to be controlled by analyzing the text data; identify information about an intelligent agent supported by the identified device based on the information about the at least one device registered in the account and the information about the intelligent agents supported by the at least one device stored in the memory; and determine, based on the identified information, whether to send the voice data to an external server supporting a second voice-based intelligent agent via the communication circuit.
[0008] According to various exemplary embodiments of the present disclosure, an electronic device configured to support a first voice-based intelligent agent may include: a communication circuit; a microphone; a processor operably connected to the communication circuit and the microphone; and a memory operably connected to the processor. According to various embodiments, the memory is configured to store information about at least one device registered in an account and information about intelligent agents supported by the at least one device, and the memory stores instructions that, when executed, cause the processor to control the electronic device to: receive voice data through the microphone; process the voice data to generate text data; identify a device to be controlled by analyzing the text data; identify information about an intelligent agent supported by the identified device based on the information about the at least one device registered in the account and the information about the intelligent agents supported by the at least one device stored in the memory; and determine whether to send the voice data to an external electronic device supporting a second voice-based intelligent agent through the communication circuit based on the information about the intelligent agents supported by the identified device.
[0009] According to various exemplary embodiments of the present disclosure, an electronic device can control a device intended to be controlled by a user by processing a received voice command (e.g., speech) to identify the device and sending the voice data to an intelligent agent, which can control the identified device based on information about the identified device. For example, even when a voice command is input to a voice-based intelligent agent, the control command can be transmitted to a device that supports a heterogeneous voice-based intelligent agent. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:
[0011] Figure 1 is a block diagram illustrating an example integrated intelligent system according to an embodiment;
[0012] Figure 2 is a diagram showing an example form of storing concept-action relationship information in a database according to an embodiment;
[0013] Figure 3 is a diagram illustrating an example user terminal according to an embodiment, the example user terminal displaying a screen for processing a voice input received through a smart application (app);
[0014] Figure 4 is a diagram illustrating an example environment for supporting various voice-based intelligent agents according to an embodiment;
[0015] Figure 5 is a diagram illustrating an example environment for controlling an electronic device supporting a second voice-based intelligent agent using a user terminal supporting a first voice-based intelligent agent according to an embodiment;
[0016] Figure 6 is a block diagram illustrating an example operating environment between a user terminal, a first smart server, and a second smart server according to an embodiment;
[0017] Figure 7 is a block diagram illustrating an example operating environment between a first user terminal and a second user terminal according to an embodiment;
[0018] Figure 8 is a flow chart illustrating an example operation of a first intelligent server according to an embodiment;
[0019] Figure 9 is a flow chart illustrating an example operation of a first smart server according to an embodiment;
[0020] Figure 10 is a signal flow diagram illustrating example operations between a user terminal, a first smart server, an IoT server, and a second smart server according to an embodiment;
[0021] Figure 11 is a signal flow diagram illustrating an example operation between a user terminal, a first smart server, and a second smart server according to an embodiment;
[0022] Figure 12 is a signal flow diagram illustrating example operations between a first user terminal and a second user terminal according to an embodiment;
[0023] Figure 13 is a diagram illustrating an example environment for controlling an electronic device supporting a second voice-based intelligent agent through a user terminal supporting a first voice-based intelligent agent according to an embodiment;
[0024] Figure 14 is a diagram illustrating an example environment for controlling an electronic device supporting a second voice-based intelligent agent through a user terminal supporting a first voice-based intelligent agent according to an embodiment;
[0025] Figure 15 is a diagram illustrating an example environment for controlling an electronic device supporting a second voice-based intelligent agent through a user terminal supporting a first voice-based intelligent agent according to an embodiment; and
[0026] Figure 16 is a block diagram illustrating an example electronic device in a network environment according to an embodiment.
[0027] Regarding the description of the drawings, the same or similar reference numerals may be used for the same or similar constituent elements. DETAILED DESCRIPTION
[0028] Figure 1 is a block diagram illustrating an example integrated intelligent system according to an embodiment.
[0029] Reference Figure 1 , the integrated intelligent system according to the embodiment may include a user terminal 100 , an intelligent server 200 , and a service server 300 .
[0030] The user terminal 100 according to an embodiment may be a terminal device (or electronic device) capable of connecting to the Internet, and may include, for example, a mobile phone, a smart phone, a personal digital assistant (PDA), a laptop computer, a television, a white appliance, a wearable device, a head-mounted device (HMD), or a smart speaker.
[0031] According to an embodiment, the user terminal 100 may include a communication interface 110, a microphone 120, a speaker 130, a display 140, a memory 150, or a processor 160. The listed elements may be operatively or electrically connected to each other.
[0032] The communication interface 110 according to an embodiment can be connected to an external device and configured to send and receive data. The microphone 120 according to an embodiment can receive sound (e.g., user voice) and convert it into an electrical signal. The speaker 130 according to an embodiment can output an electrical signal in the form of sound (e.g., voice). The display 140 according to an embodiment can be configured to display an image or video. The display 140 according to an embodiment can display a graphical user interface (GUI) of the running app (or application).
[0033] The memory 150 according to an embodiment stores a client module 151, a software development kit (SDK) 153, and a plurality of apps 155. The client module 151 and the SDK 153 may configure a framework (or program) for performing general functions. In addition, the client module 151 or the SDK 153 may configure a framework for processing voice input.
[0034] Multiple apps 150 may be programs for performing predetermined functions. Depending on the embodiment, multiple apps 155 may include a first app 155_1 and a second app 155_3. Depending on the embodiment, each of the multiple apps 155 may include multiple operations for performing the predetermined functions. For example, these apps may include an alarm clock app, a messaging app, and / or a calendar app. Depending on the embodiment, the multiple apps 155 may be executed by the processor 160 to sequentially perform at least some of the multiple operations.
[0035] The processor 160 according to an embodiment may control the overall operation of the user terminal 100. For example, the processor 160 may be electrically connected to the communication interface 110, the microphone 120, the speaker 130, and the display 140 to perform predetermined operations.
[0036] The processor 160 according to an embodiment can execute a predetermined function by running a program stored in the memory 150. For example, the processor 160 can execute the following operation for processing a voice input by running at least one of the client module 151 or the SDK 153. The processor 160 can control the operation of, for example, a plurality of apps 155 through the SDK 153. The following operations as operations of the client module 151 or the SDK 153 can be executed by the processor 160.
[0037] The client module 151 according to an embodiment may receive voice input. For example, the client module 151 may receive a voice signal corresponding to a user's voice detected by the microphone 120. The client module 151 may transmit the received voice input to the intelligent server 200. The client module 151 may transmit state information of the user terminal 100 and the received voice input to the intelligent server 200. The state information may be, for example, execution state information of an app.
[0038] The client module 151 according to an embodiment may receive a result corresponding to the received voice input. For example, if the smart module 200 obtains a result corresponding to the received voice input, the client module 151 may receive the result corresponding to the received voice input. The client module 151 may display the received result on the display 140.
[0039] The client module 151 according to an embodiment may receive a plan corresponding to the received voice input. The client module 151 may display the results obtained by executing multiple operations of the app on the display 140 according to the plan. The client module 151 may sequentially display the execution results of multiple operations on the display, for example. In another example, the user terminal 100 may display only some of the results of multiple operations (only the result of the last operation) on the display.
[0040] According to an embodiment, the client module 151 may receive a request for obtaining information required for obtaining a result corresponding to the voice input from the intelligent server 200. According to an embodiment, the client module 151 may send the required information to the intelligent server 200 in response to the request.
[0041] The client module 151 according to an embodiment may transmit result information of the execution of the plurality of operations according to the plan to the smart server 200. The smart server 200 may use the result information to recognize that the received voice input is correctly processed.
[0042] The client module 151 according to an embodiment may include a voice recognition module. According to an embodiment, the client module 151 may recognize a voice input for executing a limited function through the voice recognition module. For example, the client module 151 may execute a smart app for processing voice input to perform an organic operation through a predetermined input (e.g., wake up!).
[0043] According to an embodiment, the intelligent server 200 may receive information related to a user voice input from the user terminal 100 via a communication network. According to an embodiment, the intelligent server 200 may convert the data related to the received voice input into text data. According to an embodiment, the intelligent server 200 may generate a plan for executing a task corresponding to the user voice input based on the text data.
[0044] According to an embodiment, the plan may be generated by an artificial intelligence (AI) system. The intelligent system may be a rule-based system, a neural network-based system (e.g., a feedforward neural network (FNN) or a recurrent neural network (RNN)). Alternatively, the intelligent system may be a combination of the above systems or an intelligent system different from the above systems. According to an embodiment, the plan may be selected from a combination of predefined plans or generated in real time in response to a user request. For example, the intelligent system may select at least one plan from a plurality of predetermined plans.
[0045] According to an embodiment, the intelligent server 200 may transmit the result of the generated plan to the user terminal 100, or may transmit the generated plan to the user terminal 100. According to an embodiment, the user terminal 200 may display the result of the plan on a display. According to an embodiment, the user terminal 100 may display the result of the operation according to the plan on a display.
[0046] The intelligent server 200 according to an embodiment may include a front end 210 , a natural language platform 220 , a capsule database (DB) 230 , an execution engine 240 , and a terminal user interface 250 , a management platform 260 , a big data platform 270 , or an analysis platform 280 .
[0047] According to an embodiment, the front end 210 may receive a voice input received from the user terminal 100. The front end 210 may send a response to the voice input.
[0048] According to an embodiment, the natural language platform 220 may include an automatic speech recognition module (ASR module) 221, a natural language understanding (NLU) module 223, a planner module 225, a natural language generator (NLG) module 227, or a text to speech (TTS) module 229.
[0049] The automatic speech recognition module 221 according to an embodiment can convert the speech input received from the user terminal 100 into text data. The natural language understanding module 223 according to an embodiment can detect the user's intention based on the text data of the speech input. For example, the natural language understanding module 223 can detect the user's intention by performing grammatical analysis or semantic analysis. The natural language understanding module 223 according to an embodiment can detect the meaning of the word extracted from the speech input based on the language characteristics of morphemes or phrases (e.g., grammatical elements), and match the meaning and intention of the detected word to determine the user's intention.
[0050] According to an embodiment, the planner module 225 may generate a plan based on the intent and parameters determined by the natural language understanding module 223. According to an embodiment, the planner module 225 may determine multiple domains required to perform the task based on the determined intent. The planner module 225 may determine multiple operations included in the multiple domains determined based on the intent. According to an embodiment, the planner module 225 may determine the parameters required to perform the multiple determined operations or the result values output by performing the multiple operations. The parameters and result values may be defined by concepts of predetermined types (or classes). According to an embodiment, the plan may include multiple operations determined by the user's intent and multiple concepts. The planner module 225 may gradually (or hierarchically) determine the relationship between the multiple operations and the multiple concepts. For example, the planner module 225 may determine the execution order of the multiple operations determined based on the user's intent based on the multiple concepts. In other words, the planner module 225 may determine the execution order of the multiple operations based on the parameters required to perform the multiple operations and the results output by performing the multiple operations. Thus, the planner module 225 may generate a plan that includes information about the relationship (ontology) between the multiple operations and the multiple concepts. The planner module 225 may generate a plan based on information stored in the capsule database 230 , which stores a set of relationships between concepts and operations.
[0051] The natural language generator module 227 according to an embodiment can convert predetermined information into text form. The information converted into text form can be in the form of natural language speech. The text-to-speech module 229 can convert the information in text form into information in speech form.
[0052] According to an embodiment, some or all functions of the natural language platform 220 may also be implemented by the user terminal 100 .
[0053] The capsule database 230 may store information about the relationship between multiple concepts and operations corresponding to multiple domains. Capsules according to embodiments may include multiple operation objects (action objects or action information) and concept objects (or concept information). Capsule database 230 may store multiple capsules in the form of a concept-action network (CAN). Capsules may be stored in a function registry included in capsule database 230.
[0054] The capsule database 230 may include a policy registry that stores policy information required for determining a plan corresponding to a voice input. When multiple plans exist corresponding to a voice input, the policy information may include reference information for determining a plan. Depending on the embodiment, the capsule database 230 may include a subsequent registry that stores the following actions to suggest subsequent actions to the user in predetermined situations. The subsequent actions may include, for example, the next voice message. Depending on the embodiment, the capsule database 230 may include a layout registry that stores layout information corresponding to information output by the user terminal 100. Depending on the embodiment, the capsule database 230 may include a vocabulary registry that stores vocabulary information included in capsule information. Depending on the embodiment, the capsule database 230 may include a conversation registry that stores conversation (or interaction) information with the user. Objects stored in the capsule database 230 may be updated via developer tools. The developer tools may include a function editor for updating, for example, operation objects or concept objects. The developer tools may include a vocabulary editor for updating the vocabulary. The developer tools may include a policy editor for generating and registering policies to determine a plan. The developer tools may include a conversation editor for generating conversations with the user. The developer tool may include a subsequent editor for activating subsequent goals and editing the next voice prompt. The subsequent goal may be determined based on the current goal, user preferences, or environmental conditions. According to an embodiment, the capsule database 230 may be implemented within the user terminal 100.
[0055] The execution engine 240 according to an embodiment can obtain a result based on the generated plan. The terminal user interface 250 can send the obtained result to the user terminal 100. Therefore, the user terminal 100 can receive the result and provide the received result to the user. The management platform 260 according to an embodiment can manage the information used by the intelligent server 200. The big data platform 270 according to an embodiment can collect user data. The analysis platform 280 according to an embodiment can manage the quality of service (QoS) of the intelligent server 200. For example, the analysis platform 280 can manage the elements and processing speed (or efficiency) of the intelligent server 200.
[0056] The service server 300 according to an embodiment may provide a predetermined service (e.g., food ordering or hotel reservation) to the user terminal 100. According to an embodiment, the service server 300 may be a server operated by a third party. The service server 300 according to an embodiment may provide the intelligent server 200 with information for generating a plan corresponding to the received voice input. The provided information may be stored in the capsule database 230. Furthermore, the service server 300 may provide the intelligent server 200 with information on the results of the plan.
[0057] In the integrated intelligent system 10, the user terminal 100 can provide various intelligent services to the user in response to user input. The user input can include, for example, input through a physical button, touch input, or voice input.
[0058] According to an embodiment, the user terminal 100 may provide a voice recognition service through a smart app (or voice recognition app) stored in the user terminal 100. In this case, for example, the user terminal 100 may recognize a user's voice (speech) or voice input received through a microphone and provide a service corresponding to the recognized voice input to the user.
[0059] According to an embodiment, the user terminal 100 can perform a predetermined operation together with the intelligent server and / or service server based on the received voice input. For example, the user terminal 100 can execute an app corresponding to the received voice input and perform a predetermined operation through the executed app.
[0060] According to an embodiment, when the user terminal 100 provides a service together with the intelligent server 200 and / or the service server, the user terminal can detect the user's voice through the microphone 120 and generate a signal (or voice data) corresponding to the detected user's voice. The user terminal can send the voice data to the intelligent server 200 through the communication interface 110.
[0061] According to an embodiment, the intelligent server 200 can generate a plan for performing a task corresponding to the voice input, or the result of an operation according to the plan, in response to the voice input received from the user terminal 100. The plan may include, for example, multiple operations for performing the task corresponding to the user's voice input and multiple concepts related to the multiple operations. These concepts can be parameters input to perform the multiple operations, or can be defined as result values output by performing the multiple operations. The plan may include the relationship between the multiple operations and the multiple concepts.
[0062] The user terminal 100 according to the embodiment may receive the response through the communication interface 110. The user terminal 100 may output a voice signal generated by the user terminal 100 to the outside through the speaker 130 or output an image generated by the user terminal 100 to the outside through the display 140.
[0063] Figure 2 is a diagram illustrating an example form of storing relationships between concepts and operations in a database according to various embodiments.
[0064] The capsule database (e.g., capsule database 230) of the intelligent server 200 may store capsules in the form of concept-action networks (CANs). The capsule database may store operations for processing tasks corresponding to user voice input and parameters required for the operations in the form of concept-action networks (CANs).
[0065] The capsule database may store multiple capsules (Capsule A 401 and Capsule B 404) corresponding to multiple domains (e.g., applications). Depending on the embodiment, one capsule (e.g., Capsule A 401) may correspond to one domain (e.g., location (geographic location) or application). Furthermore, one capsule may correspond to at least one service provider (e.g., CP#1 402 and CP#2 403) for performing the function of the domain associated with the capsule. Depending on the embodiment, one capsule may include one or more operations 410 and one or more concepts 420 for performing a predetermined function.
[0066] The natural language platform 220 can generate a plan for performing a task corresponding to the received speech input using capsules stored in the capsule database. For example, the natural language platform's planner module 225 can generate a plan using capsules stored in the capsule database. For example, plan 407 can be generated using operations 4011 and 4013 and concepts 4012 and 4014 of capsule A 401, and operation 4041 and concept 4042 of capsule B 404.
[0067] Figure 3 is a diagram illustrating example screens on which a user terminal processes a voice input received through a smart app according to various embodiments.
[0068] The user terminal 100 may execute a smart app to process user input through the smart server 200 .
[0069] According to an embodiment, in screen 310, when a predetermined voice input (e.g., wake up!) is recognized or input is received through a hardware key (e.g., a dedicated hardware key), the user terminal 100 may execute a smart app for processing the voice input. The user terminal 100 may execute the smart app in a state where, for example, a schedule app is executed. According to an embodiment, the user terminal 220 may display an object 311 (e.g., an icon) corresponding to the smart app on the display 140. According to an embodiment, the user terminal 100 may receive a voice input through the user's voice. For example, the user terminal 100 may receive a voice input "Let me know my schedule for this week!" According to an embodiment, the user terminal 100 may display a user interface (UI) 313 (e.g., an input window) on the display in which the smart app displays text data of the received voice input.
[0070] According to an embodiment, the user terminal 100 may display a result corresponding to the received voice input on the display in screen 320. For example, the user terminal 100 may receive a plan corresponding to the received user input and display "this week's schedule" on the display according to the plan.
[0071] Figure 4 is a diagram illustrating an example environment for supporting various voice-based intelligent agents according to an embodiment.
[0072] Reference Figure 4 According to various embodiments, devices supporting various voice-based intelligent agents can coexist in a user's living environment, for example. The voice-based intelligent agent can be, for example, general-purpose terminology call software that performs a series of operations with autonomy and a predetermined level of independence based on voice recognition. The voice-based intelligent agent can be represented as, for example, a voice-based intelligent assistant, an AI voice assistant, or an intelligent virtual assistant.
[0073] According to an embodiment of the present disclosure, the first user terminal 100a, the second user terminal 100b, the first electronic device 41, the second electronic device 43 and the third electronic device 45 can correspond to, for example, various Internet-connectable terminal devices, such as but not limited to, voice recognition speakers (or AI speakers), smart phones, personal digital assistants (PDAs), notebook computers and electronic devices applying IoT technology (for example, smart TVs, smart refrigerators, smart lights or smart air purifiers).
[0074] According to an embodiment of the present disclosure, the first user terminal 100a, the second user terminal 100b, the first electronic device 41, the second electronic device 43 and the third electronic device 45 can provide the services required by the user through the corresponding apps (or applications) stored therein. For example, the first user terminal 100a and the second user terminal 100b can receive user input for controlling external electronic devices through a voice-based intelligent agent service (or voice recognition application) stored therein. User input can be received through, for example, physical buttons, touchpads, voice input or remote input. For example, the first user terminal 100a and the second user terminal 100b can receive the user's speech and can generate voice data corresponding to the user's speech. The first user terminal 100a can send the generated voice data to the first intelligent server. The second user terminal 100b can send the generated voice data to the second intelligent server. The first user terminal 100a and the second user terminal 100b can use a cellular network or a local area network (e.g., Wi-Fi or LAN) to send the generated voice data to the first intelligent server and the second intelligent server.
[0075] For ease of explanation, it is assumed by way of non-limiting example that the first user terminal 100a is a user terminal that supports a first voice-based intelligent agent, and the second user terminal 100b is a user terminal that supports a second voice-based intelligent agent that is different from the first voice-based intelligent agent. It is assumed that the first electronic device 41 supports the first voice-based intelligent agent and the second voice-based intelligent agent, the second electronic device 43 only supports the second voice-based intelligent agent, and the third electronic device 45 only supports the first voice-based intelligent agent.
[0076] To control the second electronic device 43, the user may select the second user terminal 100b that supports the second voice-based intelligent agent, and may input a voice command for controlling the second electronic device 43 into the second user terminal 100b. To control the third electronic device 45, the user may select the first user terminal 100a that supports the first voice-based intelligent agent, and may input a voice command for controlling the third electronic device 45 into the first user terminal 100a.
[0077] According to various embodiments of the present disclosure, if a user's voice command is delivered to a certain user terminal, the user terminal may generate voice data corresponding to the user's voice command. The user terminal may send the voice data to an intelligent server associated with a voice-based intelligent agent supported by the user terminal. The intelligent server may determine whether the voice data is voice data that can be processed by the voice-based intelligent agent. If the voice-based intelligent agent can process the voice data, the intelligent server may process the voice data. If the voice-based intelligent agent cannot process the voice data, the intelligent server may transmit the voice data to an intelligent server associated with another voice-based intelligent agent that can process the voice data.
[0078] For example, when a user inputs a voice command for controlling the second electronic device 43 into the first user terminal 100a, the first voice-based intelligent agent supported by the first user terminal 100a may not be able to directly control the second electronic device 43 that supports the second voice-based intelligent agent. In this case, the first voice-based intelligent agent can send the user's voice command to the second voice-based intelligent agent so that the second voice-based intelligent agent controls the second electronic device 43. A system can be provided that can use a voice-based intelligent agent to control all electronic devices in the user's living environment for the user.
[0079] Figure 5 is a diagram illustrating an example environment for controlling an electronic device supporting a second voice-based intelligent agent using a user terminal supporting a first voice-based intelligent agent according to an embodiment.
[0080] Reference Figure 5 For ease of explanation, it is assumed by way of non-limiting example that the user terminal 100 supports a first voice-based intelligent agent and the electronic device 40 supports a second voice-based intelligent agent.
[0081] According to an embodiment of the present disclosure, the user terminal 100 may receive a wake-up voice for executing a first voice-based intelligent agent through a microphone. For example, if a pre-specified voice (e.g., "Hi Bixby") is input through the microphone, the user terminal 100 may execute (call) the first voice-based intelligent agent. After inputting the wake-up voice, the user terminal 100 may receive various types of user voice commands. The user terminal 100 may process the received user's voice commands into voice data. For example, the user terminal 100 may generate voice data corresponding to the user's voice command by pre-processing various operations on the received user's voice command (such as an operation of removing an echo included in the received user's voice command, an operation of removing background noise included in the voice command, and an operation of adjusting the volume included in the voice command). The user terminal 100 may send the voice data to the first intelligent server 200 that supports the first voice-based intelligent agent.
[0082] The first intelligent server 200 according to an embodiment may include a communication circuit, a memory and / or a processor. The first intelligent server 200 may also include Figure 1 All or part of the configuration included in the first intelligent server 200. The various platforms included in the first intelligent server 200 can run under the control of the processor of the first intelligent server 200.
[0083] According to an embodiment, the first intelligent server 200 can process the received voice data to generate text data. For example, the first intelligent server 200 can process the voice data received from the user terminal 100 to generate text data through an automatic speech recognition (ASR) module (e.g., automatic speech recognition module 221). For example, the automatic speech recognition module may include a speech recognition module. The speech recognition module may include an acoustic model and a language model. For example, the acoustic model may include information related to utterance, and the language model may include unit phoneme information and information about a combination of unit phoneme information. The speech recognition module can use the information related to speech and the unit phoneme information to convert the user's voice into text data. Information about the acoustic model and the language model can be stored in an automatic speech recognition database (ASR DB).
[0084] According to an embodiment, the first intelligent server 200 can determine the user's intention and the operation corresponding to the user's intention by analyzing the text data. For example, the first intelligent server 200 can determine the user's intention and the operation corresponding to the user's intention by analyzing the text data. Figure 1The natural language understanding module 223 of the user's voice command processes the text data. The natural language understanding module can determine the domain and intent of the user's voice command and the parameters required to understand the intent by analyzing the processed text data. For example, a domain (e.g., coffee shop) can include multiple intents (e.g., coffee order and coffee order cancellation), and an intent can include multiple parameters (e.g., iced Americano and latte).
[0085] According to an embodiment, the first intelligent server 200 may determine a domain by processing text data. A domain may refer to, for example, an application or capsule that should process a voice command input by a user. A domain may refer to, for example, a device that should process a voice command input by a user.
[0086] According to an embodiment, the first intelligent server 200 can identify the device intended to be controlled by the user by analyzing the text data. The first intelligent server 200 can process the text data through, for example, a natural language understanding module to sufficiently determine the domain. The first intelligent server 200 can determine the domain by processing the text data, and can identify the device intended to be controlled by the user based on the determined domain. According to an embodiment, the first intelligent server 200 can identify the device to be controlled by classifying the domain of the received voice data. For example, referring to Figure 5 , the first intelligent server 200 can identify that the device intended to be controlled by the user is the electronic device 40. By way of non-limiting example, assume that the user's voice command is "turn on the TV (TV)." In this case, the first intelligent server 200 can process the text data corresponding to the user's voice command and can determine that the domain of the user's voice command is "TV." Therefore, the first intelligent server 200 can identify that the device intended to be controlled by the user is "TV." Even if the first intelligent server 200 does not perform a complete natural language understanding process (determining the user's intent and parameters), it can identify the device to be controlled.
[0087] According to an embodiment, the first smart server 200 may receive information about a smart agent supported by the identified device from an Internet of Things (IoT) server 500. For example, the IoT server 500 may store device information of at least one electronic device registered in a user's account, information about a smart agent supported by at least one device, and user information. For example, the IoT server 500 may store device information of the electronic device stored in the user's account, and information about a smart agent supported by the electronic device. For example, referring to Figure 5 , the first intelligent server 200 can identify that the electronic device 40 intended to be controlled by the user supports the second voice-based intelligent agent.
[0088] According to an embodiment, the first smart server 200 may request information about smart agents supported by the identified device from the IoT server 500, and may receive information about smart agents supported by the identified device. According to an embodiment, the first smart server 200 may pre-receive and store information about at least one device registered in the user account and information about smart agents supported by the at least one device from the IoT server 500, and may identify information about smart agents supported by the identified device based on the information.
[0089] According to an embodiment, the first intelligent server 200 can determine whether to send the user's voice data to the second intelligent server 600 that supports the second voice-based intelligent agent based on information about the intelligent agent supported by the identified device.
[0090] For example, the first intelligent server 200 may determine that because the intelligent agent supported by the electronic device 40 is the second voice-based intelligent agent, the first intelligent server 200 cannot directly control the electronic device 40, but the second intelligent server 600 can control the electronic device 40. In this case, the first intelligent server 200 may send the user's voice data to the second intelligent server 600 that supports the second voice-based intelligent agent. According to an embodiment, the first intelligent server 200 may send text data obtained by processing the user's voice data to the second intelligent server 600. According to an embodiment, the second intelligent server 600 may control the electronic device 40 to meet the user's intention by processing the received user's voice data.
[0091] For example, if the intelligent agent supported by the electronic device 40 is a first voice-based intelligent agent, not a second voice-based intelligent agent, the first intelligent server 200 can determine that it can directly control the electronic device 40. In this case, the first intelligent server 200 can determine that it has not sent the user's voice data to the second intelligent server 600. In this case, in addition to the domain, the first intelligent server 200 can also determine the user's intention and parameters by processing text data. For example, the first intelligent server 200 can determine the user's intention by performing syntactic analysis or semantic analysis on the text data through a natural language understanding module. Syntactic analysis can divide text data into syntactic units (e.g., words, phrases, or morphemes) and can grasp what syntactic elements the divided unit has. Semantic analysis can be performed using semantic matching, rule matching, or formula matching. The first intelligent server 200 can determine the domain, intention, and parameters required to grasp the intention for distinguishing services that match the intention corresponding to the text data. The first intelligent server 200 can identify the user's intention determined by natural language understanding and the operation suitable for the user's intention, and can identify the parameters required to perform the identified operation. The first smart server 200 can perform an operation corresponding to the user's intention based on the parameter. For example, the first smart server 200 can control the electronic device 40 to meet the user's intention.
[0092] In the following, reference is made to Figure 5 , an example scenario of controlling the electronic device 40 using the user terminal 100 will be described.
[0093] According to an exemplary embodiment of the present disclosure, the user terminal 100 supporting the first voice-based intelligent agent can receive a wake-up voice through a microphone to execute the first voice-based intelligent agent. After inputting the wake-up voice, the user terminal 100 can receive a voice command "Show me the channel guide on TV" from the user.
[0094] According to an example embodiment, the user terminal 100 may process the received voice command into voice data and may transmit the voice data to the first intelligent server 200 supporting the first voice-based intelligent agent.
[0095] According to an example embodiment, the first intelligent server 200 may process the received voice data to generate text data.
[0096] According to an exemplary embodiment, the first intelligent server 200 may identify that the device to be controlled by the user is "TV" by analyzing the text data. For example, the first intelligent server may determine the domain as "TV" by processing the text data, and may identify that the device to be controlled is "TV" based on the determined domain.
[0097] According to an exemplary embodiment, the first smart server 200 may receive information about the TV (e.g., device information) and information about the smart agent supported by the TV from the IoT server 500. The smart server 200 may receive information from the IoT server 500 notifying that the TV supports the second voice-based smart agent.
[0098] According to an exemplary embodiment, in response to the situation that the "TV" supports the second voice-based intelligent agent, the first intelligent server 200 may determine to transmit the user's voice data to the second intelligent server 600 that supports the second voice-based intelligent agent. For example, the first intelligent server 200 may transmit the user's voice data to the second intelligent server 600. According to an embodiment, the first intelligent server 200 may transmit the processed text data to the second intelligent server 600.
[0099] According to an exemplary embodiment, the second intelligent server 600 may determine the user's intention and an operation corresponding to the intention by processing the user's voice data, and may send a control command to the electronic device 40 to display a channel guide, so that the operation may be performed by the electronic device 40 corresponding to "TV". For example, the second intelligent server 600 may send a control command to the electronic device 40 to display a channel guide through a deep link.
[0100] According to an example embodiment, the electronic device 40 may display a channel guide based on the received control command.
[0101] Figure 6 is a block diagram illustrating an example operating environment between a user terminal, a first smart server, and a second smart server according to an embodiment.
[0102] Reference Figure 6 According to an embodiment of the present disclosure, the user terminal 100 may receive a user's wake-up voice through a microphone to execute the first voice-based intelligent client 101. In response to receiving the wake-up voice, the first voice-based intelligent agent may be invoked. After the first voice-based intelligent agent is invoked, the user terminal 100 may receive a voice command from the user. The user terminal 100 may process the received voice command into voice data. The user terminal 100 may transmit the voice data to the first intelligent server 200 that supports the first voice-based intelligent agent.
[0103] The database 235 according to an embodiment may match and store information about at least one device registered in a user account and information about an intelligent agent supported by the at least one device.
[0104] According to an embodiment, the first intelligent server 200 may include an automatic speech recognition module (e.g., including a processing circuit and / or an executable program element) 221, a multi-intelligent agent determination module (e.g., including a processing circuit and / or an executable program element) 290, a multi-intelligent agent interface module (e.g., including a processing circuit and / or an executable program element) 291, a first natural language understanding module (e.g., including a processing circuit and / or an executable program element) 223 and a database 235. The automatic speech recognition module 221, the multi-intelligent agent determination module 290, the multi-intelligent agent interface module 291 and the first natural language understanding module 223 may be software modules stored in a memory (not shown) of the first intelligent server 200 and may operate under the control of a processor. According to various embodiments, the memory may store various instructions for the operation of the processor of the first intelligent server 200. The first intelligent server 200 may also include Figure 1 As mentioned above, the above reference will be omitted. Figure 5 According to an embodiment, the automatic speech recognition module 221 of the first intelligent server 200 can process the user's speech data to generate text data. The automatic speech recognition module 221 can send the processed text data to the multi-intelligent agent determination module 290.
[0105] According to an embodiment, the multi-intelligent agent determination module 290 may be a module that determines whether the user's voice data can be processed by the first voice-based intelligent agent by processing text data. The multi-intelligent agent determination module 290 may, for example, identify the device intended to be controlled by the user by processing the text data. The multi-intelligent agent determination module 290 may identify information about the identified device and information about the intelligent agents supported by the identified device from the database 235. For example, according to an embodiment, the multi-intelligent agent determination module 290 may be a module that only performs operations to determine the domain of the natural language understanding processing being processed by the natural language understanding module. A domain may refer to, for example, an application or capsule that should process the user's input voice command, or a device that should process the user's input voice command. Based on, for example, information about the intelligent agents supported by the identified device, the multi-intelligent agent determination module 290 may determine whether to transmit the user's voice data to the second intelligent server 600 that supports the second voice-based intelligent agent. According to an embodiment, the multi-intelligent agent determination module 290 may be, for example, a module that identifies the device intended to be controlled by the user.
[0106] For example, if the identified device supports the first voice-based intelligent agent, the multi-intelligent agent determination module 290 may determine not to send the user's voice data to the second intelligent server 600, and may determine that the first intelligent server 200 processes the user's voice data. The multi-intelligent agent determination module 290 may send the text data to the first natural language understanding module 223. The first natural language understanding module 223 may determine an operation corresponding to the user's intention based on the determined parameters. The first intelligent server 200 may send a control command to the identified device to cause the identified device to perform the determined operation.
[0107] For example, if the identified device supports the second voice-based intelligent agent, the multi-intelligent agent determination module 290 may determine to send the user's voice data to the second intelligent server 600. In this case, the processor of the first intelligent server 200 may send the user's voice data to the second intelligent server 600. According to an embodiment, the processor of the first intelligent server 200 may send the processed text data to the second intelligent server 600.
[0108] The second intelligent server 600 according to the embodiment can process the user's voice data received from the first intelligent server 200 to generate text data through the automatic speech recognition module 610. If the text data is received from the first intelligent server 200, the above operation can be omitted. The second intelligent server 600 according to the embodiment can determine the user's intention and the parameters corresponding to the intention by processing the text data by the second natural language understanding module 620, and can determine the operation corresponding to the user's intention based on the parameters. According to the embodiment, the second intelligent server 600 can send a control command for causing the identified device to perform the determined operation to the identified device.
[0109] According to an embodiment, the first intelligent server 200 can transmit and receive various messages with the second intelligent server 600 through the multi-intelligent agent interface module 291. According to an embodiment, the first intelligent server 200 can receive a message including an additional question from the second intelligent server 600. In this case, the first intelligent server 200 can transmit the message including the additional question to the user terminal, and can receive a response message to the additional question from the user terminal 100.
[0110] Figure 7 is a block diagram illustrating an example operating environment between a first user terminal and a second user terminal according to an embodiment.
[0111] Reference Figure 7 , assuming that the first user terminal 70 supports a first voice-based intelligent agent, and the second user terminal 80 supports a second voice-based intelligent agent. Figure 5 and Figure 6 A description of the content that the content overlaps.
[0112] According to an embodiment, the first user terminal 70 may include a first voice-based intelligent agent client 710, a natural language platform (e.g., including a processing circuit and / or an executable program element) 720, a multi-intelligent agent interface module (e.g., including a processing circuit and / or an executable program element) 730, a database 740 (e.g., Figure 1 Memory 150), a microphone (not shown) (e.g., Figure 1 a microphone 120), a display (not shown) (e.g., Figure 1 display 140), a speaker (not shown) (e.g., Figure 1 Speaker 130), communication circuit (not shown) (e.g., Figure 1 Communication interface 110), memory (not shown) (eg, Figure 1 Memory 150) or a processor (not shown) (e.g., Figure 1 The natural language platform 720 may include, for example, an automatic speech recognition module (ASR) (e.g., including processing circuitry and / or executable program elements) 721, a natural language understanding module (NLU) (e.g., including processing circuitry and / or executable program elements) 723, and a text-to-speech module (TTS) (e.g., including processing circuitry and / or executable program elements) 725.
[0113] According to the embodiment, the second user terminal 80 may include a second voice-based intelligent agent client 810, a natural language platform (e.g., including a processing circuit and / or an executable program element) 820, a multi-intelligent agent interface module (e.g., including a processing circuit and / or an executable program element) 830, a database 840, a communication circuit (not shown) or a processor (not shown). The natural language platform 820 may include, for example, an automatic speech recognition module (ASR) (e.g., including a processing circuit and / or an executable program element) 821, a natural language understanding module (NLU) (e.g., including a processing circuit and / or an executable program element) 823, and a text-to-speech module (TTS) (e.g., including a processing circuit and / or an executable program element) 825. The natural language understanding module 723 of the first user terminal 70 and the natural language understanding module 823 of the second user terminal 80 may be different natural language understanding modules from each other.
[0114] According to an embodiment, the database 740 of the first user terminal 70 and the database 840 of the second user terminal 80 may store device information 750 registered in the first voice-based intelligent agent and device information 850 registered in the second voice-based intelligent agent. For example, the first user terminal 70 and the second user terminal 80 may integrally store information about at least one device registered in the first voice-based intelligent agent and information about at least one device registered in the second voice-based intelligent agent.
[0115] According to an embodiment, the first user terminal 70 may receive the user's voice data through a microphone.
[0116] According to an embodiment, the first user terminal 70 may process the received voice data to generate text data.
[0117] According to an embodiment, the first user terminal 70 may identify a device intended to be controlled by the user by analyzing text data.
[0118] According to an embodiment, the first user terminal 70 may identify information about at least one device registered in a pre-stored user account and information about an intelligent agent supported by the at least one device. According to an embodiment, the first user terminal 70 may determine whether to transmit the user's voice data to the second user terminal 80 supporting a second voice-based intelligent agent based on the information about the intelligent agent supported by the identified device.
[0119] For example, if the identified device supports the second voice-based intelligent agent, the first user terminal 70 can transmit the user's voice data to the second user terminal 80 through the communication circuit. For example, if the identified device supports the first voice-based intelligent agent, the first user terminal 70 can process the text data through the natural language platform 720.
[0120] According to an embodiment, the device information 750 registered in the first voice-based intelligent agent and the device information 850 registered in the second voice-based intelligent agent may include, for example, information about the name of the registered device (e.g., the registered name) and the intelligent agents supported by the registered device. In this case, even if the name of the same device registered in the first voice-based intelligent agent is different from the name registered in the second voice-based intelligent agent, the device name can be changed based on the stored information. For example, in a state where an electronic device is registered as "master bedroom TV (TV)" in the first voice-based intelligent agent and as "large TV" in the second voice-based intelligent agent, even if a voice command "turn on the large TV" is received in the first user terminal 70, the first user terminal 70 can also recognize that the device intended to be controlled by the user is "master bedroom TV". According to an embodiment, even if any of the registered device names of the same device is input from the user, the first user terminal 70 and the second user terminal 80 can accurately identify the device intended to be controlled by the user.
[0121] According to an embodiment, the first user terminal 70 and the second user terminal 80 may transmit and receive various messages to and from each other through the multi-intelligent agent interface modules 730 and 830. According to an embodiment, the first user terminal 70 may receive a message including an additional question from the second user terminal 80.
[0122] Figure 8 is a diagram showing a first intelligent server (eg, Figure 5 A flowchart of an example operation of the first intelligent server 200). Figure 8 , the first intelligent server is described as an electronic device.
[0123] Referring to flowchart 800, the processor of the electronic device according to the embodiment may, at operation 810, receive data from a user terminal (eg, Figure 5 The electronic device may be a server (eg, a user terminal 100) that supports a first voice-based intelligent agent. Figure 5 first intelligent server 200).
[0124] In operation 820, the processor of the electronic device according to the embodiment may process the voice data to generate text data.
[0125] In operation 830, the processor of the electronic device according to an embodiment may identify a device intended to be controlled by the user by analyzing text data.
[0126] At operation 840, the processor of the electronic device according to the embodiment may receive information about the intelligent agent supported by the identified device. For example, the processor may receive information from a first external server (e.g., Figure 5The IoT server 500 receives information about the smart agent supported by the identified device. For example, the first external server may store device information of at least one electronic device registered in the user's account, information about the smart agent supported by the at least one device, and user information.
[0127] At operation 850, the processor of the electronic device according to the embodiment may determine whether to transmit the user's voice data to an external server (eg, Figure 5 The external server may be a server that supports the second voice-based intelligent agent. For example, if the identified device supports the first voice-based intelligent agent, the processor may determine not to send the user's voice data to the external server. For example, if the identified device supports the second voice-based intelligent agent, the processor may determine to send the user's voice data to the external server.
[0128] Figure 9 is a diagram showing a first intelligent server (eg, Figure 5 A flowchart of an example operation of the first intelligent server 200). Figure 9 , the first intelligent server is described as an electronic device.
[0129] Referring to the operation flow chart 900, the processor of the electronic device according to the embodiment may, at operation 901, receive a signal from a user terminal (eg, Figure 5 The user terminal 100 receives the user's voice data. The electronic device can be, for example, a server that supports the first voice-based intelligent agent.
[0130] In operation 903 , the processor of the electronic device according to an embodiment may process the received voice data to generate text data.
[0131] In operation 905 , the processor of the electronic device according to an embodiment may determine the domain of the voice data by analyzing the text data.
[0132] In operation 907 , the processor of the electronic device according to the embodiment may identify a device intended to be controlled by the user based on the determined domain.
[0133] At operation 909, the processor of the electronic device according to the embodiment may receive a message from a first external server (eg, Figure 5 The IoT server 500 receives information about the smart agent supported by the identified device. For example, the first external server may store device information of at least one electronic device registered in the user's account, information about the smart agent supported by the at least one device, and user information.
[0134] In operation 911, the processor of the electronic device according to an embodiment may identify whether the identified device supports the first voice-based intelligent agent based on the received information.
[0135] If it is identified that the identified device does not support the first voice-based intelligent agent, the operation branches to operation 919 (911-No), and the processor of the electronic device may transmit the user's voice data to a second external server (e.g., a server that supports an intelligent agent capable of controlling the identified device) Figure 5 For example, if the identified device supports a second voice-based intelligent agent, the processor of the electronic device may send the user's voice data to a second external server that supports the second voice-based intelligent agent.
[0136] If it is recognized that the identified device supports the first voice-based intelligent agent, the operation branches to operation 913 ( 911 —Yes), and the processor of the electronic device may determine the user's intention and parameters corresponding to the intention by analyzing text data.
[0137] At operation 915 , the processor of the electronic device according to the embodiment may determine an operation corresponding to the user's intention based on the determined parameters.
[0138] In operation 917 , the processor of the electronic device according to the embodiment may transmit a control command to the identified device based on the determined operation.
[0139] Figure 10 1 is a signal flow diagram illustrating an example operation among the user terminal 100, the first intelligent server 200, the IoT server 500, and the second intelligent server 600 according to an embodiment. According to an embodiment, the user terminal 100 and the first intelligent server 200 can support a first voice-based intelligent agent, and the second intelligent server 600 can support a second voice-based intelligent agent.
[0140] Referring to signal flow chart 1000, the user terminal 100 according to an embodiment may receive a voice command from a user at operation 1001. For example, the user terminal 100 may receive the user's speech through a microphone.
[0141] In operation 1003 , the user terminal 100 according to an embodiment may process the received voice command to generate voice data.
[0142] In operation 1005 , the user terminal 100 according to an embodiment may transmit voice data to the first intelligent server 200 .
[0143] In operation 1007 , the first intelligent server 200 according to an embodiment may process the received voice data to generate text data.
[0144] In operation 1009 , the first smart server 200 according to an embodiment may identify a device intended to be controlled by a user by analyzing text data.
[0145] In operation 1011, the first smart server 200 according to an embodiment may request information about the identified device from the IoT server 500. The information about the identified device may include, for example, information about a smart agent supported by the identified device.
[0146] In operation 1013 , the first smart server 200 according to an embodiment may receive information about the identified device transmitted from the IoT server 500 .
[0147] In operation 1015, the first intelligent server 200 according to an embodiment may determine whether to transmit the user's voice data to a second external server that supports a second voice-based intelligent agent based on the information about the identified device. For example, the first intelligent server 200 may determine whether to transmit the user's voice data to a second external server that supports a second voice-based intelligent agent based on the information about the intelligent agent supported by the identified device.
[0148] For example, if the identified device supports the first voice-based intelligent agent, the first intelligent server 200 can determine not to send the user's voice data to the second intelligent server 600, and can determine that the first intelligent server 200 processes the user's voice data.
[0149] For example, if the identified device supports the second voice-based intelligent agent, the first intelligent server 200 can determine to send the user's voice data to the second intelligent server 600.
[0150] In operation 1017, the first intelligent server 200 according to an embodiment may transmit the user's voice data to the second intelligent server 600 in response to the identified device supporting the second voice-based intelligent agent.
[0151] Figure 11 1 is a signal flow diagram illustrating an example operation among the user terminal 100, the first intelligent server 200, and the second intelligent server 600 according to an embodiment. According to an embodiment, the user terminal 100 and the first intelligent server 200 can support a first voice-based intelligent agent, and the second intelligent server 600 can support a second voice-based intelligent agent.
[0152] Referring to the signal flow chart 1100, in operation 1101, the first intelligent server 200 according to the embodiment may store data in a memory (e.g., Figure 6 In operation 1103, the second intelligent server 600 according to the embodiment may also store information about at least one device registered in the user account and information about the intelligent agent supported by the at least one device in a memory (e.g., Figure 6 The first intelligent server 200 stores information about at least one device registered in the user account and information about the intelligent agent supported by the at least one device in a database 630. For example, the first intelligent server 200 may store not only the device information registered in the first voice-based intelligent agent, but also the device information registered in the second voice-based intelligent agent. The information about the at least one device registered in the user account may include, for example, the name of the at least one device registered in the first voice-based intelligent agent and the second voice-based intelligent agent, respectively.
[0153] In operation 1105 , the user terminal 100 according to an embodiment may receive a voice command from a user.
[0154] In operation 1107 , the user terminal 100 according to an embodiment may process the received voice command to generate voice data.
[0155] In operation 1109 , the user terminal 100 according to an embodiment may transmit voice data to the first intelligent server 200 .
[0156] In operation 1111 , the user terminal 100 according to an embodiment may process received voice data to generate text data.
[0157] In operation 1113, the first intelligent server 200 according to an embodiment may identify a device intended to be controlled by the user by analyzing the text data. For example, the first intelligent server 200 may determine the domain of the voice data by analyzing the text data, and may identify a device intended to be controlled by the user based on the determined domain.
[0158] In operation 1115, the first intelligent server 200 according to an embodiment may determine whether to transmit the user's voice data to the second intelligent server 600. For example, the first intelligent server 200 may identify information about an intelligent agent supported by the identified device based on information about at least one device registered in the user account and information about intelligent agents supported by the at least one device stored in the memory. The first intelligent server 200 may determine whether to transmit the user's voice data to the second intelligent server 600 supporting the second voice-based intelligent agent based on the identified information.
[0159] In operation 1117, the first intelligent server 200 according to an embodiment may transmit the user's voice data to the second intelligent server 600 in response to the identified device supporting the second voice-based intelligent agent.
[0160] In operation 1119 , the first smart server 200 according to an embodiment may transmit a transmission completion message including information that the voice data has been completely transmitted to the second smart server 600 to the user terminal 100 .
[0161] In operation 1121, in response to receiving the transmission completion message, the user terminal 100 according to an embodiment may output a message to the second voice-based intelligent agent notifying the second voice-based intelligent agent that the transmission of the user's voice command is complete. For example, the user terminal 100 may output the transmission completion message as sound through a speaker. For example, the user terminal 100 may display the transmission completion message on a display.
[0162] In operation 1123, the second intelligent server 600 according to an embodiment may control the identified device by processing the user's voice data in response to receiving the user's voice data from the first intelligent server 200. For example, the second intelligent server 600 may determine the user's intention and an operation corresponding to the user's intention by processing the user's voice data, and may send a control command to the identified device based on the determined operation.
[0163] In operation 1125, the second intelligent server 600 according to an embodiment may send a voice command completion message including information that the second voice-based intelligent agent has completed the voice command to the first intelligent server 200 in response to control of the identified device.
[0164] In operation 1127 , the first intelligent server 200 according to an embodiment may transmit a voice command completion message to the user terminal 100 .
[0165] In operation 1129, the user terminal 100 according to an embodiment may output a message notifying the completion of the voice command through the second voice-based intelligent agent in response to receiving the voice command completion message. For example, the user terminal 100 may output the voice command completion message as sound through a speaker. For example, the user terminal 100 may display the voice command completion message on a display.
[0166] Figure 12 is a signal flow diagram illustrating an example operation between a first user terminal 70 and a second user terminal 80 according to an embodiment. According to an embodiment, the first user terminal 70 may store information about at least one device registered in a user account and information about an intelligent agent supported by the at least one device in a memory (e.g., Figure 7in the database 740).
[0167] Reference Figure 12 , in operation 1201, the first user terminal 70 according to an embodiment may receive a user's voice command.
[0168] In operation 1203 , the first user terminal 70 according to an embodiment may process the received voice command to generate voice data.
[0169] In operation 1205 , the first user terminal 70 according to an embodiment may process voice data to generate text data.
[0170] In operation 1207 , the first user terminal 70 according to an embodiment may identify a device intended to be controlled by the user by analyzing text data.
[0171] In operation 1209 , the first user terminal 70 according to an embodiment may identify information about an intelligent agent supported by the identified device based on pre-stored information.
[0172] In operation 1211 , the first user terminal 70 according to an embodiment may transmit the user's voice data to the second user terminal 80 in response to a recognized case that the device supports the second voice-based intelligent agent.
[0173] At operation 1213, the first user terminal 70 according to an embodiment may output a transmission completion message including information that the user's voice data has been transmitted to the second user terminal 80 supporting the second voice-based intelligent agent. For example, the first user terminal 70 may output the transmission completion message as sound through a speaker. For example, the first user terminal 70 may display the transmission completion message on a display.
[0174] At operation 1215, the second user terminal 80 according to an embodiment may control the identified device by processing the user's voice data in response to receiving the user's voice data from the first user terminal 70. For example, the second user terminal 80 may determine the user's intention and an operation corresponding to the user's intention by processing the user's voice data, and may send a control command to the identified device based on the determined operation.
[0175] In operation 1217, the second user terminal 80 according to an embodiment may transmit a voice command completion message including information that the second voice-based intelligent agent has completed the voice command to the first user terminal 70 in response to control of the identified device.
[0176] In operation 1219, in response to receiving the voice command completion message, the first user terminal 70 according to an embodiment may output a message notifying the second voice-based intelligent agent of the completion of the voice command. For example, the first user terminal 70 may output the voice command completion message as sound through a speaker. For example, the first user terminal 70 may display the voice command completion message on a display.
[0177] Figure 13 is a diagram illustrating an example environment for controlling an electronic device 45 supporting a second voice-based intelligent agent 1303 through a user terminal 100a supporting a first voice-based intelligent agent 1301 according to an embodiment.
[0178] Reference Figure 13 , assuming that the first user terminal 100a is a device that supports the first voice-based intelligent agent 1301, and the electronic device 45 is an electronic device that supports the second voice-based intelligent agent 1303.
[0179] According to an embodiment, a user may input a voice command "Hey Bixby, turn on the living room light" into the first user terminal 100a. "Hey Bixby" may be, for example, a wake-up call for invoking the first voice-based intelligent agent 1301 client. The first voice-based intelligent agent 1301 may process the input user voice command "turn on the living room light."
[0180] According to an embodiment, the first voice-based intelligent agent 1301 may recognize that the device to be controlled by the user is a "living room lamp". The first voice-based intelligent agent 1301 may recognize that the intelligent agent supported by the "living room lamp" is the second voice-based intelligent agent 1303. For example, the first voice-based intelligent agent 1301 may receive a message from an external server (e.g., Figure 5 The IoT server 500 receives information about the smart agent supported by the "living room lamp". For example, the first voice-based smart agent 1301 can pre-store information about the smart agent supported by the "living room lamp" in the database.
[0181] According to an embodiment, the first voice-based intelligent agent 1301 can send the user's voice command to the second voice-based intelligent agent 1303.
[0182] For example, the first user terminal 100a may transmit the user's voice data obtained by processing the user's voice command to the first intelligent server 200 supporting the first voice-based intelligent agent 1301. The first intelligent server 200 may identify the device intended to be controlled by the user by processing the user's voice data. The first intelligent server 200 may identify information about the intelligent agent supported by the device intended to be controlled by the user and may transmit the user's voice data to the second intelligent server 600 supporting the second voice-based intelligent agent 1303.
[0183] For example, the first user terminal 100a can identify information about the intelligent agent supported by the device intended to be controlled by the user by processing the user's voice data, and in response to the device intended to be controlled by the user supporting the second voice-based intelligent agent 1303, the user's voice data can be sent directly to the second user terminal 100b supporting the second voice-based intelligent agent 1303.
[0184] According to an embodiment, the first user terminal 100a may output a message indicating that the first user terminal 100a has sent the user's voice command to the second voice-based intelligent agent 1303. For example, the first user terminal 100a may output the message "I have commanded Google to process the living room light" as a voice through a speaker. For example, Google may be the name of the second voice-based intelligent agent 1303.
[0185] According to an embodiment, second voice-based intelligent agent 1303 may process a received user voice command. By processing the user's voice command, second voice-based intelligent agent 1303 may recognize that the device to be controlled by the user is a living room lamp, determine that the user intends to increase the brightness of the living room lamp, and determine parameters for setting the brightness of the living room lamp to maximum brightness. Second voice-based intelligent agent 1303 may transmit a control command for setting the living room lamp to maximum brightness to electronic device 45 corresponding to the living room lamp.
[0186] For example, the second intelligent server 600 can process the received user voice data and send a control command to the electronic device 45 intended to be controlled by the user to perform an operation corresponding to the user's intention included in the user's voice data. For example, the second intelligent server 600 can send a control command to set the electronic device 45 intended to be controlled by the user to maximum brightness.
[0187] For example, the second user terminal 100b can directly process the user's voice data and send a control command to the electronic device 45 intended to be controlled by the user to perform an operation corresponding to the user's intention included in the user's voice data. For example, the second user terminal 100b can send a control command to the electronic device 45 to set the electronic device 45 intended to be controlled by the user to maximum brightness.
[0188] According to an embodiment, the second user terminal 100b may output a message notifying the user of the completion of the received voice command. For example, the second user terminal 100b may output the message "I have set the living room light to maximum brightness" as a sound. According to an embodiment, the message notifying the user of the completion of the voice command may be sent to the first user terminal 100a via the second voice-based intelligent agent 1303, so that the first user terminal 100a may output the message.
[0189] For example, the second intelligent server 600 may control the second user terminal 100b to output a message indicating that the second user terminal 100b supporting the second voice-based intelligent agent 1303 has received and processed the user's voice command in response to the user's voice command "brighten the living room light." For example, the second user terminal 100b may output a message "I have set the living room light to maximum brightness" as a sound through a speaker.
[0190] For example, if the second user terminal 100b directly processes the user's voice data, the second user terminal 100b may output a message indicating that the user's voice command has been transmitted and processed in response to processing of the user's voice command "brighten the living room light." For example, the second user terminal 100b may output a message "I have set the living room light to maximum brightness" as sound through a speaker.
[0191] Figure 14 is a diagram illustrating an example environment for controlling an electronic device 47 supporting a second voice-based intelligent agent 1305 through a user terminal 100a supporting a first voice-based intelligent agent 1301 according to an embodiment.
[0192] Reference Figure 14 , assuming that the user terminal 100a is a device that supports the first voice-based intelligent agent 1301, and the electronic device 47 is an electronic device that supports the second voice-based intelligent agent 1305.
[0193] According to an embodiment, a user may input a voice command into user terminal 100a, "Hey Bixby, turn on the car's air conditioning." "Hey Bixby" may be, for example, a wake-up call for invoking the first voice-based intelligent agent 1301 client. The first voice-based intelligent agent 1301 may process the input user voice command, "Turn on the car's air conditioning."
[0194] According to an embodiment, the first voice-based intelligent agent 1301 may recognize that the device intended to be controlled by the user is a "car". The first voice-based intelligent agent 1301 may recognize that the intelligent agent supported by the "car" is the second voice-based intelligent agent 1305. For example, the first voice-based intelligent agent 1301 may receive a message from an external server (e.g., Figure 5 The IoT server 500 receives information about the intelligent agents supported by the “car.” For example, the first voice-based intelligent agent 1301 may pre-store information about the intelligent agents supported by the “car” in a database.
[0195] According to an embodiment, the first voice-based intelligent agent 1301 may transmit the user's voice command to the electronic device 47 supporting the second voice-based intelligent agent 1305. According to an embodiment, the user terminal 100a may output an indication that the user terminal 100a has transmitted the user's voice command to the electronic device 47 supporting the second voice-based intelligent agent 1305. For example, the user terminal 100a may output a message "I have transmitted the command to the car" as a sound.
[0196] According to an embodiment, the second voice-based intelligent agent 1305 may process the received user's voice command. By processing the user's voice command, the second voice-based intelligent agent 1305 may recognize that the device to be controlled by the user is a car, and may determine that the user's intention is to turn on the car's air conditioning. The second voice-based intelligent agent 1305 may control the electronic device 47 corresponding to the car to turn on the car's air conditioning.
[0197] According to an embodiment, the second voice-based intelligent agent 1305 may send a message to the user terminal 100a notifying the user of the completion of the voice command. According to an embodiment, the user terminal 100a may output a message notifying the user of the completion of the received voice command. For example, the user terminal 100a may output the message "The car has completed the command" as a sound.
[0198] Figure 15 is a diagram illustrating an example environment for controlling an electronic device 40 supporting a second voice-based intelligent agent 1303 through a user terminal 100c supporting a first voice-based intelligent agent 1301 according to an embodiment.
[0199] Reference Figure 15 , assuming that the user terminal 100c is a device that supports the first voice-based intelligent agent 1301, and the electronic device 40 is an electronic device that supports the second voice-based intelligent agent 1303.
[0200] According to an embodiment, a user may input a voice command "Hi Bixby, turn off the TV" to the user terminal 100c. "Hi Bixby" may be, for example, a wake-up voice for invoking the first voice-based intelligent agent 1301 client. The first voice-based intelligent agent 1301 may process the input user's voice command "turn off the TV."
[0201] According to an embodiment, the first voice-based intelligent agent 1301 may recognize that the device intended to be controlled by the user is “TV.” The first voice-based intelligent agent 1301 may recognize that the intelligent agent supported by “TV” is the second voice-based intelligent agent 1303 .
[0202] According to an embodiment, the first voice-based intelligent agent 1301 may transmit the user's voice command to the second voice-based intelligent agent 1303. According to an embodiment, the user terminal 100c may output a message to the user interface (UI) 1510 indicating that the user terminal 100c has transmitted the user's voice command to the second voice-based intelligent agent 1303. For example, the user terminal 100c may display the intelligent agent to which the voice command has been transmitted as a logo, and may output the transmitted voice command through the UI 1510.
[0203] According to an embodiment, the second voice-based intelligent agent 1303 may process the received user's voice command. By processing the user's voice command, the second voice-based intelligent agent 1303 may recognize that the device intended to be controlled by the user is a TV and may determine that the user's intention is to turn off the TV. The second voice-based intelligent agent 1303 may control the TV to turn off the power of the TV.
[0204] According to an embodiment, the second voice-based intelligent agent 1303 may send a message to the user terminal 100c notifying the user that the voice command has been completed. According to an embodiment, the user terminal 100c may output a message through the UI notifying the second voice-based intelligent agent 1303 that the user's voice command has been completed. According to an embodiment, the user terminal 100c may output the second voice-based intelligent agent and the content of the sent voice command through the display. According to an embodiment, the user terminal 100c may output a message as a sound notifying the user that the voice command has been completed by the second voice-based intelligent agent 1303. For example, the user terminal 100c may output the message "Turn off the power" as a sound.
[0205] Figure 16 1 is a block diagram illustrating an example electronic device 1601 in a network environment 1600 according to various embodiments. Figure 16 , the electronic device 1601 in the network environment 1600 may communicate with the electronic device 1602 via a first network 1698 (e.g., a short-range wireless communication network), or may communicate with the electronic device 1604 or the server 1608 via a second network 1699 (e.g., a long-range wireless communication network). According to an embodiment, the electronic device 1601 may communicate with the electronic device 1604 via the server 1608. According to an embodiment, the electronic device 1601 may include a processor 1620, a memory 1630, an input device 1650, a sound output device 1655, a display device 1660, an audio module 1670, a sensor module 1676, an interface 1677, a haptic module 1679, a camera module 1680, a power management module 1688, a battery 1689, a communication module 1690, a subscriber identification module (SIM) 1696, or an antenna module 1697.
[0206] In some embodiments, at least one of these components (e.g., display device 1660 or camera module 1680) may be omitted from electronic device 1601, or one or more other components may be added to electronic device 1601. In some embodiments, some components may be implemented as a single integrated circuit. For example, sensor module 1676 (e.g., a fingerprint sensor, an iris sensor, or an illumination sensor) may be implemented as embedded in display device 1660 (e.g., a display).
[0207] The processor 1620 may execute, for example, software (e.g., program 1640) to control at least one other component (e.g., hardware or software component) of the electronic device 1601 coupled to the processor 1620, and may perform various data processing or calculations. According to an example embodiment, as at least part of the data processing or calculation, the processor 1620 may load commands or data received from another component (e.g., sensor module 1676 or communication module 1690) into the volatile memory 1632, process the commands or data stored in the volatile memory 1632, and store the resulting data in the non-volatile memory 1634. According to an embodiment, the processor 1620 may include a main processor 1621 (e.g., a central processing unit (CPU) or an application processor (AP)), an auxiliary processor 1623 (e.g., a graphics processing unit (GPU), an image signal processor (ISP), a sensor hub processor, or a communication processor (CP)) that may operate independently of the main processor 1621 or in conjunction with the main processor. Additionally or alternatively, the secondary processor 1623 may be adapted to consume less power than the primary processor 1621 or be dedicated to a specific function. The secondary processor 1623 may be implemented separate from the primary processor 1621 or as part of the primary processor 1621.
[0208] The auxiliary processor 1623 may control at least some of the functions or states related to at least one of the components of the electronic device 1601 (e.g., the display device 1660, the sensor module 1676, or the communication module 1690) instead of the main processor 1621 when the main processor 1621 is in an inactive (e.g., sleep) state, or control at least some of the functions or states related to at least one of the components of the electronic device 1601 (e.g., the display device 1660, the sensor module 1676, or the communication module 1690) together with the main processor 1621 when the main processor 1621 is in an active state (e.g., executing an application). Depending on the embodiment, the auxiliary processor 1623 (e.g., an image signal processor or a communication processor) may be implemented as part of another component (e.g., the camera module 1680 or the communication module 1690) that is functionally related to the auxiliary processor 1623.
[0209] The memory 1630 may store various data used by at least one component of the electronic device 1601 (e.g., the processor 1620 or the sensor module 1676). The various data may include, for example, software (e.g., the program 1640) and input data or output data for commands related thereto. The memory 1630 may include a volatile memory 1632 or a non-volatile memory 1634.
[0210] The program 1640 may be stored as software in the memory 1630 and may include, for example, an operating system (OS) 1642 , middleware 1644 , or applications 1646 .
[0211] The input device 1650 may receive commands or data from outside the electronic device 1601 (e.g., a user) to be used by other components of the electronic device 1601 (e.g., the processor 1620). The input device 1650 may include, for example, a microphone, a mouse, a keyboard, or a digital pen (e.g., a stylus).
[0212] The sound output device 1655 can output sound signals to the outside of the electronic device 1601. The sound output device 1655 can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as playing multimedia or playing records, while the receiver can be used for incoming calls. Depending on the embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0213] The display device 1660 can visually provide information to the outside of the electronic device 1601 (e.g., a user). The display device 1660 may include, for example, a display, a hologram device, or a projector, and a control circuit that controls a corresponding one of the display, the hologram device, and the projector. Depending on the embodiment, the display device 1660 may include a touch circuit adapted to detect a touch, or a sensor circuit adapted to measure the strength of the force caused by the touch (e.g., a pressure sensor).
[0214] The audio module 1670 can convert sound into an electrical signal, and vice versa. According to an embodiment, the audio module 1670 can obtain sound via the input device 1650, or output sound via the sound output device 1655 or an earphone of an external electronic device (e.g., electronic device 1602) directly (e.g., wired) or wirelessly coupled to the electronic device 1601.
[0215] The sensor module 1676 can detect the operating state (e.g., power or temperature) of the electronic device 1601 or the environmental state (e.g., the state of the user) outside the electronic device 1601, and then generate an electrical signal or data value corresponding to the detected state. Depending on the embodiment, the sensor module 1676 may include, for example, a posture sensor, a gyroscope sensor, an atmospheric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an infrared (IR) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illumination sensor.
[0216] The interface 1677 may support one or more designated protocols for the electronic device 1601, which is directly (e.g., wired) or wirelessly coupled to an external electronic device (e.g., the electronic device 1602). Depending on the embodiment, the interface 1677 may include, for example, a High-Definition Multimedia Interface (HDMI), a Universal Serial Bus (USB) interface, a Secure Digital (SD) card interface, or an audio interface.
[0217] The connection end 1678 may include a connector through which the electronic device 1601 can be physically connected to an external electronic device (e.g., the electronic device 1602). Depending on the embodiment, the connection end 1678 may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0218] The haptic module 1679 may convert electrical signals into mechanical stimulation (eg, vibration or motion) or electrical stimulation, which may be recognized by the user via their tactile or kinesthetic sense. Depending on the embodiment, the haptic module 1679 may include, for example, electrodes, piezoelectric elements, or electrical stimulators.
[0219] The camera module 1680 may capture still images or moving images. Depending on the embodiment, the camera module 1680 may include one or more lenses, image sensors, image signal processors, or flashes.
[0220] The power management module 1688 may manage power supplied to the electronic device 1601. According to an example embodiment, the power management module 1688 may be implemented as, for example, at least a portion of a power management integrated circuit (PMIC).
[0221] The battery 1689 may supply power to at least one component of the electronic device 1601. Depending on an embodiment, the battery 1689 may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0222] The communication module 1690 can support establishing a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device 1601 and an external electronic device (e.g., electronic device 1602, electronic device 1604, or server 1608), and communicate through the established communication channel. The communication module 1690 may include one or more communication processors that can operate independently of the processor 1620 (e.g., an application processor (AP)) and support direct (e.g., wired) communication or wireless communication. According to an embodiment, the communication module 1690 may include a wireless communication module 1692 (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module 1694 (e.g., a local area network (LAN) communication module or a power line communication (PLC) module). A corresponding one of these communication modules can communicate via a first network 1698 (e.g., a short-range communication network such as Bluetooth TM , Wireless Fidelity (Wi-Fi) Direct, or Infrared Data Association (IrDA)) or a second network 1699 (for example, a long-distance communication network such as a cellular network, the Internet, or a computer network (for example, a LAN or a wide area network (WAN)) to communicate with an external electronic device. These various types of communication modules can be implemented as a single component (for example, a single chip), or these various types of communication modules can be implemented as multiple components separated from each other (for example, multiple chips). The wireless communication module 1692 can identify and authenticate the electronic device 1601 in a communication network (such as the first network 1698 or the second network 1699) using user information (for example, an International Mobile Subscriber Identity (IMSI)) stored in the user identification module 1696.
[0223] Antenna module 1697 can transmit or receive signals or power to or from the outside of electronic device 1601 (e.g., an external electronic device). Depending on the embodiment, antenna module 1697 may include an antenna comprising a radiating element formed of a conductive material or conductive pattern formed in or on a substrate (e.g., a PCB). Depending on the embodiment, antenna module 1697 may include multiple antennas. In this case, at least one antenna suitable for the communication scheme used in a communication network (such as first network 1698 or second network 1699) may be selected from the multiple antennas by, for example, communication module 1690 (e.g., wireless communication module 1692). Signals or power can then be transmitted or received between communication module 1690 and the external electronic device via the selected at least one antenna. Depending on the embodiment, additional components (e.g., a radio frequency integrated circuit (RFIC)) in addition to the radiating element may also be formed as part of antenna module 1697.
[0224] At least some of the above components can be connected to each other via an inter-peripheral communication scheme (e.g., a bus, general-purpose input output (GPIO), serial peripheral interface (SPI), or mobile industry processor interface (MIPI)) and communicatively transmit signals (e.g., commands or data) therebetween.
[0225] According to an embodiment, commands or data may be transmitted or received between electronic device 1601 and external electronic device 1604 via server 1608 coupled to second network 1699. Each of electronic device 1602 and electronic device 1604 may be of the same type as electronic device 1601, or of a different type. According to an embodiment, all or some operations to be executed on electronic device 1601 may be executed on one or more of external electronic device 1602, external electronic device 1604, or server 1608. For example, if electronic device 1601 is to automatically execute a function or service or to execute a function or service in response to a request from a user or another device, electronic device 1601 may request one or more external electronic devices to execute at least part of the function or service instead of executing the function or service, or may request one or more external electronic devices to execute at least part of the function or service in addition to executing the function or service. The one or more external electronic devices that receive the request may execute at least part of the requested function or service, or execute another function or service related to the request, and transmit the result of the execution to the electronic device 1601. The electronic device 1601 may provide the result as at least a partial response to the request, either by further processing the result or without further processing the result. To this end, for example, cloud computing technology, distributed computing technology, or client-server computing technology may be used.
[0226] According to various exemplary embodiments of the present disclosure, an electronic device (eg, Figure 5 The first intelligent server 200 may include: a communication circuit; a processor operatively connected to the communication circuit; and a memory operatively connected to the processor. According to various exemplary embodiments, the memory may store instructions that, when executed, cause the processor to control the electronic device: to receive data from a user terminal (e.g., Figure 5 receiving voice data from a user terminal 100); processing the voice data to generate text data; identifying the device to be controlled by analyzing the text data; and transmitting the data from a first external server (eg, Figure 5IoT server 500) receives information about an intelligent agent supported by the identified device; and based on the information about the intelligent agent supported by the identified device, determines whether to send the voice data to a second external server (e.g., a second external server supporting a second voice-based intelligent agent) Figure 5 The second intelligent server 600).
[0227] In an electronic device according to various exemplary embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device by analyzing the text data, determining the domain of the voice data, and identifying the device to be controlled based on the determined domain.
[0228] In an electronic device according to various example embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device: in response to the identified device supporting the first voice-based intelligent agent, determines the intent included in the voice data and the parameters corresponding to the intent by analyzing the text data.
[0229] In an electronic device according to various example embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device: in response to the identified device supporting the second voice-based intelligent agent, the voice data is sent to the second external server through the communication circuit.
[0230] In an electronic device according to various exemplary embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device: in response to the processor sending the voice data to the second external server, a transmission completion message is sent to the user terminal through the communication circuit.
[0231] In an electronic device according to various exemplary embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device to receive a voice command completion message related to the voice data from the second external server through the communication circuit, and send the voice command completion message to the user terminal.
[0232] In an electronic device according to various exemplary embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device to receive information about at least one device registered in the account and information about an intelligent agent supported by the at least one device from the first external server.
[0233] In an electronic device according to various example embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device: based on identifying the device intended to be controlled, requests information about the intelligent agent supported by the identified device from the first external server through the communication circuit; and in response to the request, receives information about the intelligent agent supported by the identified device from the first external server through the communication circuit.
[0234] According to various exemplary embodiments of the present disclosure, an electronic device (eg, Figure 6 The first intelligent server 200 may include: a communication circuit; a processor operatively connected to the communication circuit; and a memory ( Figure 6 According to various example embodiments, the memory may store information about at least one device registered in the account and information about an intelligent agent supported by the at least one device, and the memory may store instructions that, when executed, cause the processor to control the electronic device to: receive a call from a user terminal (e.g., Figure 6 receiving voice data via a user terminal 100 of the user terminal 100; processing the voice data to generate text data; identifying a device intended to be controlled by analyzing the text data; identifying information about an intelligent agent supported by the identified device based on information about the at least one device registered in the account and information about an intelligent agent supported by the at least one device stored in the memory; and determining, based on the identified information, whether to transmit the voice data to an external server (e.g., an external server supporting a second voice-based intelligent agent) via the communication circuit. Figure 6 Second smart server).
[0235] In the electronic device according to various exemplary embodiments of the present disclosure, the information about the at least one device registered in the account includes the name of the at least one device registered in the first voice-based intelligent agent and the second voice-based intelligent agent, respectively.
[0236] In an electronic device according to various exemplary embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device by analyzing the text data, determining the domain of the voice data, and identifying a device to be controlled based on the determined domain.
[0237] In an electronic device according to various example embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device: in response to the identified device supporting the first voice-based intelligent agent, determines the intent included in the voice data and the parameters corresponding to the intent by analyzing the text data.
[0238] In an electronic device according to various example embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device: in response to the identified device supporting the second voice-based intelligent agent, the voice data is sent to the external server through the communication circuit.
[0239] In an electronic device according to various exemplary embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device: in response to the processor sending the voice data to the external server, a transmission completion message is sent to the user terminal through the communication circuit.
[0240] In an electronic device according to various exemplary embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device to receive a user voice command completion message related to the user's voice data from the external server through the communication circuit, and send the user voice command completion message to the user terminal.
[0241] According to various exemplary embodiments of the present disclosure, an electronic device (eg, Figure 7 The first user terminal 70 may include: a communication circuit (eg, Figure 1 Communication interface 110), microphone (eg, Figure 1 a microphone 120 ), a processor operatively connected to the communication circuit and the microphone (e.g., Figure 1 processor 160) and a memory operatively connected to the processor (e.g., Figure 7 Database 740 or Figure 1According to various embodiments, the memory may store information about at least one device registered in the account and information about an intelligent agent supported by the at least one device, and the memory may store instructions that, when executed, cause the processor to control the electronic device to: receive voice data through the microphone; process the voice data to generate text data; identify a device intended to be controlled by analyzing the text data; identify information about an intelligent agent supported by the identified device based on the information about the at least one device registered in the account and the information about the intelligent agent supported by the at least one device stored in the memory; and determine, based on the information about the intelligent agent supported by the identified device, whether to send the voice data to an external electronic device that supports a second voice-based intelligent agent (e.g., a second voice-based intelligent agent) through the communication circuit. Figure 7 second user terminal 80).
[0242] In an electronic device according to various example embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device: in response to the identified device supporting the second voice-based intelligent agent, the voice data is sent to the external electronic device through the communication circuit.
[0243] In the electronic device according to various exemplary embodiments of the present disclosure, the electronic device may further include a speaker (eg, Figure 1 In the electronic device according to various exemplary embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device to output, through the speaker, a message including the following information: the processor has sent the voice data to the external electronic device in response to the processor sending the voice data to the external electronic device.
[0244] In the electronic device according to various exemplary embodiments of the present disclosure, the electronic device may further include a display (eg, Figure 1 In the electronic device according to various exemplary embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device to: in response to the processor sending the voice data to the external electronic device, output, through the display, a message including the following information: the processor has sent the voice data to the external electronic device.
[0245] In an electronic device according to various exemplary embodiments of the present disclosure, when the instructions are executed, the processor controls the electronic device to output, through the display, information about the second voice-based intelligent agent supported by the external electronic device and the content of the sent voice data.
[0246] The electronic device according to various embodiments may be one of various types of electronic devices. The electronic device may include, for example, a portable communication device (e.g., a smart phone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, a home appliance, etc. According to an embodiment of the present disclosure, the electronic device is not limited to those described above.
[0247] It should be understood that the various embodiments of the present disclosure and the terms used therein are not intended to limit the technical features set forth herein to specific embodiments, but rather include various changes, equivalents or alternative forms for the corresponding embodiments. For the description of the accompanying drawings, similar reference numerals may be used to refer to similar or related elements. It will be understood that the nouns in the singular form corresponding to the term may include one or more things, unless the relevant context clearly indicates otherwise. As used herein, each of the phrases such as "A or B", "at least one of A and B", "at least one of A or B", "A, B or C", "at least one of A, B and C" and "at least one of A, B or C" may include any one or all possible combinations of the items listed together with the corresponding phrase in the multiple phrases. As used herein, terms such as "1st" and "2nd" or "first" and "second" may be used to simply distinguish a corresponding component from another component, and do not limit the component in other aspects (e.g., importance or order). It will be understood that if an element (e.g., a first element) is referred to as being “coupled to another element (e.g., a second element)”, “coupled to another element (e.g., a second element)”, “connected to another element (e.g., a second element)”, or “connected to another element (e.g., a second element)”, whether or not the terms “operably” or “communicatively” are used, it means that the element may be directly (e.g., wired) connected to the other element, wirelessly connected to the other element, or connected to the other element via a third element.
[0248] As used herein, the term "module" may include units implemented in hardware, software, or firmware, and may be used interchangeably with other terms (e.g., "logic," "logic block," "component," or "circuit"). A module may be a single integrated component adapted to perform one or more functions or the smallest unit or portion of the single integrated component. For example, depending on an embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0249] The various embodiments described herein can be implemented as software (e.g., program 1640) comprising one or more instructions stored in a storage medium (e.g., internal memory 1636 or external memory 1638) that can be read by a machine (e.g., electronic device 1601). For example, under the control of a processor, a processor (e.g., processor 1620) of the machine (e.g., electronic device 1601) can call at least one of the one or more instructions stored in the storage medium and execute the at least one instruction with or without the use of one or more other components. This enables the machine to be operable to perform at least one function according to the called at least one instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The machine-readable storage medium can be provided in the form of a non-transitory storage medium. The term "non-transitory" only means that the storage medium is a tangible device and does not include signals (e.g., electromagnetic waves), but the term does not distinguish between data being semi-permanently stored in the storage medium and data being temporarily stored in the storage medium.
[0250] According to an embodiment, the method according to various embodiments of the present disclosure may be included and provided in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be released in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be downloaded via an application store (e.g., Play Store). TM ) The computer program product may be published online (e.g., downloaded or uploaded) or may be distributed (e.g., downloaded or uploaded) directly between two user devices (e.g., smartphones). If published online, at least part of the computer program product may be temporarily generated or at least part of the computer program product may be at least temporarily stored in a machine-readable storage medium (such as a memory of a manufacturer's server, an application store's server, or a forwarding server).
[0251] According to various embodiments, each component (e.g., module or program) in the above-mentioned parts may include a single entity or multiple entities. According to various embodiments, one or more components in the above-mentioned parts may be omitted, or one or more other components may be added. Alternatively or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, according to various embodiments, the integrated component may still perform the one or more functions of each component in the multiple components in the same or similar manner as a corresponding component in the multiple components performed one or more functions before integration. According to various embodiments, the operations performed by a module, program or another component may be performed sequentially, in parallel, repeatedly or in a heuristic manner, or one or more operations in the operations may be run in different orders or omitted, or one or more other operations may be added.
[0252] Although the present disclosure has been shown and described with reference to various exemplary embodiments thereof, it should be understood that the various exemplary embodiments are intended to be illustrative rather than restrictive. Those skilled in the art will further understand that various changes in form and details may be made without departing from the true spirit and full scope of the present disclosure, including the appended claims and their equivalents.
Claims
1. An electronic device configured to support a first voice-based intelligent agent, the electronic device comprising: Communication circuits; a processor operatively connected to the communication circuit; as well as a memory operatively connected to the processor, The memory stores instructions, which, when executed, enable the processor to control the electronic device: receiving voice data from a user terminal via the communication circuit; processing the speech data to generate text data; identifying a device intended to be controlled by analyzing the text data; receiving, via the communication circuit, information from a first external server regarding an intelligent agent supported by the identified device; and Based on the information about the intelligent agent supported by the identified device, a determination is made whether to transmit the voice data to a second external server supporting a second voice-based intelligent agent.
2. The electronic device according to claim 1, wherein When the instructions are executed, the processor further causes the electronic device to: Determining the domain of the voice data by analyzing the text data; and Based on the determined domain, the device intended to be controlled is identified.
3. The electronic device according to claim 2, wherein When the instructions are executed, the processor also controls the electronic device: in response to the identified device supporting the first voice-based intelligent agent, the processor determines the intent included in the voice data and the parameters corresponding to the intent by analyzing the text data.
4. The electronic device according to claim 1, wherein When the instructions are executed, the processor further causes the electronic device to: in response to the identified device supporting the second voice-based intelligent agent, send the voice data to the second external server through the communication circuit.
5. The electronic device according to claim 4, wherein When the instructions are executed, the processor controls the electronic device to send a transmission completion message to the user terminal through the communication circuit in response to the processor sending the voice data to the second external server. The electronic device according to claim 4 , wherein: When the instructions are executed, the processor further controls the electronic device to receive a voice command completion message related to the voice data from the second external server through the communication circuit, and send the voice command completion message to the user terminal.
7. The electronic device according to claim 1, wherein When executed, the instructions further cause the processor to control the electronic device to receive, from the first external server, information about at least one device registered in the account and information about an intelligent agent supported by the at least one device.
8. The electronic device according to claim 1, wherein When the instructions are executed, the processor further causes the electronic device to: upon identifying the device intended to be controlled, requesting information about the intelligent agent supported by the identified device from the first external server through the communication circuit; and In response to the request, the information about the intelligent agent supported by the identified device is received from the first external server through the communication circuit.
9. An electronic device configured to support a first voice-based intelligent agent, the electronic device comprising: Communication circuits; a processor operatively connected to the communication circuit; as well as a memory operatively connected to the processor, wherein the memory stores information about at least one device registered in the account and information about an intelligent agent supported by the at least one device, and The memory stores instructions, which, when executed, cause the processor to control the electronic device: receiving voice data from a user terminal via the communication circuit; processing the speech data to generate text data, By analyzing said text data, identifying the device intended to be controlled, identifying information about a smart agent supported by the identified device based on information about the at least one device registered in the account and information about a smart agent supported by the at least one device stored in the memory, and Based on the recognized information, a determination is made as to whether to transmit the voice data to an external server supporting a second voice-based intelligent agent via the communication circuit.
10. The electronic device according to claim 9, wherein The information about the at least one device registered in the account includes a name of the at least one device registered in the first voice-based intelligent agent and the second voice-based intelligent agent, respectively.
11. The electronic device according to claim 9, wherein When the instructions are executed, the processor further causes the electronic device to: Determining the domain of the speech data by analyzing the text data, and Based on the determined domain, devices intended to be controlled are identified.
12. The electronic device according to claim 11, wherein When the instructions are executed, the processor also controls the electronic device: in response to the identified device supporting the first voice-based intelligent agent, the processor determines the intent included in the voice data and the parameters corresponding to the intent by analyzing the text data.
13. The electronic device according to claim 9, wherein When executed, the instructions further cause the processor to control the electronic device: in response to the identified device supporting the second voice-based intelligent agent, send the voice data to the external server through the communication circuit.
14. The electronic device according to claim 13, wherein: When the instructions are executed, the processor controls the electronic device to send a transmission completion message to the user terminal through the communication circuit in response to the processor sending the voice data to the external server.
15. The electronic device according to claim 13, wherein When the instructions are executed, the processor further controls the electronic device to receive a voice command completion message related to the voice data from the external server through the communication circuit, and send the voice command completion message to the user terminal.
16. An electronic device configured to support a first voice-based intelligent agent, the electronic device comprising: Communication circuits; microphone; a processor operatively connected to the communication circuit and the microphone; as well as a memory operatively connected to the processor, wherein the memory is configured to store information about at least one device registered in the account and information about an intelligent agent supported by the at least one device, and The memory stores instructions that, when executed, cause the processor to control the electronic device: receiving voice data via the microphone, processing the speech data to generate text data, By analyzing said text data, identifying the device intended to be controlled, identifying information about a smart agent supported by the identified device based on information about the at least one device registered in the account and information about the smart agent supported by the at least one device stored in the memory, and Based on information about intelligent agents supported by the identified device, it is determined whether to transmit the voice data to an external electronic device supporting a second voice-based intelligent agent through the communication circuit.
17. The electronic device according to claim 16, wherein: When executed, the instructions further cause the processor to control the electronic device: in response to the identified device supporting the second voice-based intelligent agent, send the voice data to the external electronic device through the communication circuit.
18. The electronic device according to claim 17, further comprising a speaker, in, When the instructions are executed, the processor controls the electronic device: in response to the processor sending the voice data to the external electronic device, the speaker outputs the following message, the message including information that the processor has sent the voice data to the external electronic device.
19. The electronic device according to claim 17, further comprising a display, in, When the instructions are executed, the processor controls the electronic device: in response to the processor sending the voice data to the external electronic device, the display outputs the following message, the message including information that the processor has sent the voice data to the external electronic device.
20. The electronic device according to claim 19, wherein When the instructions are executed, the processor further causes the electronic device to output, through the display, information about the second voice-based intelligent agent supported by the external electronic device and the content of the transmitted voice data.
Citation Information
Patent Citations
Smart home appliance voice control method, apparatus and system
CN108183844A
Household appliance, control method thereof, server and storage medium
CN109799719A