Hub device and method performed by a hub device
The voice signal is received through the hub device and analyzed using ASR and NLU models, and the operation execution device is automatically determined and the corresponding operations are performed, solving the problems of inconvenience and slow response speed of users when inputting intentions through voice in a multi-device system, achieving a fast and economical voice control effect.
Patent Information
- Application Number
- CN202510384489.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-04
- Filing Date
- 2020-04-29
- Publication Date
- 2025-05-30
AI Technical Summary
In a multi-device system, when a user enters an intention through voice, the prior art requires registration and use of voice assistant services, resulting in inconvenience and slow response.
The voice signal is received through the hub device, the voice signal is converted into text using automatic speech recognition (ASR), the text is analyzed in combination with the natural language understanding (NLU) model, the operation execution device is automatically determined, and the text is provided to the identified device to perform the corresponding operation.
It realizes automatic identification and execution of voice input intentions without user registration, improves response speed and reduces network usage fees.
Smart Images

Figure CN120071934A_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application for an invention titled "Hub Device, Multi-Device System Including a Hub Device and a Plurality of Devices, and Operating Method Thereof", with an application date of April 29, 2020, an application number of 202080031822.X. Technical Field
[0002] The present disclosure relates to a hub device, a multi-device system including the hub device, and a method of operating the hub device. For example, according to an embodiment, the hub device may determine an operation execution device (e.g., an Internet of Things (IoT) device) for performing an operation according to a user's intention (e.g., included in a voice input received from the user) in a multi-device environment from among the hub device itself and a plurality of other electronic devices. According to an embodiment, the hub device may control the determined operation execution device. Background Art
[0003] With the development of multimedia technology and network technology, users can receive various services by using IoT devices (such as air purifiers or televisions (TVs)). In particular, with the development of voice recognition technology such as virtual personal assistant technology, users can input voice (e.g., utterances) to a device (e.g., a listening device) and can receive a response message to the voice input through a service providing agent (e.g., a virtual assistant).
[0004] However, in a multi-device system such as a home network environment including a plurality of IoT devices, when a user wants to receive a service by using an unregistered IoT device that has not been registered to interact through voice input or the like, the user has to inconveniently register the IoT device (e.g., including selecting the IoT device to provide the service). Specifically, because the types of services provided by a plurality of IoT devices are different, a technology that can recognize the intention included in the user's voice input and effectively provide a corresponding service is needed.
[0005] To recognize the intention included in the user's voice input, artificial intelligence (AI) technology may be used, and rule-based natural language understanding (NLU) technology may also be used. When a user's voice input is received through the hub device, since the hub device may not directly select a device for providing a service according to the voice input and has to control the device by using a separate voice assistant service providing server, the user may have to pay network usage fees, and the response speed is reduced because of using the voice assistant service providing server. Summary of the Invention
[0006] Solution to the Problem
[0007] The present disclosure relates to a method for controlling a device by a hub device, the method comprising: receiving a voice signal from a listener device; converting the received voice signal into text by performing automatic speech recognition (ASR); analyzing the text by using a first natural language understanding (NLU) model, and determining an operation execution device corresponding to the analyzed text by using a device determination model; identifying a device storing a function determination model corresponding to the determined operation execution device from the determined operation execution device and the listener device; and providing at least a part of the text to the identified device.
[0008] The present disclosure relates to a hub device for controlling a device, the hub device comprising: a communication interface configured to perform data communication with at least one of a voice assistant server or a plurality of devices including a listener device; a voice signal receiver configured to receive a voice signal from the listener device; a memory configured to store a program including one or more instructions; and a processor configured to execute the one or more instructions of the program stored in the memory, wherein the processor is further configured to: convert the received voice signal into text by performing automatic speech recognition (ASR), analyze the text by using a first natural language understanding (NLU) model, and determine an operation execution device corresponding to the analyzed text by using a device determination model, identify a device storing a function determination model corresponding to the determined operation execution device from the determined operation execution device and the listener device, and send at least a part of the text to the identified device by using the communication interface.
[0009] The present disclosure relates to a hub device, a multi-device system including the hub device and a plurality of devices, and an operation method thereof, and more particularly, to a hub device, a multi-device system, and an operation method thereof, in which the hub device receives a voice input of a user, automatically determines a device for performing an operation according to the received voice input according to the user's intention, and provides a plurality of pieces of information required for performing a service according to the determined device.
[0010] Additional aspects will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the embodiments presented in the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other aspects, features, and advantages of specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which like reference numerals represent like structural elements, wherein:
[0012] Figure 1 is a block diagram illustrating some elements of a multi-device system including a hub device, a voice assistant server, an Internet of Things (IoT) server, and a plurality of devices according to an embodiment of the present disclosure;
[0013] Figure 2 is a block diagram illustrating elements of a hub device according to an embodiment of the present disclosure;
[0014] Figure 3 is a block diagram illustrating elements of a voice assistant server according to an embodiment of the present disclosure;
[0015] Figure 4 is a block diagram illustrating elements of an IoT server according to an embodiment of the present disclosure;
[0016] Figure 5 is a block diagram illustrating some elements of multiple devices according to an embodiment of the present disclosure;
[0017] Figure 6 is a flowchart of a method for controlling a device based on a voice input performed by a hub device according to an embodiment of the present disclosure;
[0018] Figure 7 is a flowchart of a method for providing at least a part of text to one of a hub device, a voice assistant server, and an operation execution device according to a user's voice input performed by the hub device according to an embodiment of the present disclosure;
[0019] Figure 8 is a flowchart of a method for operating a hub device, a voice assistant server, an IoT server, and an operation execution device according to an embodiment of the present disclosure;
[0020] Figure 9 is a flowchart of a method for operating a hub device and an operation execution device according to an embodiment of the present disclosure;
[0021] Figure 10 is a flowchart of a method for operating a hub device and an operation execution device according to an embodiment of the present disclosure;
[0022] Figure 11 is a flowchart of a method for operating a hub device, a voice assistant server, an IoT server, a third - party IoT server, and a third - party device according to an embodiment of the present disclosure;
[0023] Figure 12A is a conceptual diagram illustrating operations of a hub device and multiple devices according to an embodiment of the present disclosure;
[0024] Figure 12B is a conceptual diagram illustrating operations of a hub device and multiple devices according to an embodiment of the present disclosure;
[0025] Figure 13It is a flowchart illustrating a method of a device in which a hub device determines an operation execution device based on a voice signal received from a listening device and sends text to a device storing a function determination model corresponding to the operation execution device according to an embodiment of the present disclosure;
[0026] Figure 14 It is a flowchart illustrating a method of a device in which a hub device sends text to a device storing a function determination model corresponding to an operation execution device according to an embodiment of the present disclosure;
[0027] Figure 15 It is a flowchart illustrating a method of operating a hub device, a voice assistant server, and a listening device according to an embodiment of the present disclosure;
[0028] Figure 16 It is a flowchart illustrating a method of operating a hub device, a voice assistant server, a listening device, and an operation execution device according to an embodiment of the present disclosure;
[0029] Figure 17 It is a diagram illustrating an example in which an operation execution device updates a function determination model according to an embodiment of the present disclosure;
[0030] Figure 18 It is a flowchart illustrating a method of operating a hub device, a voice assistant server, an IoT server, and an operation execution device according to an embodiment of the present disclosure;
[0031] Figure 19 It is a flowchart illustrating a method of operating a hub device, a voice assistant server, an IoT server, and a new device according to an embodiment of the present disclosure;
[0032] Figure 20 It is a diagram illustrating a multi-device system environment including a hub device, a voice assistant server, and multiple devices; and
[0033] Figure 21A and Figure 21B It is a diagram illustrating a voice assistant model executable by a hub device and a voice assistant server according to an embodiment of the present disclosure.
[0034] Best Mode
[0035] According to an embodiment of the present disclosure, a method of controlling a device based on voice input by a hub device includes: receiving a voice input of a user; converting the received voice input into text by performing automatic speech recognition (ASR); determining an operation execution device based on the text by using a device determination model; identifying a device storing a function determination model corresponding to the determined operation execution device among a plurality of devices connected to the hub device; and providing at least a part of the text to the identified device.
[0036] The device determination model may include a first natural language understanding (NLU) model configured to analyze text and determine an operation execution device based on the analysis result of the text.
[0037] The function determination model may include a second NLU model configured to analyze at least a part of the text and obtain operation information related to an operation to be performed by the determined operation execution device based on the analysis result of at least the part of the text.
[0038] The method may further include obtaining information about the function determination model stored in at least one device among a plurality of devices, where the function determination model is used to determine functions related to each of the plurality of devices.
[0039] The identification of the device may include: based on the obtained information about the function determination model, identifying the device storing the function determination model corresponding to the determined operation execution device.
[0040] According to another embodiment of the present disclosure, a hub device for controlling a device based on a voice input includes: a communication interface configured to perform data communication with at least one of a plurality of devices, a voice assistant server, and an Internet of Things (IoT) server; a microphone configured to receive a voice input of a user; a memory configured to store a program including one or more instructions; and a processor configured to execute one or more instructions of the program stored in the memory, where the processor is further configured to execute one or more instructions to convert the voice input received through the microphone into text by performing automatic speech recognition (ASR), determine an operation execution device from among the plurality of devices based on the text by using the device determination model, identify the device storing the function determination model corresponding to the determined operation execution device by using the function determination device determination module, and control the communication interface to provide at least a part of the text to the identified device.
[0041] The device determination model may include a first natural language understanding (NLU) model configured to analyze text and determine an operation execution device based on the analysis result of the text.
[0042] The function determination model may include a second NLU model configured to analyze at least a part of the text and obtain operation information related to an operation to be performed by the determined operation execution device based on the analysis result of at least the part of the text.
[0043] The processor may also be configured to execute one or more instructions to control the communication interface to obtain information about a function determination model stored in at least one device from among a plurality of devices, where the function determination model is used to determine functions associated with each of the plurality of devices.
[0044] The processor may also be configured to execute one or more instructions to identify a device storing a function determination model corresponding to a determined operation execution device based on the obtained information about the function determination model.
[0045] According to another embodiment of the present disclosure, a method of an operating system, the system including a hub device and a first device storing a function determination model, the method including: receiving, by the hub device, a voice input of a user; converting the received voice input into text by performing automatic speech recognition (ASR) using data about an ASR module stored in a memory of the hub device; determining, based on the text, the first device as an operation execution device by using data about a device determination model stored in the memory of the hub device; obtaining, by the hub device, information about the function determination model stored in the first device from the first device; and sending, by the hub device, at least a portion of the text to the first device based on the obtained information about the function determination model.
[0046] The device determination model may include a first natural language understanding (NLU) model configured to analyze the text and determine the first device as the operation execution device from among a plurality of devices based on an analysis result of the text.
[0047] The function determination model may include a second NLU model configured to analyze at least a portion of the text received from the hub device and obtain operation information related to an operation to be performed by the first device based on an analysis result of at least a portion of the text.
[0048] The method may further include: analyzing, by the first device, at least a portion of the text by using the second NLU model of the function determination model and obtaining operation information related to an operation to be performed by the first device based on an analysis result of at least a portion of the text.
[0049] The method may further include: generating, by the first device, a control command for controlling an operation of the first device based on the operation information; and performing, by the first device, the operation based on the control command.
[0050] According to another embodiment of the present disclosure, a multi-device system includes a hub device and a first device storing a function determination model. The hub device includes: a communication device configured to perform data communication with the first device; a microphone configured to receive a voice input from a user; a memory configured to store a program including one or more instructions; and a processor configured to execute one or more instructions of the program stored in the memory. The processor is further configured to execute one or more instructions to convert the voice input received through the microphone into text by performing automatic speech recognition (ASR), determine the first device as an operation execution device based on the text by using a device determination model, control a communication interface to obtain information about the function determination model stored in the first device from the first device, and based on the obtained information about the function determination model, control the communication interface to send at least a portion of the text to the first device.
[0051] The device determination model may include a first natural language understanding (NLU) model configured to analyze the text and determine the first device as an operation execution device from multiple devices based on the analysis result of the text.
[0052] The first device may include a communication interface configured to receive at least a portion of the text from the hub device. The function determination model includes a second NLU model configured to analyze at least a portion of the received text and obtain operation information related to an operation to be performed by the first device based on the analysis result of at least a portion of the text.
[0053] The first device may further include a processor configured to analyze at least a portion of the text by using the second NLU model and obtain operation information to be performed by the first device based on the analysis result of at least a portion of the text.
[0054] The processor of the first device may further be configured to control at least one element of the first device to generate a control command for controlling the operation of the first device based on the operation information and perform the operation based on the control command.
[0055] According to another embodiment of the present disclosure, a method for controlling a device by a hub device includes: receiving a voice signal from a listening device; converting the received voice signal into text by performing automatic speech recognition (ASR); analyzing the text by using a first natural language understanding (NLU) model and determining an operation execution device corresponding to the analyzed text by using a device determination model; identifying a device storing a function determination model corresponding to the determined operation execution device from the determined operation execution device and the listening device; and providing at least a portion of the text to the identified device.
[0056] The device determination model may include a first NLU model configured to analyze text and determine an operation execution device based on the analysis result of the text.
[0057] The function determination model may include a second NLU model configured to analyze at least a part of the text and obtain operation information related to an operation to be executed by the determined operation execution device based on the analysis result of at least a part of the text.
[0058] The method may further include determining whether the determined operation execution device is the same as the listening device.
[0059] Based on determining that the operation execution device is the same as the listening device, identifying the device storing the function determination model may include: obtaining function determination model information on whether the listening device stores the function determination model in an internal memory; and based on the obtained function determination model information, determining the listening device as the device storing the function determination model.
[0060] Based on determining that the operation execution device is a device different from the listening device, identifying the device storing the function determination model may include: obtaining function determination model information on whether the determined operation execution device stores the function determination model in an internal memory; and based on the obtained function determination model information, determining whether the operation execution device is the device storing the function determination model.
[0061] The method may further include: receiving update data of the device determination model from a voice assistant server; and updating the device determination model by using the received update data.
[0062] The update data may include data for updating the device determination model based on update information of the function determination model included in at least one of the operation execution device or the listening device to determine an updated function from the text and determine an operation execution device corresponding to the updated function.
[0063] The method may further include: receiving device information of a new device from a voice assistant server, where the device information of the new device includes at least one of device identification information of the new device, storage information of the device determination model, and storage information of the function determination model; and updating the device determination model by adding the new device to device candidates that can be determined as operation execution devices by the device determination model by using the received device information of the new device.
[0064] According to another embodiment of the present disclosure, a hub device for controlling a device includes: a communication interface configured to perform data communication with at least one of a voice assistant server and a plurality of devices including listening devices; a voice signal receiver configured to receive a voice signal from the listening device; a memory configured to store a program including one or more instructions; and a processor configured to execute one or more instructions of the program stored in the memory, wherein the processor is further configured to convert the received voice signal into text by performing automatic speech recognition (ASR), analyze the text by using a first natural language understanding (NLU) model, and determine an operation execution device corresponding to the analyzed text by using a device determination model, identify a device storing a function determination model corresponding to the determined operation execution device from the determined operation execution device and the listening device, and send at least a part of the text to the identified device by using the communication interface.
[0065] The device determination model may include a first NLU model configured to analyze the text and determine an operation execution device based on the analysis result of the text.
[0066] The function determination model may include a second NLU model configured to analyze at least a part of the text and obtain operation information related to an operation to be performed by the determined operation execution device based on the analysis result of at least the part of the text.
[0067] The processor may further be configured to determine whether the determined operation execution device is the same as the listening device.
[0068] Based on determining that the operation execution device is the same as the listening device, the processor may further be configured to obtain function determination model information on whether the listening device stores the function determination model in an internal memory, and determine the listening device as the device storing the function determination model based on the obtained function determination model information.
[0069] Based on determining that the operation execution device is a device different from the listening device, the processor may further be configured to obtain function determination model information on whether the determined operation execution device stores the function determination model in an internal memory, and determine whether the operation execution device is the device storing the function determination model based on the obtained function determination model information.
[0070] The processor may further be configured to receive update data of the device determination model from the voice assistant server by using the communication interface, and update the device determination model by using the received update data.
[0071] The updated data may include the following, where the data is used to update a device determination model based on update information of a model determined based on functions included in at least one of an operation execution device and a listening device, to determine an updated function from text and determine an operation execution device corresponding to the updated function.
[0072] The processor may also be configured to update the device determination model by receiving, via a communication interface, device information of a new device including at least one of device identification information of the new device, storage information of the device determination model, and storage information of the function determination model from a voice assistant server, and adding the new device to a device candidate that can be determined as an operation execution device by the device determination model by using the received device information of the new device.
[0073] According to an embodiment of the present disclosure, a method may include: based on a voice input of a user received by a hub device: converting the received voice input into text by the hub device by performing automatic speech recognition (ASR); identifying a device capable of performing an operation corresponding to the text by the hub device; identifying, from the hub device and a plurality of other devices connected to the hub device, which device stores a function determination model corresponding to the device capable of performing the operation corresponding to the text; and sending at least a part of the text to the identified device based on the identified device storing the function determination model being a device different from the hub device, where the hub device includes a hardware processor.
[0074] The method may further include using a first natural language understanding (NLU) model to analyze the text and determining, based on an analysis result of the text, a device capable of performing the operation, where the NLU is the device determination model.
[0075] The method may further include using a second natural language understanding (NLU) model included in the device determination model to analyze at least a part of the text and obtaining operation information related to the operation corresponding to the text based on an analysis result of at least the part of the text.
[0076] The method may further include obtaining information about the function determination model stored in at least one device from at least one device storing the function determination model.
[0077] Identifying which device stores the function determination model may include: based on the obtained information about the function determination model, identifying a device that stores the function determination model corresponding to the identified device capable of performing the operation.
[0078] According to an embodiment, a hub device for controlling a device based on voice input may include: a communication interface configured to perform data communication with at least one of a plurality of devices, a voice assistant server, and an Internet of Things (IoT) server; a microphone configured to receive a voice input from a user; a memory configured to store a program including one or more instructions; and a processor configured to execute one or more instructions of the program stored in the memory to: convert the voice input received through the microphone into text by performing automatic speech recognition (ASR), identify a device capable of performing an operation corresponding to the text; identify, from the hub device and a plurality of other devices connected to the hub device, which device stores a function determination model corresponding to the device capable of performing the operation corresponding to the text; and based on the identified device storing the function determination model being a device different from the hub device, control the communication interface to send at least a portion of the text to the identified device storing the function determination model, wherein the hub device includes a hardware processor.
[0079] The processor may also be configured to use a device determination model including a first natural language understanding (NLU) model configured to analyze the text to identify a device capable of performing an operation corresponding to the text, and determine the device capable of performing the operation corresponding to the text based on the analysis result of the text.
[0080] The function determination model may include a second NLU model configured to analyze at least a portion of the text and obtain operation information related to an operation to be performed by a device capable of performing an operation corresponding to the text based on the analysis result of at least a portion of the text.
[0081] The processor may also be configured to execute one or more instructions to control the communication interface to obtain information about the function determination model stored in at least one device storing the function determination model among a plurality of devices, wherein the function determination model is used to determine functions related to each of the plurality of devices.
[0082] The processor may also be configured to execute one or more instructions to identify, based on the obtained information about the function determination model, which device stores a function determination model corresponding to the device capable of performing the operation corresponding to the text.
[0083] According to an embodiment, a method for an operating system, the system including a hub device and a first device storing a function determination model, the method including: based on a voice input of a user received by the hub device: converting the received voice input into text by the hub device performing automatic speech recognition (ASR) by using data of an ASR module stored in a memory of the hub device; identifying, by the hub device, the first device as a device capable of performing an operation corresponding to the text by using data stored in a memory of the hub device; obtaining, by the hub device, information about the function determination model stored in the first device from the first device; and sending, by the hub device, at least a part of the text to the first device based on the obtained information about the function determination model.
[0084] The method may further include: analyzing the text by using a first natural language understanding (NLU) model, and determining, based on an analysis result of the text, the first device from among a plurality of devices as a device capable of performing an operation corresponding to the text.
[0085] The method may further include performing analysis by using the function determination model, including: analyzing at least a part of the text received from the hub device by using a second natural language understanding (NLU) model included in the device determination model, and obtaining operation information related to an operation to be performed by the first device based on an analysis result of at least the part of the text.
[0086] The method may further include: analyzing at least a part of the text by using a second NLU model of the function determination model by the first device, and obtaining operation information related to the operation to be performed by the first device based on an analysis result of at least the part of the text.
[0087] The method may further include: generating, by the first device, a control command for controlling an operation of the first device based on the operation information; and performing an operation by the first device based on the control command.
[0088] According to an embodiment, a multi-device system may include a hub device and a first device storing a function determination model. The hub device includes: a communication interface configured to perform data communication with the first device storing the function determination model; a microphone configured to receive a voice input of a user; a memory configured to store a program including one or more instructions; and a processor configured to execute one or more instructions of the program stored in the memory to: convert the voice input received through the microphone into text by performing automatic speech recognition (ASR), identify the first device as a device capable of performing an operation corresponding to the text, control the communication interface to obtain information about the function determination model stored in the first device from the first device, and based on the obtained information about the function determination model, control the communication interface to send at least a portion of the text to the first device, wherein the hub device includes a hardware processor.
[0089] The processor may also be configured to analyze the text by using a device determination model including a first natural language understanding (NLU) model, identify the first device as a device capable of performing an operation corresponding to the text, and based on the analysis result of the text, determine the first device among a plurality of devices as a device capable of performing an operation corresponding to the text.
[0090] The first device may include a communication interface configured to receive at least a portion of the text from the hub device. The function determination model includes a second NLU model configured to analyze at least a portion of the received text and obtain operation information related to an operation to be performed by the first device based on the analysis result of at least a portion of the text.
[0091] The first device may further include a processor configured to analyze at least a portion of the text by using the second NLU model and obtain operation information to be performed by the first device based on the analysis result of at least a portion of the text.
[0092] The processor of the first device may also be configured to control at least one element of the first device to generate a control command for controlling the operation of the first device based on the operation information and perform an operation based on the control command.
[0093] According to an embodiment, a method may include: based on a voice signal received by a hub device from a listening device: converting, by the hub device, the received voice signal into text by performing automatic speech recognition (ASR); analyzing the text by using a first natural language understanding (NLU) model and identifying, by using a device determination model, a device capable of performing an operation corresponding to the analyzed text; identifying, from among the device capable of performing the operation and the listening device, which device stores a function determination model corresponding to the device capable of performing the operation; and sending at least a portion of the text to the identified device.
[0094] The method may further include analyzing text using a first NLU model included in the device determination model and determining a device capable of performing an operation based on the analysis result of the text.
[0095] The method may further include analyzing at least a part of the text using a second natural language understanding (NLU) model included in the device determination model and obtaining operation information related to an operation to be performed by a device capable of performing an operation based on the analysis result of at least the part of the text.
[0096] The method may further include determining whether the device capable of performing an operation is the same as the listening device.
[0097] The method may further include, based on determining that the device capable of performing an operation is the same as the listening device, identifying a device storing the function determination model, including: obtaining function determination model information on whether the listening device stores the function determination model in an internal memory; and determining the listening device as the device storing the function determination model based on the obtained function determination model information.
[0098] The method may further include, based on determining that the device capable of performing an operation is a device different from the listening device, identifying a device storing the function determination model, including: obtaining function determination model information on whether the device capable of performing an operation stores the function determination model in an internal memory; and determining whether the device capable of performing an operation is the device storing the function determination model based on the obtained function determination model information.
[0099] The method may further include: receiving update data of the device determination model from a voice assistant server; and updating the device determination model by using the received update data.
[0100] The update data may include data for updating the device determination model to determine updated functions from text and determining a device capable of performing an operation corresponding to the updated functions based on update information of the function determination model included in at least one of the device capable of performing an operation and the listening device.
[0101] The method may further include: receiving device information of a new device from a voice assistant server, the device information of the new device including at least one of device identification information of the new device, storage information of the device determination model, and storage information of the function determination model; and updating the device determination model by adding the new device to a device candidate of devices that can be determined by the device determination model to be capable of performing an operation by using the received device information of the new device.
[0102] According to an embodiment, a hub device may include: a communication interface configured to perform data communication with at least one of a voice assistant server and a plurality of devices including listening devices; a voice signal receiver configured to receive a voice signal from a listening device; a memory configured to store a program including one or more instructions; and a processor configured to execute one or more instructions of the program stored in the memory to: convert the received voice signal into text by performing automatic speech recognition (ASR), analyze the text by using a first natural language understanding (NLU) model, and identify, by using a device determination model, a device capable of performing an operation corresponding to the analyzed text, identify, from the devices capable of performing the operation and the listening device, which device stores a function determination model corresponding to the device capable of performing the operation, and send at least a portion of the text to the identified device storing the function determination model by using the communication interface.
[0103] The device determination model may include a first NLU model configured to analyze the text and determine, based on the analysis result of the text, a device capable of performing an operation.
[0104] The function determination model may include a second NLU model configured to analyze at least a portion of the text and obtain, based on the analysis result of at least a portion of the text, operation information related to an operation to be performed by the device capable of performing the operation.
[0105] The processor may further be configured to determine whether the device capable of performing the operation is the same as the listening device.
[0106] The processor may further be configured to: based on determining that the device capable of performing the operation is the same as the listening device, the processor is further configured to obtain function determination model information regarding whether the listening device stores the function determination model in an internal memory, and determine the listening device as the device storing the function determination model based on the obtained function determination model information.
[0107] The processor may further be configured to: based on determining that the device capable of performing the operation is a device different from the listening device, the processor is further configured to obtain function determination model information regarding whether the device capable of performing the operation stores the function determination model in an internal memory, and determine whether the device capable of performing the operation is the device storing the function determination model based on the obtained function determination model information.
[0108] The processor may further be configured to: receive, by using the communication interface, update data of the device determination model from the voice assistant server, and update the device determination model by using the received update data.
[0109] The updated data may include the following data for updating a device determination model to determine an updated function from text and determining, based on the update information of the function determination model included in at least one of a device capable of performing an operation and a listening device, a device capable of performing an operation corresponding to the updated function.
[0110] The processor may also be configured to: receive, via a communication interface, device information of a new device from a voice assistant server, the device information of the new device including at least one of device identification information of the new device, storage information of a device determination model, and storage information of a function determination model, and update the device determination model by adding the new device to a device candidate of devices determined by the device determination model to be capable of performing an operation by using the received device information of the new device.
[0111] According to an embodiment, a method may include: based on a voice of a user detected by a hub device: converting, by the hub device, a received voice input into text by performing automatic speech recognition (ASR); identifying, by the hub device, an intention of the user; identifying, by the hub device, an Internet of Things (IoT) device capable of performing an operation corresponding to the text; identifying, among the hub device and a plurality of other devices connected to the hub device, which device stores a function determination model corresponding to the IoT device capable of performing the operation corresponding to the text; and sending at least a portion of the text to the identified device based on the identified device storing the function determination model being different from the hub device, wherein the hub device includes a hardware processor.
[0112] The method may further include: by the hub device, storing information as to whether a function determination model of each of a plurality of IoT devices previously registered using a user account is associated with information on a storage location of the function determination model of each of the plurality of IoT devices in the form of a lookup table (LUT).
[0113] The method may further include: obtaining, by using the function determination model, operation information related to an operation corresponding to the text to be performed by the IoT device. Detailed Description
[0114] This application is based on and claims the benefit of U.S. Provisional Patent Application Nos. 62 / 862,201 filed on Jun. 17, 2019 and 62 / 905,707 filed on Sep. 25, 2019, and claims priority to Korean Patent Application Nos. 10-2019-0051824 filed on May 2, 2019, 10-2019-0123310 filed on Oct. 4, 2019, and 10-2020-0027217 filed on Mar. 4, 2020, the disclosures of which are incorporated herein by reference in their entireties.
[0115] Although the terms used herein are selected from common terms currently widely used in view of their functions in the present disclosure, these terms may change according to the intention of those of ordinary skill in the art, precedent, or the emergence of new technologies. In addition, in specific cases, the terms are arbitrarily selected by the applicant of the present disclosure, and the meanings of these terms will be described in detail in the corresponding parts of the specific embodiments. Therefore, the terms used in the present disclosure are not merely designations of terms, but rather define the terms based on the meanings of the terms and the content throughout the present disclosure.
[0116] Throughout the disclosure, the expression "at least one of a, b, and c" means only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or variations thereof.
[0117] As used herein, unless the context clearly dictates otherwise, the singular forms are intended to also include the plural forms. The terms used herein, including technical and scientific terms, have the same meanings as those commonly understood by those of ordinary skill in the art to which this disclosure pertains.
[0118] Throughout this application, when a component "includes" an element, it should be understood that, unless there is a specific contrary statement, the component additionally includes other elements rather than excluding other elements. In addition, terms such as "... unit", "module", etc. used in the present disclosure indicate a unit that processes at least one function or movement, and the unit can be implemented as hardware or software, or a combination of hardware and software.
[0119] The expression "configured to (or set to)" used herein can be replaced, depending on the circumstances, with, for example, "suitable for", "capable of...", "designed to", "adapted to", "manufactured to", or "able to". The expression "configured to (or set to)" does not necessarily mean "specially designed for" in hardware. Instead, in some cases, the expression "a system configured to..." can mean that the system "is capable of..." together with other devices or components. For example, "a processor configured to (or set to) execute A, B, and C" can refer to a dedicated processor for executing the corresponding operations (e.g., an embedded processor), or a general-purpose processor (e.g., a central processing unit (CPU) or an application processor (AP)) capable of executing the corresponding operations by executing one or more software programs stored in a memory.
[0120] According to an embodiment of the present disclosure, the term "first natural language understanding (NLU) model" used herein may refer to a model trained to analyze text converted from a speech input and determine an operation execution device based on the analysis result. The first NLU model can be used to determine an intention by interpreting the text and determine an operation execution device based on the intention.
[0121] According to an embodiment of the present disclosure, the term "second NLU model" as used herein may refer to a model trained to analyze text related to a specific device. The second NLU model may be a model trained to obtain operation information related to an operation to be performed by a specific device by interpreting at least a part of the text. The storage capacity of the second NLU model may be greater than the storage capacity of the first NLU model.
[0122] According to an embodiment of the present disclosure, the term "intention" as used herein may refer to information indicating a user's intention determined by interpreting text. The intention as information indicating the intention of the user's utterance may be information indicating an operation of a device that executes a requested operation by the user. The intention may be determined by interpreting the text using an NLU model. For example, based on the text converted from a user's voice input being "Play the movie Avengers on TV", the intention may be determined to be "content playback". Alternatively, based on the text converted from a user's voice input being "Lower the air conditioner temperature to 18 °C", the intention may be determined to be "temperature control".
[0123] According to an embodiment of the present disclosure, the intention may include not only information indicating the intention of the user's utterance (hereinafter, referred to as intention information), but also a numerical value corresponding to the information indicating the user's intention. The numerical value may indicate the probability that the text is related to the information indicating a specific intention. After interpreting the text using an NLU model, when multiple pieces of intention information indicating the user's intention are obtained, the intention information having the maximum numerical value among the multiple numerical values corresponding to the multiple pieces of intention information may be determined as the intention.
[0124] According to an embodiment of the present disclosure, the term "operation" of a device as used herein may refer to at least one action performed by the device when the device executes a specific function. According to an embodiment of the present disclosure, the operation may indicate at least one action performed by the device when the device executes an application. For example, when the device executes an application, the operation may indicate, for example, one of the following: video playback, music playback, email creation, weather information reception, news information display, playing a game, and performing photography. However, the operation is not limited to the above examples.
[0125] According to an embodiment of the present disclosure, the operation of the device may be performed based on information about a detailed operation output from an action plan management module. According to an embodiment of the present disclosure, the device may perform at least one action by executing a function corresponding to the detailed operation output from the action plan management module. According to an embodiment of the present disclosure, the device may store instructions for executing a function corresponding to the detailed operation, and when the detailed operation is determined, the device may determine an instruction corresponding to the detailed operation and may execute a specific function by executing the instruction.
[0126] In addition, according to an embodiment, the device may store instructions for executing an application corresponding to a detailed operation. According to an embodiment of the present disclosure, the instructions for executing an application may include instructions for executing the application itself and instructions for executing detailed functions constituting the application. Based on determining the detailed operation, the device may execute the application by executing the instructions for executing the application corresponding to the detailed operation, and may execute the detailed function by executing the instructions for executing the detailed function of the application corresponding to the detailed operation.
[0127] According to an embodiment of the present disclosure, the term "operation information" used herein may refer to information related to a detailed operation to be determined by the device, a relationship between each detailed operation and another detailed operation, and an execution order of the detailed operation. According to an embodiment of the present disclosure, when a first operation is to be executed, the relationship between each detailed operation and another detailed operation may include information about a second operation that has to be executed before executing the first operation. For example, when the operation to be executed is "music playback", "power on" may be another detailed operation that has to be executed before executing "music playback". According to an embodiment of the present disclosure, the operation information may include, but is not limited to, one or more of the following items: functions to be executed by the operation execution device to execute a specific operation, an execution order of the functions, input values required to execute the functions, and output values output as a result of the execution of the functions.
[0128] According to an embodiment of the present disclosure, the term "operation execution device" used herein may refer to a device determined to execute an operation based on an intention obtained from text among a plurality of devices. The text may be analyzed by using a first NLU model, and the operation execution device may be determined based on the analysis result. According to an embodiment of the present disclosure, the operation execution device may execute at least one action by executing a function corresponding to a detailed operation output from the action plan management module. According to an embodiment of the present disclosure, the operation execution device may execute an operation based on the operation information.
[0129] According to an embodiment of the present disclosure, the term "action plan management module" used herein may refer to a module for managing detailed operations to be executed by the operation execution device and operation information related to the detailed operations of the device to generate an execution order of the detailed operations. According to an embodiment of the present disclosure, the action plan management module may manage operation information about detailed operations of the device according to the device type and the relationship between the detailed operations.
[0130] According to an embodiment of the present disclosure, the term "Internet of Things (IoT) server" as used herein may refer to a server that obtains, stores, and manages IoT device information regarding each of a plurality of devices (e.g., including IoT devices, mobile phones, etc.). The IoT server may obtain, determine, or generate a control command for controlling a device (e.g., an IoT device) by using the stored device information. According to an embodiment of the present disclosure, the IoT server may send a control command to a device determined to perform an operation based on operation information. According to an embodiment of the present disclosure, the IoT server may be implemented as, but not limited to, a hardware device independent of the "server" of the present disclosure. According to an embodiment, the IoT server may be an element of the "voice assistant server" of the present disclosure, or may be a server designed to be classified as software. Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those of ordinary skill in the art can easily implement and practice the present disclosure. However, according to an embodiment of the present disclosure, the present disclosure may be implemented in many different forms and should not be construed as limited to the embodiments of the present disclosure set forth herein.
[0131] Reference will now be made in detail to embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings.
[0132] Figure 1 is a block diagram illustrating some elements of a multi-device system including a hub device 1000, a voice assistant server 2000, an IoT server 3000, and a plurality of devices 4000 according to an embodiment of the present disclosure.
[0133] In Figure 1 the embodiment, elements for describing the operations of the hub device 1000, the voice assistant server 2000, the IoT server 3000, and the plurality of devices 4000 are illustrated. The elements included in the hub device 1000, the voice assistant server 2000, the IoT server 3000, and the plurality of devices 4000 are not limited to Figure 1 those illustrated in
[0134] Figure 1 The reference numerals S1 to S16 marked by arrows in
[0135] Reference Figure 1, according to an embodiment of the present disclosure, the hub device 1000, the voice assistant server 2000, the IoT server 3000, and the plurality of devices 4000 may be connected to each other by using a wired communication or wireless communication method and may perform communication. In an embodiment of the present disclosure, the hub device 1000 and the plurality of devices 4000 may be directly connected to each other or connected via a communication network, but the present disclosure is not limited thereto. According to an embodiment, the hub device 1000 and the plurality of devices 4000 may be connected to the voice assistant server 2000, and the hub device 1000 may be connected to the plurality of devices 4000 through the voice assistant server 2000. In addition, according to an embodiment, the hub device 1000 and the plurality of devices 4000 may be connected to the IoT server 3000. In another embodiment of the present disclosure, each of the hub device 1000 and the plurality of devices 4000 may be connected to the voice assistant server 2000 through a communication network and may be connected to the IoT server 3000 through the voice assistant server 2000.
[0136] According to an embodiment, the hub device 1000, the voice assistant server 2000, the IoT server 3000, and the plurality of devices 4000 may be connected by one or more of the following: a local area network (LAN), a wide area network (WAN), a value-added network (VAN), a mobile radio communication network, a satellite communication network, or a combination thereof. Examples of wireless communication methods may include but are not limited to Wi-Fi (Wireless Fidelity), Bluetooth, Bluetooth Low Energy (BLE), Zigbee, Wi-Fi Direct (WFD), Ultra Wideband (UWB), Infrared Data Association (IrDA), or Near Field Communication (NFC).
[0137] According to an embodiment, the hub device 1000 may be a device that receives a voice input from a user and controls at least one of the plurality of devices 4000 based on the received voice input. According to an embodiment of the present disclosure, the hub device 1000 may be a listening device that receives a voice input from a user.
[0138] According to an embodiment, at least one of the plurality of devices 4000 may be an operation execution device that performs a specific operation by receiving a control command from the hub device 1000 or the IoT server 3000. According to an embodiment of the present disclosure, the plurality of devices 4000 may be IoT devices that log in using the same user account as the user account of the hub device 1000 and are previously registered with the IoT server 3000 using the user account of the hub device 1000.
[0139] According to an embodiment, at least one of the plurality of devices 4000 may be a listening device that receives voice data (e.g., a voice input from a user). The listening device may be, but is not limited to, a device designed to process user voice (e.g., a device that only receives and processes voice inputs from a human user or from a specific registered user). In an embodiment of the present disclosure, the listening device may be an operation execution device that receives a control command from the hub device 1000 and performs an operation for a specific function.
[0140] According to an embodiment of the present disclosure, at least one of the plurality of devices 4000 may receive a control command (S12, S14, and S16) from the IoT server 3000, or may receive at least a part of the text converted from the input voice from the hub device 1000 (S3 and S5). According to an embodiment of the present disclosure, at least one of the plurality of devices 4000 may receive a control command (S12, S14, and S16) from the IoT server 3000 without receiving at least a part of the text from the hub device 1000.
[0141] According to an embodiment, the hub device 1000 may include a device determination model 1330 that determines a device for performing an operation based on a user's voice input. According to an embodiment, the device determination model 1330 may determine an operation execution device from among the plurality of devices 4000 registered according to the user account. In an embodiment of the present disclosure, the hub device 1000 may receive, from the voice assistant server 2000, device information (S2) including at least one of identification information (e.g., device id information) of each of the plurality of devices 4000, the device type of each of the plurality of devices 4000, the function execution ability of each of the plurality of devices 4000, location information, and status information. According to an embodiment, each of the plurality of devices 4000 is an IoT device. According to an embodiment, the hub device 1000 may determine, based on the received device information and by using data regarding the device determination model 1330, a device (e.g., an IoT device) for performing an operation according to the user's voice input from among the plurality of devices 4000.
[0142] According to an embodiment of the present disclosure, a function determination model corresponding to the operation execution device determined by the hub device 1000 may be stored in the memory 1300 of the hub device 1000 (see Figure 2 ), may be stored in the operation execution device itself, or may be stored in the memory 2300 of the voice assistant server 2000 (see Figure 3 ). According to an embodiment of the present disclosure, the term "function determination model" corresponding to each device refers to a model for obtaining operation information regarding detailed operations for performing an operation according to the determined function of the device and the relationships between the detailed operations.
[0143] According to an embodiment of the present disclosure, the function determination device determination module 1340 of the hub device 1000 may identify, from among the hub device 1000, the voice assistant server 2000, and the operation execution device, a device in which a function determination model corresponding to the operation execution device is stored by using a database 1360 including information on the function determination model of the device stored in the memory. According to an embodiment of the present disclosure, the database 1360 may include information on a plurality of devices 4000 registered with a user account associated with the hub device 1000. Specifically, identification information (e.g., device identifier (ID) information) of each of the plurality of devices 4000, information on whether a function determination model of each of the plurality of devices 4000 exists, and information on the storage location of the function determination model of each of the plurality of devices 4000 (e.g., identification information of the stored device / server, Internet protocol (IP) address of the stored device / server, or media access control (MAC) address of the stored device / server) may be stored. In an embodiment of the present disclosure, the function determination device determination module 1340 may search the database 1360 according to the device identification information of the operation execution device output by the device determination model 1330, and may obtain information on the storage location of the function determination model corresponding to the operation execution device based on the search result of the database 1360.
[0144] According to an embodiment of the present disclosure, the hub device 1000 may transmit at least a part of the text converted from the user's voice input to the device identified as storing the function determination model corresponding to the operation execution device by using the function determination device determination module 1340.
[0145] For example, based on the voice input received by the hub device 1000 that the user says "Increase the temperature by 1°C", the hub device 1000 can determine, through the device determination model 1330, that the device for increasing the temperature is an air conditioner. Next, according to an embodiment, the function determination device determination module 1340 can check whether the function determination model corresponding to the air conditioner is stored in the air conditioner (IoT device), and based on determining that it is stored in the air conditioner, can send the text corresponding to "Increase the temperature by 1°C" to the air conditioner. According to an embodiment, the IoT device (e.g., an air conditioner) capable of performing the operation of increasing the air temperature can analyze the received text through the stored function determination model corresponding to the air conditioner, and perform a temperature control operation by using the text analysis result. That is, based on the device determination model 1330, it is determined that the operation execution device capable of performing the user's voice input is the first device 4100 as an "air conditioner". Since the function determination model corresponding to the air conditioner is stored in the first device 4100 itself, the function determination device determination module 1340 can send at least a part of the text to the first device 4100 (e.g., an air conditioner) (S3).
[0146] For example, based on the voice input received by the hub device 1000 that the user says "Change the channel", the hub device 1000 can determine, through the device determination model 1330, that the device for changing the channel is a TV. Next, according to an embodiment, the function determination device determination module 1340 can check whether the function determination model corresponding to the TV is stored in the hub device 1000, and can analyze the text corresponding to "Change the channel" through the stored function determination model corresponding to the TV. According to an embodiment, the hub device 1000 can determine the channel change operation as an operation to be performed by the TV by using the text analysis result, and can send the operation information about the channel change operation to the TV. That is, based on the device determination model 1330, it is determined that the operation execution device is the second device 4200 as a "TV". Since the function determination model corresponding to the TV is stored in the hub device 1000, the function determination device determination module 1340 can provide (e.g., send) at least a part of the text to the TV function determination model 1354, so that the hub device 1000 itself can process at least a part of the text.
[0147] For example, based on the voice input that the user says "Execute the deodorization mode" received by the hub device 1000, the hub device 1000 can determine, through the device determination model 1330, that the device for executing the deodorization mode is an air purifier. Next, according to an embodiment, the function determination device determination module 1340 can check whether the function determination model corresponding to the air purifier is stored in the voice assistant server 2000, and based on determining that the function determination model corresponding to the air purifier is stored in the voice assistant server 200, can send the text corresponding to "Execute the deodorization mode" to the voice assistant server 2000. According to an embodiment of the present disclosure, the voice assistant server 2000 can analyze the received text through the function determination model corresponding to the air purifier, and can determine the deodorization mode execution operation as an operation to be performed by the air purifier by using the text analysis result. According to an embodiment of the present disclosure, the voice assistant server 2000 can send the operation information regarding the deodorization mode execution operation to the air purifier, and in this case, the operation information related to the deodorization mode execution operation can be sent through the IoT server 3000. That is, according to an embodiment, based on the operation execution device determined by the device determination model 1330 being the third device 4300 as an "air purifier", since the function determination model corresponding to the air purifier is stored in the voice assistant server 2000, the function determination device determination module 1340 can send at least a part of the text to the voice assistant server 2000 (S1).
[0148] In an embodiment of the present disclosure, the hub device 1000 itself can store the function determination model corresponding to at least one of the plurality of devices 4000. For example, when the hub device 1000 is a voice assistant speaker, the hub device 1000 can store the speaker function determination model 1352, and the speaker function determination model 1352 is used to obtain the operation information regarding the detailed operations for executing the functions of the voice assistant speaker and the relationships between the detailed operations.
[0149] According to an embodiment, the hub device 1000 can also store the function determination model corresponding to another device. For example, the hub device 1000 can store the TV function determination model 1354 for obtaining the operation information regarding the detailed operations corresponding to the TV and the relationships between the detailed operations. According to an embodiment of the present disclosure, the TV can be a device previously registered to the IoT server 3000 by using the same user account as the user account of the hub device 1000.
[0150] According to an embodiment, the speaker function determination model 1352 and the TV function determination model 1354 can respectively include a second NLU model 1352a and 1354a and an action plan management module 1532b and 1354b. Refer to Figure 2Describe the second NLU models 1352a and 1354a and the action plan management modules 1352b and 1354b.
[0151] According to an embodiment, the voice assistant server 2000 may determine an operation execution device for performing an operation for a user intention based on text received from the hub device 1000. According to an embodiment, the voice assistant server 2000 may receive user account information from the hub device 1000 (S1). According to an embodiment of the present disclosure, based on the user account information received by the voice assistant server 2000 from the hub device 1000, the voice assistant server 2000 may send a query for requesting device information about a plurality of devices 4000 previously registered according to the received user account information to the IoT server 3000 (S9), and may receive device information about the plurality of devices 4000 from the IoT server 3000 (S10). According to an embodiment, the voice device information may include at least one of identification information (e.g., device id information) of each of the plurality of devices 4000, device types of each of the plurality of devices 4000, function execution capabilities of each of the plurality of devices 4000, location information, and status information. According to an embodiment, the voice assistant server 2000 may send the device information received from the IoT server 3000 to the hub device 1000 (S2).
[0152] Reference will be made to Figure 2 Describe the components of the hub device 1000 according to an embodiment.
[0153] According to an embodiment, the voice assistant server 2000 may include a device determination model 2330 and a plurality of function determination models 2342, 2344, 2346, and 2348. According to an embodiment, the voice assistant server 2000 may select, by using the device determination model 2330, a function determination model corresponding to at least a part of the text received from the hub device 1000 from among the plurality of function determination models 2342, 2344, 2346, and 2348, and may obtain operation information required for the operation execution device to perform an operation by using the selected function determination model. According to an embodiment, the voice assistant server 2000 may send the operation information to the IoT server 3000 (S9).
[0154] Reference is made to Figure 3 Describe the components of the voice assistant server 2000 according to an embodiment of the present disclosure.
[0155] According to an embodiment, the IoT server 3000 may be connected via a network and may store information about a plurality of devices 4000 previously registered with a user account using the hub device 1000. In an embodiment of the present disclosure, the IoT server 3000 may receive at least one of user account information used by each of the plurality of devices 4000 to log in, identification information (e.g., device ID information) of each of the plurality of devices 4000, the device type of each of the plurality of devices 4000, and information on the function execution capabilities of each of the plurality of devices 4000 (S11, S13, and S15). In an embodiment of the present disclosure, the IoT server 3000 may receive status information on power-on / off or operations being performed of each of the plurality of devices 4000 from the plurality of devices 4000 (S11, S13, and S15). The IoT server 3000 may store the device information and status information received from the plurality of devices 4000.
[0156] According to an embodiment, the IoT server 3000 may generate a control command readable and executable by an operation execution device based on operation information received from the voice assistant server 2000. According to an embodiment of the present disclosure, the IoT server 3000 may send the control command to a device determined to be the operation execution device among the plurality of devices 4000 (S12, S14, and S16).
[0157] Reference Figure 4 Elements of the IoT server 3000 according to an embodiment are described.
[0158] In Figure 1 the embodiment, the plurality of devices 4000 may include a first device 4100, a second device 4200, and a third device 4300. Although in Figure 1 the first device 4100 may be an air conditioner, the second device 4200 may be a TV, and the third device 4300 may be an air purifier, the present disclosure is not limited thereto. The plurality of devices 4000 may include not only air conditioners, TVs, and air purifiers, but also other IoT devices, such as home appliances such as robotic cleaners, washing machines, ovens, microwave ovens, scales (e.g., weighing scales), refrigerators, or digital photo frames, and mobile devices such as smart phones, tablet personal computers (PCs), mobile phones, video phones, e-book readers, desktop PCs, laptop PCs, netbook computers, workstations, servers, personal digital assistants (PDAs), portable multimedia players (PMPs), MP3 players, mobile medical devices, cameras, or wearable devices.
[0159] According to an embodiment, at least one of the plurality of devices 4000 may store a function determination model by itself. For example, a function determination model 4132 for obtaining operation information on detailed operations required for a first device 4100 to perform an operation determined from a user's voice input and relationships between the detailed operations, and generating a control command based on the operation information may be stored in a memory 4130 of the first device 4100.
[0160] According to an embodiment, each of a second device 4200 and a third device 4300 among the plurality of devices 4000 may not store a function determination model.
[0161] According to an embodiment, at least one of the plurality of devices 4000 may send information on whether the device itself stores a function determination model to the hub device 1000 (S4, S6, and S8).
[0162] According to an embodiment, the plurality of devices 4000 may further include third - party devices (e.g., the third - party devices are manufactured by different manufacturers from the manufacturer of the hub device 1000), a voice assistant server 2000, and an IoT server 3000 that are not manufactured by the same manufacturer as the manufacturer of the hub device 1000, and are not directly controlled by the hub device 1000, the voice assistant server 2000, and the IoT server 3000. The third - party devices will be described in detail with reference to Figure 11 Describe the third - party devices in detail.
[0163] Figure 2 is a block diagram illustrating elements of a hub device 1000 according to an embodiment of the present disclosure.
[0164] According to an embodiment, the hub device 1000 may be a device that receives a user's voice input and controls at least one of the plurality of devices 4000 based on the received voice input. According to an embodiment, the hub device 1000 may be a listening device that receives a voice input from a user.
[0165] Refer to Figure 2 , according to an embodiment, the hub device 1000 may include a microphone 1100, a processor 1200, a memory 1300, and a communication interface 1400. According to an embodiment, the hub device 1000 may receive a voice input (e.g., a user's words) from a user through the microphone 1100 and may obtain a voice signal from the received voice input. In an embodiment of the present disclosure, the processor 1200 of the hub device 1000 may convert the sound received through the microphone 1100 into an acoustic signal and may obtain a voice signal by removing noise (e.g., non - voice components) from the acoustic signal.
[0166] However, the present disclosure is not limited thereto, and the hub device 1000 may receive a voice signal from a listening device.
[0167] According to an embodiment, the hub device 1000 may include a voice recognition module having a function of detecting a specified voice input (e.g., a wake-up input such as "Hi, Bixby" or "OK, Google") or a function of preprocessing a voice signal obtained from a part of the voice input.
[0168] According to an embodiment of the present disclosure, the processor 1200 may execute one or more instructions of a program stored in the memory 1300. According to an embodiment of the present disclosure, the processor 1200 may include hardware components that perform arithmetic, logical, and input / output operations, as well as signal processing. According to an embodiment, the processor 1200 may include, but is not limited to, at least one of a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), and a field programmable gate array (FPGA).
[0169] According to an embodiment, the program may include one or more instructions for controlling the plurality of devices 4000 based on a user's voice input received through the microphone 1100, and the program may be stored in the memory 1300. According to an embodiment, instructions and / or program code readable by the processor 1200 may be stored in the memory 1300. In an exemplary embodiment of the present disclosure, the processor 1200 may be implemented by executing instructions or code stored in the memory.
[0170] According to an embodiment, one or more or all of the data regarding the automatic speech recognition (ASR) module 1310, the data regarding the natural language generator (NLG) module 1320, the data regarding the device determination model 1330, the data regarding the function determination device determination module 1340, the data corresponding to each of the plurality of function determination models 1350, and the data corresponding to the database 1360 may be stored in the memory 1300.
[0171] According to an embodiment, the memory 1300 may include at least one type of storage medium such as a flash type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD (Secure Digital) memory or XD (eXtreme Digital) memory), a random access memory (RAM), a static random access memory (SRAM), a read only memory (ROM), an electrically erasable programmable read only memory (EEPROM), a programmable read only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk.
[0172] According to an embodiment, the processor 1200 may perform ASR by using data about the ASR module 1310 stored in the memory 1300, and may convert a voice signal received through the microphone 1100 into text.
[0173] However, the present disclosure is not limited thereto. In an embodiment of the present disclosure, the processor 1200 may receive a voice signal from a listening device through the communication interface 1400, and may convert the voice signal into text by performing ASR by using the ASR module 1310.
[0174] According to an embodiment, the processor 1200 may analyze text by using data about the device determination model 1330 stored in the memory 1300, and may determine an operation execution device (e.g., an IoT device) from among a plurality of devices 4000 (e.g., a plurality of IoT devices) based on the analysis result of the text. According to an embodiment of the present disclosure, the device determination model 1330 may include a first NLU model 1332. In an embodiment of the present disclosure, the processor 1200 may analyze text by using data about the first NLU model 1332 included in the device determination model 1330, and may determine an operation execution device for performing an operation according to a user's intention from among a plurality of devices 4000 based on the analysis result of the text.
[0175] According to an embodiment, the first NLU model 1332 may be a model trained to analyze text converted from voice input and determine an operation execution device based on the analysis result. According to an embodiment of the present disclosure, the first NLU model 1332 may be used to determine an intention by interpreting text and determine an operation execution device based on the intention.
[0176] In an embodiment of the present disclosure, the processor 1200 may parse text in units of morphemes, words, or phrases by using data about the first NLU model 1332 stored in the memory 1300, and may infer the meaning of a word extracted from the parsed text by using language features (e.g., grammatical components) of the morphemes, words, or phrases. According to an embodiment, the processor 1200 may compare the inferred meaning of the word with a predefined intention provided by the first NLU model 1332, and may determine an intention corresponding to the inferred meaning of the word.
[0177] According to an embodiment, the processor 1200 may determine, as an operation execution device, a device related to an intention recognized from text based on a matching model for determining a relationship between an intention and a device. In an embodiment of the present disclosure, the matching model may be included in data about the device determination model 1330 stored in the memory 1300, and may be obtained through learning via a rule-based system, but the present disclosure is not limited thereto.
[0178] In an embodiment of the present disclosure, the processor 1200 may obtain a plurality of numerical values indicating the degree of relationship between an intention and a plurality of devices 4000 by applying a matching model to the intention, and may determine, as a final operation execution device, a device having the maximum value among the obtained plurality of numerical values (e.g., an IoT device capable of performing an operation intended by a user's voice (voice input)). For example, based on the intention being related to each of a first device 4100 (see Figure 1 ) and a second device 4200 (see Figure 1 ), the processor 1200 may obtain a first numerical value indicating the degree of relationship between the intention and the first device 4100 and a second numerical value indicating the degree of relationship between the intention and the second device 4200, and may determine the first device 4100 having the larger value among the first numerical value and the second numerical value as the final operation execution device.
[0179] For example, based on the hub device 1000 receiving a voice input from a user saying "Lower the set temperature by 2°C because it's hot", the processor 1200 may perform ASR to convert the voice input into text, and may analyze the text by using data related to the first NLU model 1332 to obtain (e.g., by inference) an intention corresponding to "set temperature adjustment". According to an embodiment, the processor 1200 may obtain a first numerical value indicating the degree of relationship between the intention of "set temperature adjustment" and the first device 4100 which is an air conditioner, a second numerical value indicating the degree of relationship between the intention of "set temperature adjustment" and the second device 4200 which is a TV, and a third numerical value indicating the degree of relationship between the intention of "set temperature adjustment" and a third device 4300 which is an air purifier (see Figure 1 ) by applying the matching model. According to an embodiment, the processor 1200 may determine the first device 4100 as an operation execution device related to "set temperature adjustment" by using the first numerical value which is the maximum value among the obtained numerical values.
[0180] According to an embodiment, based on the hub device 1000 receiving a voice input from a user saying "Play the movie Avengers", the processor 1200 may analyze the text converted from the voice input, and may obtain an intention corresponding to "content playback". According to an embodiment, the processor 1200 may determine the second device 4200 as an operation execution device related to "content playback" based on the second numerical value information being the maximum value among a first numerical value indicating the degree of relationship between the intention of "content playback" and the first device 4100 which is an air conditioner, a second numerical value indicating the degree of relationship between the intention of "content playback" and the second device 4200 which is a TV, and a third numerical value indicating the degree of relationship between the intention of "content playback" and the third device 4300 which is an air purifier, calculated by using the matching model.
[0181] However, the present disclosure is not limited to the above examples, and the processor 1200 may arrange in ascending order numerical values indicating the degree of relationship between the intention and multiple devices, and may determine a predetermined number of devices as operation execution devices. In an embodiment of the present disclosure, the processor 1200 may determine a device whose numerical value indicating the degree of relationship is equal to or greater than a predetermined threshold value as an operation execution device related to the intention. In this case, multiple devices may be determined as operation execution devices.
[0182] According to an embodiment, the processor 1200 may train a matching model between the intention and the operation execution device by using, for example, a rule-based system, but the present disclosure is not limited thereto. The AI model used by the processor 1200 may be, for example, a neural network-based system (e.g., a convolutional neural network (CNN) or a recurrent neural network (RNN)), a support vector machine (SVM), linear regression, logistic regression, naive Bayes, random forest, decision tree, or k-nearest neighbor algorithm. Alternatively, the AI model may be a combination of the above examples or any other AI model. According to an embodiment, the AI model used by the processor 1200 may be stored in the device determination model 1330.
[0183] According to an embodiment, the device determination model 1330 stored in the memory 1300 of the hub device 1000 may determine an operation execution device from among multiple devices 4000 registered according to the user account of the hub device 1000. According to an embodiment, the hub device 1000 may receive device information about each of the multiple devices 4000 from the voice assistant server 2000 by using the communication interface 1400. According to an embodiment, the device information may include, for example, at least one of identification information (e.g., device id information) of each of the multiple devices 4000, the device type of each of the multiple devices 4000, the function execution ability of each of the multiple devices 4000, location information, and status information. According to an embodiment, the processor 1200 may determine, based on the device information, a device for performing an operation according to the intention from among the multiple devices 4000 by using data about the device determination model 1330 stored in the memory 1300.
[0184] In an embodiment of the present disclosure, the processor 1200 may analyze numerical values indicating the degree of relationship between the intention and multiple devices 4000 by using the device determination model 1330, where the multiple devices 4000 were previously registered using the same user account as the user account of the hub device 1000, and may determine as the operation execution device the device having the maximum value among the numerical values indicating the degree of relationship between the intention and the multiple devices 4000.
[0185] Since the device determination model 1330 is configured to determine an operation execution device by using only the plurality of devices 4000 as device candidates, and the plurality of devices 4000 are logged in and registered by using the same user account as the user account of the hub device 1000, the following technical effects may exist: the amount of computation executed by the processor 1200 to determine the degree of relationship with the intention can be reduced to less than the amount of computation of the processor 2200 of the voice assistant server 2000. In addition, due to the reduction in the amount of computation, the processing time required to determine the operation execution device can be reduced, and thus the response speed can be improved.
[0186] In an embodiment of the present disclosure, the processor 1200 may obtain the name of a device from text by using the first NLU model 1332, and may determine an operation execution device based on the name of the device by using data about the device determination model 1330 stored in the memory 1300. In an embodiment of the present disclosure, the processor 1200 may extract a general name related to a device and words or phrases about the installation location of the device from text by using the first NLU model 1332, and may determine an operation execution device based on the extracted general name and the installation location of the device. For example, when the text converted from a voice input is "Play the movie Avengers on the TV", the processor 1200 may parse the text in units of words or phrases by using the first NLU model 1332, and may identify the name of the device corresponding to "TV" by comparing the words or phrases with pre-stored words or phrases. According to an embodiment, the processor 1200 may determine the second device 4200, which is a TV among the plurality of devices 4000 that are logged in by using the same account as the user account of the hub device 1000 and are connected to the hub device 1000, as the operation execution device.
[0187] In an embodiment of the present disclosure, any one of the plurality of devices 4000 may be a listening device that receives a voice input from a user, and the processor 1200 may determine the listening device as the operation execution device by using the device determination model 1330. However, the present disclosure is not limited thereto, and the processor 1200 may determine a device other than the listening device as the operation execution device from among the plurality of devices 4000 by using the device determination model 1330.
[0188] According to an embodiment, the processor 1200 may receive updated data of the device determination model 1330 from the voice assistant server 2000 by using the communication interface 1400. In an embodiment of the present disclosure, the latest functions of the update may be included in at least one of the plurality of devices 4000, or a new device with new functions may be added to the user account. In this case, the device determination model 1330 of the hub device 1000 may not be able to determine the new functions of the plurality of devices 4000, or may not be able to determine the newly added device as an operation execution device. According to an embodiment, the processor 1200 may update the device determination model 1330 to the latest version by using the received updated data. According to an embodiment, the latest version may be the same as the version of the device determination model 2330 of the voice assistant server 2000 (see Figure 3 ).
[0189] According to an embodiment, the updated data of the device determination model 1330 may include updated data such that the device determination model 1330 can determine detailed operation information of the latest functions of each of the plurality of devices 4000 connected through the user account, and can determine the operation execution device that executes the latest functions from the text. In an embodiment of the present disclosure, according to an embodiment, the updated data of the device determination model 1330 may include information about the functions of a new device newly added to the user account.
[0190] Reference will be made to Figure 17 and 18 to describe embodiments of updating the device determination model 1330.
[0191] According to an embodiment, the processor 1200 may receive device information of a new device from the voice assistant server 2000 by using the communication interface 1400. According to an embodiment, the device information obtained by the processor 1200 of the hub device 1000 from the voice assistant server 2000 may include at least one of, for example, device identification information of the new device (e.g., device ID information), storage information of the device determination model of the new device, and storage information of the function determination model of the new device. According to an embodiment, the new device may be a target device expected to be registered in the user account of the voice assistant server 2000.
[0192] According to an embodiment, the processor 1200 may add the new device to the device candidates that can be determined as operation execution devices by the device determination model 1330 by using the device information of the new device obtained from the voice assistant server 2000. In an embodiment of the present disclosure, since the device determination model 1330 prompts the new device to be included in the device candidates, the processor 1200 may enhance the device determination ability to determine the operation execution device from the text.
[0193] Reference Figure 19Describe an embodiment of registering a new device.
[0194] According to an embodiment, the NLG model 1320 can be used to provide a response message during an interaction between the hub device 1000 and the user. For example, the processor 1200 can generate a response message such as "I will play a movie on the TV" or "I will lower the set temperature of the air conditioner by 2°C" by using the NLG model 1320.
[0195] According to an embodiment, when there are multiple operation execution devices determined by the processor 1200 or there are multiple devices with a similar degree of relationship to the intent, the NLG model 1320 can store data for generating a query message for determining a specific operation execution device. In an embodiment of the present disclosure, the processor 1200 can generate a query message for selecting one operation execution device from among multiple device candidates by using the NLG model 1320. According to an embodiment, the query message can be a message for requesting a response from the user regarding which of the multiple device candidates will be identified (determined) as the operation execution device.
[0196] According to an embodiment, the function determination device determination module 1340 can be a module for identifying a device in which a function determination model corresponding to the operation execution device is stored and for determining a target device to which at least part of the text or the entire text is to be sent. According to an embodiment, the function determination device determination module 1340 can identify, by using the database 1360 including information on the function determination model of the device, a device in which a function determination model corresponding to the operation execution device is stored from among the hub device 1000, the voice assistant server 2000, and the operation execution device. In an embodiment of the present disclosure, when the hub device 1000 receives a voice signal from the listening device, the function determination device determination module 1340 can identify, by using the database 1035, a device in which a function determination model corresponding to the operation execution device is stored from among the internal memory of the hub device 1000, the internal memory of the operation execution device, the internal memory of the listening device, and the voice assistant server 2000.
[0197] According to an embodiment, information on whether it is possible to store information on the function determination model for each of the plurality of devices 4100 previously registered using a user account, and information on the storage location of the function determination model for each of the plurality of devices 4000 (e.g., device identification information, IP address, or MAC address) may be stored in the database 1360 in the form of a lookup table (LUT). In an embodiment of the present disclosure, the function determination device determination module 1340 may search the database 1360 for the lookup table based on the device identification information of the operation execution device output by the device determination model 1330, and may obtain information on the storage location of the function determination model corresponding to the operation execution device based on the search result of the lookup table.
[0198] In an embodiment of the present disclosure, the function determination device determination module 1340 itself may store data for determining the target device to which the entire text or at least a part of the entire text is to be sent. According to an embodiment, the processor 1200 may identify the device storing the function determination model corresponding to the operation execution device by using the data on the function determination device determination module 1340.
[0199] According to an embodiment, the function determination model corresponding to the operation execution device may be stored in the memory 1300 of the hub device 1000, may be stored in the memory of the operation execution device itself, or may be stored in the memory 2300 of the voice assistant server 2000.
[0200] According to an embodiment, the term "function determination model corresponding to the operation execution device" refers to a model for obtaining operation information on detailed operations for performing operations according to the determined function of the operation execution device and the relationships between the detailed operations. In an embodiment of the present disclosure, the speaker function determination model 1352 and the TV function determination model 1354 stored in the memory 1300 of the hub device 1000 may correspond to the plurality of devices that are logged in using the same account as the user account and connected to the hub device 1000 via a network, respectively.
[0201] For example, the speaker function determination model 1352, which is a first function determination model, may be a model for obtaining operation information on detailed operations for performing operations according to the function of the first device 4100 (see Figure 1 ) and the relationships between the detailed operations. In an embodiment of the present disclosure, the speaker function determination model 1352 may be, but is not limited to, a model for obtaining operation information according to the function of the hub device 1000. Similarly, the TV function determination model 1354, which is a second function determination model, may be a model for obtaining operation information on detailed operations for performing operations according to the function of the second device 4200 (see Figure 1 ) and the relationships between the detailed operations.
[0202] According to an embodiment, the speaker function determination model 1352 and the TV function determination model 1354 may respectively include a second NLU model 1352a and 1354a, configured to analyze at least a part of the text and obtain operation information about an operation to be performed by an operation execution device determined based on the analysis result of at least a part of the text. According to an embodiment, the speaker function determination model 1352 and the TV function determination model 1354 may respectively include action plan management modules 1352b and 1354b, configured to manage operation information related to detailed operations of the device so as to generate detailed operations to be performed by the device and an execution order of the detailed operations. According to an embodiment, the action plan management modules 1352b and 1354b may manage operation information about detailed operations of the device according to the relationship between the device and the detailed operations. According to an embodiment, the action plan management modules 1352b and 1354b may plan detailed operations to be performed by the device and an execution order of the detailed operations based on the analysis result of at least a part of the text.
[0203] According to an embodiment, multiple function determination models of multiple devices, for example, the speaker function determination model 1352 and the TV function determination model 1354, may be stored in the memory 1300 of the hub device 1000.
[0204] As described above, according to an embodiment, the processor 1200 may check whether a function determination model corresponding to the operation execution device is stored in the memory 1300 by using data about the function determination device determination module 1340 to search a look-up table stored in the database 1360. For example, based on determining that the operation execution device is the first device 4100, the function determination model corresponding to the first device 4100 may not be stored in the memory 1300 of the hub device 1000. In this case, according to an embodiment, the processor 1200 may check that the function determination model corresponding to the operation execution device is not stored in the hub device 1000. As another example, based on determining that the operation execution device is the second device 4200, the TV function determination model 1354 corresponding to the second device 4200 may be stored in the memory 1300 of the hub device 1000. In this case, according to an embodiment, the processor 1200 may check that the function determination model corresponding to the operation execution device is stored in the hub device 1000.
[0205] According to an embodiment, when the processor 1200 checks that a function determination model corresponding to an operation execution device is stored in the hub device 1000, the processor 1200 may provide at least a part of the text to the function determination model corresponding to the operation execution device stored in the hub device 1000 by using data on the function determination device determination module 1340. For example, when the operation execution device is the second device 4200, the TV function determination model 1354 corresponding to the TV as the second device 4200 may be stored in the memory 1300 of the hub device 1000, so the processor 1200 may provide at least a part of the text to the TV function determination model 1354 by using the function determination device determination module 1340.
[0206] In an embodiment of the present disclosure, the processor 1200 may send only a part (instead of the whole text) of the whole text to the TV function determination model 1354. For example, when the text converted from voice input is "Play the movie Avengers on the TV", "on the TV" specifies the name of the operation execution device, so it may be unnecessary information for the TV function determination model 1354. According to an embodiment, the processor 1200 may parse the text in units of words or phrases by using the first NLU model 1332, may identify words or phrases specifying the name, common name, or installation location of the device, and may provide the remaining part of the text except for the words or phrases identified in the whole text to the TV function determination model 1354.
[0207] According to an embodiment, the processor 1200 may obtain operation information related to an operation to be performed by the operation execution device by using the function determination model corresponding to the operation execution device stored in the memory 1300, for example, the second NLU model 1354a of the TV function determination model 1354. According to an embodiment, the second NLU model 1354a, which is a model dedicated to a specific device (e.g., TV), may be an AI model trained to obtain an intention related to a device corresponding to the operation execution device determined by the first NLU model 1332 and corresponding to the text. In addition, the second NLU model 1354a may be a model trained to determine an operation of a device related to a user's intention by interpreting the text. According to an embodiment, an operation may refer to at least one action performed by a device when the device executes a specific function. An operation may refer to at least one action performed by a device when the device executes an application.
[0208] In an embodiment of the present disclosure, the processor 1200 may analyze text by using a second NLU model 1354a of the TV function determination model 1354 corresponding to a determined operation execution device (e.g., TV). According to an embodiment, the processor 1200 may parse the text in units of morphemes, words, or phrases by using the second NLU model 1354a, may identify the meanings of the morphemes, words, or phrases parsed through syntactic or semantic analysis, and may determine an intention and parameters by matching the identified meanings with predefined words. According to an embodiment, the term "parameters" used herein refers to variable information for determining detailed operations of an operation execution device related to an intention. For example, when the text sent to the TV function determination model 1354 is "Play the movie Avengers on the TV", the intention may be "content playback", and the parameter may be "movie Avengers" as information about the content to be played.
[0209] According to an embodiment, the processor 1200 may obtain operation information about at least one detailed operation related to an intention and parameters by using an action plan management module 1354b of the TV function determination model 1354. According to an embodiment, the action plan management module 1354b may manage information about detailed operations of a device according to the relationship between the device and the detailed operations. According to an embodiment, the processor 1200 may plan detailed operations to be executed by an operation execution device (e.g., TV) and the execution order of the detailed operations based on the intention and parameters by using the action plan management module 1354b, and may obtain operation information. According to an embodiment, the operation information may be information related to the detailed operations to be executed by the device and the execution order of the detailed operations. According to an embodiment, the operation information may include information related to the detailed operations to be executed by the device, the relationship between each detailed operation and another detailed operation, and the execution order of the detailed operations. The operation information may include, but is not limited to, functions executed by the operation execution device to execute a specific operation, the execution order of the functions, input values required to execute the functions, and output values output as execution results of the functions.
[0210] According to an embodiment, the processor 1200 may generate a control command for controlling the operation execution device based on the operation information. According to an embodiment, the control command may refer to instructions readable and executable by the operation execution device, such that the operation execution device executes the detailed operations included in the operation information. According to an embodiment, the processor 1200 may control the communication interface 1400 to send the generated control command to the operation execution device.
[0211] In an embodiment of the present disclosure, the processor 1200 may obtain information about a function determination model from a plurality of devices 4000 that are logged in using the same account as a user account of the hub device 1000 and connected to the hub device 1000 through a network through the communication interface 1400. According to an embodiment, the information about the function determination model may include information about whether each of the plurality of devices 4000 stores a function determination model. Figure 1 , despite Figure 1 In an embodiment of the present invention, the first device 4100 among the plurality of devices 4000 may store the function determination model 4132 in the memory 4130, but the second device 4200 and the third device 4300 may not have the function determination model (or may not have the function determination model 4132). According to an embodiment, when information about whether at least one of the first device 4100, the second device 4200, or the third device 4300 stores the function determination model is sent to the hub device 1000 (S4, S6, and S8), the hub device 1000 may provide the information about the function determination model obtained through the communication interface 1400 to the processor 1200.
[0212] In an embodiment of the present disclosure, the processor 1200 may receive the voice assistant server 2000 (see Figure 1 ) obtain information about multiple devices 4000 (see Figure 1 ) of the function determination model. According to an embodiment, each of the plurality of devices 4000 may be registered according to a user account logged in when a user inputs a user id and a password, and the user account information and device information of each of the plurality of devices 4000 may be sent to the IoT server 3000. In this case, information on whether each of the plurality of devices 4000 itself stores a device determination model and stores a function determination model may also be sent to the IoT server 3000. According to an embodiment, the IoT server 3000 may send device information about the plurality of devices 4000 registered in the user account information, information about the device determination model, and information about the function determination model to the voice assistant server 2000. According to an embodiment, the voice assistant server 2000 may send device information about the plurality of devices 4000, information about the device determination model, and information about the function determination model to the hub device 1000 having the same user account information. According to an embodiment, the hub device 1000 may store information about the function determination model of each of the plurality of devices 4000 in the form of a lookup table in the database 1360.
[0213] However, the present disclosure is not limited thereto, and the hub device 1000 may receive information about the function determination model from at least one of the plurality of devices 4000. In an embodiment of the present disclosure, the processor 1200 may control the communication interface 1400 to receive information about the function determination model from at least one device storing the function determination model for determining the function of each of the plurality of devices 4000. According to an embodiment, based on the received information about the function determination model of the plurality of devices 4000, the processor 1200 may control the communication interface 1400 to further receive information about the storage location of the function determination model of each of the plurality of devices 4000 (e.g., device identification information, IP address, or MAC address).
[0214] According to an embodiment, the processor 1200 may store the received information about the function determination model and the received information about the storage location of the function determination model of each device in the database 1360. In an embodiment of the present disclosure, the processor 1200 may store the information about the storage location of the function determination model and the information about whether the function determination model is stored in the form of a lookup table according to the name or identification information of the device.
[0215] According to an embodiment, the processor 1200 may obtain the information about the function determination model of each device by searching the lookup table stored in the database 1360, and may identify the device storing the function determination model corresponding to the operation execution device based on the obtained information about the function determination model. For example, according to an embodiment, when the operation execution device determined based on the text is the first device 4100, the processor 1200 may determine the first device 4100 itself as the device storing the function determination model 4132 (see Figure 1 ). In this case, the processor 1200 may determine the first device 4100 as the target device to which at least a part of the text is to be sent by using the data about the function determination device determination module 1340. According to an embodiment, the processor 1200 may control the communication interface 1400 to send at least a part of the text to the function determination model 4132 of the first device 4100.
[0216] In an embodiment of the present disclosure, the processor 1200 may control the communication interface 1400 to separate the part regarding the name of the operation execution device from the text and send only the remaining part of the text to the first device 4100. For example, when the first device 4100 is an air conditioner and the text is "lower the set temperature by 2°C in the air conditioner", when the text is sent to the air conditioner, it is not necessary to send "in the air conditioner". In this case, the processor 1200 may parse the text in units of words or phrases, may identify the words or phrases specifying the name, general name, installation location, etc. of the first device 4100, and may provide the remaining part of the text except for the words or phrases identified in the entire text to the first device 4100.
[0217] According to an embodiment, when the processor 1200 is to identify the device storing the function determination model corresponding to the operation execution device based on the obtained information regarding the function determination model, the processor 1200 may check whether the function determination model is not stored in the operation execution device. For example, based on determining that the operation execution device is the third device 4300, according to an embodiment, the function determination model is not stored in the third device 4300. According to an embodiment, the function determination model corresponding to the third device 4300 is also not stored in the hub device 1000. In this case, the processor 1200 may check through the data regarding the function determination device determination module 1340 that the function determination model corresponding to the operation execution device is not stored in any of the hub device 1000 and the multiple devices 4000. According to an embodiment, the processor 1200 may determine through the data regarding the function determination device determination module 1340 that the target device to which at least a part of the text is to be sent is the voice assistant server 2000. According to an embodiment, the processor 1200 may control the communication interface 1400 to send at least a part of the text to the voice assistant server 2000.
[0218] In an embodiment of the present disclosure, the function determination device determination module 1340 may be configured to determine whether the operation execution device determined by the device determination model 1330 is the same as the listening device. In an embodiment of the present disclosure, when the processor 1200 receives a voice signal from the listening device through the communication interface 1400, the processor 1200 may also receive the device identification information of the listening device (e.g., the device id information of the listening device). According to an embodiment, the function determination device determination module 1340 may check the function information of the listening device from the device identification information of the listening device, and may determine whether the listening device and the operation execution device are the same by comparing the function information of the listening device with the function information of the operation execution device.
[0219] According to an embodiment, based on determining that the operation execution device is the same as the listening device, the processor 1200 may determine whether the listening device stores the function determination model in the internal memory by using the data on the function determination device determination module 1340. According to an embodiment, when the function determination model stored in the internal memory of the listening device is the same as the function determination model corresponding to the operation execution device, the processor 1200 may send at least a part of the text to the function determination model previously stored in the internal memory of the listening device by using the communication interface 1400. In an embodiment of the present disclosure, the processor 1200 may control the communication interface 1400 to separate the part regarding the name of the operation execution device from the text and send only the remaining part of the text to the listening device. For example, when the listening device is an air conditioner and the text is "lower the set temperature to 20 °C in the air conditioner", when the text is sent to the air conditioner, it is not necessary to send "in the air conditioner". In this case, the processor 1200 may parse the text in units of words or phrases, may identify the words or phrases specifying the name, general name, installation location, etc. of the listening device, and may provide the remaining part of the text except for the words or phrases identified in the entire text to the listening device.
[0220] According to an embodiment, when the function determination model stored in the internal memory of the listening device is not the same as the function determination model corresponding to the operation execution device, the processor 1200 may send at least a part of the text to the voice assistant server 2000 by using the communication interface 1400.
[0221] According to an embodiment, based on determining that the operation execution device is a separate device different from the listening device, the processor 1200 may determine whether the operation execution device stores the function determination model in the internal memory by using the data on the function determination device determination module 1340. According to an embodiment, when the function determination model is stored in the internal memory of the operation execution device, the processor 1200 may send at least a part of the text to the function determination model previously stored in the internal memory of the operation execution device by using the communication interface 1400. In an embodiment of the present disclosure, the processor 1200 may control the communication interface 1400 to separate the part regarding the name of the operation execution device from the text and send only the remaining part of the text to the operation execution device.
[0222] According to an embodiment, when the function determination model is not stored in the internal memory of the operation execution device, the processor 1200 may send at least a part of the text to the voice assistant server 2000 by using the communication interface 1400.
[0223] According to an embodiment, it may be determined whether the operation execution device and the listening device are the same, and at least a part of the text may be sent to be referred to Figure 14Any one of the described operation execution devices, listening devices, and voice assistant servers.
[0224] According to an embodiment, the communication interface 1400 may perform data communication with the voice assistant server 2000, the IoT server 3000, and the plurality of devices 4000. According to an embodiment, the communication interface 1400 may perform data communication with one or more of the voice assistant server 2000, the IoT server 3000, and the plurality of devices 4000 by using at least one of data communication methods including wireless LAN, Wi-Fi, Bluetooth, Zigbee, WFD, IrDA, BLE, NFC, wireless broadband Internet (Wibro), worldwide interoperability for microwave access (WiMAX), shared wireless access protocol (SWAP), wireless gigabit alliance (WiGig), and radio frequency (RF) communication.
[0225] According to an embodiment, the database 1360 may store identification information of each of the plurality of devices 4000, information on whether a function determination model exists, and information on the storage location of the function determination model of each of the plurality of devices 4000 (e.g., identification information of the stored device, the IP address of the stored device, or the MAC address of the stored device).
[0226] Although the database 1360 may be stored in Figure 2 the memory 1300, the database 1360 may also be stored in a memory separate from the memory 1300.
[0227] In (e.g., Figure 1 and Figure 2In an embodiment, the hub device 1000 may include a device determination model 1330 configured to determine an operation execution device by using only a plurality of devices 4000 as device candidates. The plurality of devices 4000 are logged in by using the same user account information as the user account information of the hub device 1000 and registered according to the user account information. And the function determination device determination module 1340 may be configured to identify a device storing a function determination model corresponding to the operation execution device determined by the device determination model 1330 and may determine at least a part of the text to be sent to the identified device. Since the hub device 1000 includes some of the models included in the voice assistant server 2000, determines the operation execution device by using the models, and controls the operation of the operation execution device, it is not necessary to operate the voice assistant server 2000 through the network in all processes. Therefore, the network usage fee can be reduced and the server operation efficiency can be improved. In addition, since the operation execution device is determined by using only a plurality of devices registered by using the user account of the hub device 1000 as device candidates, the calculation amount and processing time can be reduced, and the response speed can be improved.
[0228] Figure 3 FIG. is a block diagram illustrating elements of a voice assistant server 2000 according to an embodiment of the present disclosure.
[0229] According to an embodiment, the voice assistant server 2000 may be a server that receives text converted from a user's voice input from the hub device 1000, determines an operation execution device based on the received text, and obtains operation information by using a function determination model corresponding to the operation execution device.
[0230] Reference Figure 3 Referring to, according to an embodiment, the voice assistant server 2000 may at least include a communication interface 2100, a processor 2200, and a memory 2300.
[0231] According to an embodiment, the communication interface 2100 of the voice assistant server 2000 may receive device information including at least one of the following items from the IoT server 3000 (see Figure 1 ) by performing data communication with the IoT server 3000: a plurality of devices 4000 (see Figure 1identification information (e.g., device ID information) for each of them, the device type of each of the multiple devices 4000, the function execution capabilities of each of the multiple devices 4000, location information, and status information. In an embodiment of the present disclosure, the voice assistant server 2000 may receive information about the function determination model for each of the multiple devices 4000 from the IoT server 3000 through the communication interface 2100. According to an embodiment, the voice assistant server 2000 may receive user account information from the hub device 1000 through the communication interface 2100, and may send device information about the multiple devices 4000 registered according to the received user account information and information about the function determination model to the hub device 1000.
[0232] According to an embodiment, the processor 2200 and the memory 2300 of the voice assistant server 2000 may perform functions that are the same as or similar to those of the processor 1200 (see Figure 2 ) of the hub device 1000 (see Figure 2 ) and the memory 1300 (see Figure 2 ). Therefore, a description of the processor 2200 and the memory 2300 of the voice assistant server 2000 that is the same as the description of the processor 1200 and the memory 1300 of the hub device 1000 is not provided here.
[0233] According to an embodiment, the memory 2300 of the voice assistant server 2000 may store data on one or more or all of the following: the ASR module 2310, data on the NLG module 2320, data on the device determination model 2330, and data corresponding to each of the multiple function determination models 2340. According to an embodiment, the memory 2300 of the voice assistant server 2000 may store multiple function determination models 2340 corresponding to multiple devices related to multiple different user accounts, rather than the multiple function determination models 1350 (see Figure 2 ) stored in the memory 1300 of the hub device 1000. In addition, according to an embodiment, multiple function determination models 2340 for more types of devices than the multiple function determination models 1350 stored in the memory 1300 of the hub device 1000 may be stored in the memory 2300 of the voice assistant server 2000. The total capacity of the multiple function determination models 2340 stored in the memory 2300 of the voice assistant server 2000 may be greater than the capacity of the multiple function determination models 1350 stored in the memory 1300 of the hub device 1000.
[0234] According to an embodiment, when at least a part of the text is received from the hub device 1000, the communication interface 2100 of the voice assistant server 2000 may send at least a part of the received text to the processor 2200, and the processor 2200 may analyze at least a part of the text by using the first NLU model 2332 stored in the memory 2300. According to an embodiment, the processor 2200 may determine an operation execution device related to at least a part of the text based on the analysis result by using the data of the device determination model 2330 stored in the memory 2300. According to an embodiment, the processor 2200 may select a function determination model corresponding to the operation execution device from a plurality of function determination models stored in the memory 2300, and may obtain operation information about the detailed operations for executing the functions of the operation execution device and the relationships between the detailed operations by using the selected function determination model.
[0235] For example, based on determining that the operation execution device is the third device 4300 that functions as an air purifier, the processor 2200 may analyze at least a part of the text by using the second NLU model 2346a of the function determination model 2346 corresponding to the air purifier, and may obtain operation information by planning the detailed operations to be executed by the device and the execution order of the detailed operations by using the action plan management module 2346b. The detailed description is the same as the description of the hub device 1000, and thus the repeated description will be omitted.
[0236] According to an embodiment, based on the hub device 1000 sending at least a part of the entire text of the user voice, the voice assistant server 2000 may determine an operation execution device related to at least a part of the entire text, and may obtain operation information for performing the operations of the operation execution device. According to an embodiment, the part of the entire text may be less than the entire text of the user voice. According to an embodiment, the voice assistant server 2000 may send the obtained operation information to the IoT server 3000 through the communication interface 2100.
[0237] According to an embodiment, the voice assistant server 2000 may receive update request information of a function determination model and device identification information from an operation execution device among a plurality of devices 4000. According to an embodiment, the update request information of the function determination model may be information for requesting synchronization of the version of the function determination model stored in the memory of the operation execution device with the version of the function determination model corresponding to the operation execution device among the plurality of function determination models 2342, 2344, 2346, and 2348 stored in the voice assistant server 2000. According to an embodiment, the processor 2200 of the voice assistant server 2000 may identify the function determination model corresponding to the operation execution device based on the device identification information received from the operation execution device, and may check the version information of the identified function determination model. According to an embodiment, the processor 2200 may send update data for updating the function determination model of the operation execution device to the operation execution device by using the communication interface 2100.
[0238] In an embodiment of the present disclosure, based on the operation execution device receiving a control command from the IoT server 3000, the processor 2200 may control the communication interface 2100 to send update data for updating the function determination model to the operation execution device.
[0239] In an embodiment of the present disclosure, the processor 2200 may control the communication interface 2100 to periodically send update data of the function determination model to a device storing the function determination model among the plurality of devices 4000 connected to the user account through the hub device 1000. In an embodiment of the present disclosure, when the processor 2200 sends update data for updating an application or firmware to a device storing the function determination model among the plurality of devices 4000 connected to the user account through the hub device 1000, the processor 2200 may control the communication interface 2100 to also send update data of the function determination model.
[0240] In an embodiment of the present disclosure, a new device may be registered in the user account information registered with the IoT server 300. In this case, the processor 2200 may receive, from the IoT server 3000 by using the communication interface 2100, the user account information, a device list updated to include the new device, and storage information of the device determination model and the function determination model for each device candidate included in the device list. According to an embodiment, the processor 2200 may identify, based on the storage information of the device determination model, a device storing the device determination model among the plurality of devices registered according to the user account, and may determine the identified device as the hub device. According to an embodiment, the processor 2200 may send the device identification information of the new device and the storage information of the device determination model and the function determination model to the hub device by using the communication interface 2100. Refer to Figure 19Describe an embodiment of registering a new device.
[0241] Figure 4 FIG. is a block diagram illustrating components of an IoT server 3000 according to an embodiment of the present disclosure.
[0242] According to an embodiment, the IoT server 3000 may be a server that obtains, stores, and manages device information about each of a plurality of devices 4000 (see Figure 1 ). According to an embodiment, the IoT server 3000 may obtain, determine, or generate a control command for controlling a device by using the stored device information. Although the IoT server 3000 is implemented as an independent hardware device separate from the voice assistant server 2000 in Figure 1 , the present disclosure is not limited thereto. In an embodiment of the present disclosure, the IoT server 3000 may be a component of the voice assistant server 2000 (see Figure 1 ), or may be a server designed to be classified as software.
[0243] According to an embodiment, referring to Figure 4 , the IoT server 3000 may at least include a communication interface 3100, a processor 3200, and a memory 3300.
[0244] According to an embodiment, the IoT server 3000 may be connected to an operation execution device or the voice assistant server 2000 via a network through the communication interface 3100, and may receive or send data. According to an embodiment, the IoT server 3000 may send data stored in the memory 3300 to the voice assistant server 2000 or the operation execution device through the communication interface 3100 under the control of the processor 3200. In addition, the IoT server 3000 may receive data from the voice assistant server 2000 or the operation execution device through the communication interface 3100 under the control of the processor 3200.
[0245] In an embodiment of the present disclosure, the communication interface 3100 may receive device information including at least one of device identification information (e.g., device ID information), function execution ability information, location information, and status information from each of a plurality of devices 4000 (see Figure 1 ). In an embodiment of the present disclosure, the communication interface 3100 may receive user account information from each of a plurality of devices 4000. In addition, the communication interface 3100 may receive information about power on / off or an operation being performed from a plurality of devices 4000. According to an embodiment, the communication interface 3100 may provide the received device information to the memory 3300.
[0246] According to an embodiment, the memory 3300 may store device information received through the communication interface 3100. In an embodiment of the present disclosure, the memory 3300 may classify the device information according to user account information received from multiple devices 4000, and may store the classified device information in the form of a lookup table.
[0247] In an embodiment of the present disclosure, the communication interface 3100 may receive, from the voice assistant server 2000, a query for requesting user account information and device information about multiple devices 4000 previously registered by using the user account information. According to an embodiment, in response to the received query, the processor 3200 may obtain, from the memory 3300, the device information about multiple devices 4000 previously registered by using the user account, and may control the communication interface 3100 to send the obtained device information to the voice assistant server 2000.
[0248] According to an embodiment, the processor 3200 may control the communication interface 3100 to send a control command to an operation execution device determined to perform an operation based on operation information received from the voice assistant server 2000. According to an embodiment, the IoT server 3000 may receive, through the communication interface 3100, an operation execution result according to the control command from the operation execution device.
[0249] Figure 5 is a block diagram showing some elements of multiple devices 4000 according to an embodiment of the present disclosure.
[0250] According to an embodiment, the multiple devices 4000 may be devices controlled by the hub device 1000 (see Figure 1 ) or the IoT server 3000 (see Figure 1 ). In an embodiment of the present disclosure, the multiple devices 4000 may be actuator devices that perform operations based on control commands received from the hub device 1000 or the IoT server 3000. According to an embodiment, the multiple devices 4000 may be IoT devices.
[0251] Referring to Figure 5 , the multiple devices 4000 may include a first device 4100, a second device 4200, and a third device 4300. Although in the Figure 5 embodiment, the first device 4100 is an air conditioner, the second device 4200 is a TV, and the third device 4300 is an air purifier, this is merely an example, and the multiple devices 4000 of the present disclosure are not limited to Figure 5 the devices shown.
[0252] According to an embodiment, each of the first device 4100, the second device 4200, and the third device 4300 may be as shown in Figure 5The processor, memory, and communication interface shown. According to an embodiment, it may further include a plurality of devices 4000 based on the elements required to control the execution of operations. According to an embodiment, the required elements may be, for example, a fan for an air conditioner.
[0253] According to an embodiment, one or more of the plurality of devices 4000 may store a function determination model. In Figure 5 the embodiment, the first device 4100 may include a communication interface 4110, a processor 4120, and a memory 4130, and the function determination model 4132 may be stored in the memory 4130. According to an embodiment, the function determination model 4132 stored in the first device 4100 may be a model for obtaining operation information about the detailed operations for executing the operations of the first device 4100 and the relationships between the detailed operations. According to an embodiment, the function determination model 4132 may include a first NLU model 4134, configured to analyze at least a part of the text received from the hub device 1000 or the IoT server 3000, and obtain operation information about the operations to be performed by the first device 4100 based on the analysis result of at least a part of the text. According to an embodiment, the function determination model 4132 may include an action plan management module 4136, configured to manage operation information related to the detailed operations of the device so as to generate the detailed operations to be performed by the first device 4100 and the execution order of the detailed operations. The action plan management module 4136 may plan the detailed operations to be performed by the first device 4100 and the execution order of the detailed operations based on the analysis result of at least a part of the text.
[0254] According to an embodiment, the second device 4200 may include a communication interface 4210, a processor 4220, and a memory 4230. According to an embodiment, the third device 4300 may include a communication interface 4310, a processor 4320, and a memory 4330. According to an embodiment, different from the first device 4100, the second device 4200 and the third device 4300 may not store a function determination model. According to an embodiment, the second device 4200 and the third device 4300 may not receive at least a part of the text from the hub device 1000 (see Figure 1 ) or the IoT server 3000 (see Figure 1 ). According to an embodiment, the second device 4200 and the third device 4300 may receive a control command from the hub device 1000 or the IoT server 3000, and may perform operations based on the received control command.
[0255] According to an embodiment, a plurality of devices 4000 may send information about user account information, device information, and information about a function determination model to the IoT server 3000 by using communication interfaces 4110, 4210, and 4310. In an embodiment of the present disclosure, when a user logs in to the plurality of devices 4000, the plurality of devices 4000 may send user account information, information about whether each of the plurality of devices 4000 stores a device determination model itself, and information about whether each of the plurality of devices 4000 stores a function determination model itself to the IoT server 3000. According to an embodiment, the device information may include at least one of identification information (e.g., device id information) of the plurality of devices 4000, a device type of each of the plurality of devices 4000, a function execution ability of each of the plurality of devices 4000, location information, and status information.
[0256] In an embodiment of the present disclosure, the plurality of devices 4000 may send user account information, information about whether each of the plurality of devices 4000 stores a device determination model itself, and information about whether each of the plurality of devices 4000 stores a function determination model itself to the hub device 1000 (see Figure 1 ) by using communication interfaces 4110, 4210, and 4310.
[0257] According to an embodiment, one or more of the plurality of devices 4000 may be third-party devices manufactured by a manufacturer different from that of the hub device 1000. For example, the third device 4300 may be a third-party device. When the third device 4300 is a third-party device, a third-party user account may be used to log in to a third-party IoT server through the third device 4300, and the IoT server 3000 may access the third-party IoT server by using a temporary account and may obtain device information and information about a function determination model of the third device 4300. According to an embodiment, based on determining that the third device 4300 is an operation execution device, the IoT server 3000 may convert a control command into a control command readable and executable by the third device 4300 and may send the control command to the third-party IoT server. According to an embodiment, the third-party IoT server may send the converted control command to the third device 4300, and the third device 4300 may perform an operation based on the received control command.
[0258] Figure 6 is a flowchart of a method for controlling a device based on voice input executed by the hub device 1000 according to an embodiment of the present disclosure.
[0259] According to an embodiment, in operation S610, the hub device 1000 may receive voice data (e.g., a user's voice input). In an embodiment of the present disclosure, the hub device 1000 may receive a voice input (e.g., a user's utterance) from the user through the microphone 1100 (see Figure 2 ), and may obtain a voice signal or voice data from the received voice input. In an embodiment of the present disclosure, the processor 1200 of the hub device 1000 (see Figure 2 ) may convert the voice data (e.g., the sound received through the microphone 1100) into an acoustic signal, and may obtain a voice signal by removing noise (e.g., non-speech components) from the acoustic signal.
[0260] According to an embodiment, in operation S620, the hub device 1000 may convert the received voice input into text by performing ASR. In an embodiment of the present disclosure, the hub device 1000 may perform ASR, which converts a voice signal into computer-readable text by using a predefined model such as an acoustic model (AM) or a language model (LM). When an acoustic signal without noise removed is received, the hub device 1000 may obtain a voice signal by removing noise from the received acoustic signal, and may perform ASR on the voice signal.
[0261] According to an embodiment, in operation S630, the hub device 1000 may analyze the text by using a first NLU model, and may determine an operation execution device corresponding to the analyzed text by using a device determination model. In an embodiment of the present disclosure, the hub device 1000 may analyze the text by using the first NLU model included in the device determination model, and may determine, based on the analysis result of the text, an operation execution device for performing an operation according to the user's intention from multiple devices. According to an embodiment, the multiple devices may refer to devices that are logged in using the same user account as the user account of the hub device 1000 and are connected to the hub device 1000 through a network, such as IoT devices. The multiple IoT devices may be devices registered to an IoT server by using the same user account as the user account of the hub device 1000.
[0262] According to an embodiment, the first NLU model may be a model trained to analyze text converted from a voice input and determine an operation execution device based on the analysis result. According to an embodiment, the first NLU model may be used to determine an intention by interpreting text and determine an operation execution device based on the intention. According to an embodiment, the hub device 1000 may parse text in units of morphemes, words, or phrases by using the first NLU model, and may infer the meaning of words extracted from the parsed text by using the linguistic features (e.g., grammatical components) of the morphemes, words, or phrases. According to an embodiment, the hub device 1000 may compare the inferred meaning of the words with predefined intentions provided by the first NLU model and may determine the intention corresponding to the inferred meaning of the words.
[0263] According to an embodiment, the hub device 1000 may determine, based on a matching model for determining the relationship between an intention and a device, a device related to the intention recognized from the text as an operation execution device. In an embodiment of the present disclosure, the matching model may be obtained through learning via a rule-based system, but the present disclosure is not limited thereto.
[0264] In an embodiment of the present disclosure, the hub device 1000 may obtain a plurality of numerical values indicating the degree of relationship between an intention and a plurality of devices by applying the matching model to the intention, and may determine, as a final operation execution device, a device having the maximum value among the obtained plurality of numerical values. For example, when an intention is related to each of a first device and a second device, the hub device 1000 may obtain a first numerical value indicating the degree of relationship between the intention and the first device and a second numerical value indicating the degree of relationship between the intention and the second device, and may determine the first device having the larger value among the first numerical value and the second numerical value as the operation execution device.
[0265] Although the hub device 1000 may train a matching model between an intention and an operation execution device by using, for example, a rule-based system, the present disclosure is not limited thereto. The AI model used by the hub device 1000 may include, for example, a neural network-based system (e.g., CNN or RNN), SVM, linear regression, logistic regression, naive Bayes, random forest, decision tree, or k-nearest neighbor algorithm. Alternatively, the AI model may be a combination of the above examples or any other AI model.
[0266] According to an embodiment, the device determination model 1330 (see Figure 2 ) stored in the memory 1300 (see Figure 2)The operation execution device can be determined from among a plurality of devices registered according to the user account of the hub device 1000. In an embodiment of the present disclosure, the device determination model can analyze a numerical value indicating the degree of relationship between the indication intention and a plurality of previously registered devices logged in using the same user account as the user account of the hub device 1000, and can determine, as the operation execution device, the device having the maximum value among the numerical values indicating the degree of relationship between the indication intention and the plurality of devices.
[0267] In an embodiment of the present disclosure, the hub device 1000 can receive device information about each of a plurality of devices previously registered according to the user account from the voice assistant server 2000 (see Figure 1 ). The device information can include at least one of, for example, identification information (e.g., device ID information) of each of the plurality of devices, the device type of each of the plurality of devices, the function execution ability of each of the plurality of devices, location information, and status information. The hub device 1000 can determine, based on the received device information, a device for performing an operation according to the intention from among the plurality of devices.
[0268] According to an embodiment, in operation S640, the hub device 1000 can identify a device storing a function determination model corresponding to the operation execution device determined from among the plurality of devices.
[0269] According to an embodiment, the function determination model corresponding to the operation execution device can be stored in the memory of the hub device 1000, can be stored in the memory of the operation execution device itself, or can be stored in the memory of the voice assistant server 2000 (see Figure 1 ). According to an embodiment, the term "function determination model corresponding to the operation execution device" can refer to a model for obtaining operation information related to detailed operations for performing an operation according to the determined function of the operation execution device and the relationship between the detailed operations.
[0270] In an embodiment of the present disclosure, the hub device 1000 can use information about whether it is stored in the database 1360 (see Figure 2) The information of the function determination model for each of the multiple devices and the information about the storage location of the function determination model corresponding to each of the multiple devices are used to identify the device storing the function determination model from the memory of the hub device 1000, the memory of the voice assistant server 2000, or the memory of the operation execution device itself. According to an embodiment, the hub device 1000 may obtain the information about the function determination model from at least one device among the multiple devices that stores the function determination model for determining the functions related to each of the multiple devices, and may store the obtained information about the function determination model in the database 1360. According to an embodiment, based on the received information about the function determination model of the multiple devices, the hub device 1000 may also receive the information about the storage location of the function determination model for each of the multiple devices (e.g., device identification information, IP address, or MAC address). According to an embodiment, the information about whether to store the information of the function determination model for each of the multiple devices previously registered according to the user account and the information about the storage location of the function determination model for each of the multiple devices may be stored in the database 1360 in the form of a lookup table.
[0271] In an embodiment of the present disclosure, the hub device 1000 may obtain the device identification information of the operation execution device determined in operation S630, may search the database 1360 for the lookup table according to the device identification information, and may obtain the information about the storage location of the function determination model corresponding to the operation execution device based on the search result of the lookup table. By using the above method, the hub device 1000 can identify the device storing the function determination model corresponding to the operation execution device from the memory of the hub device 1000, the memory of the voice assistant server 2000, and the memory of the operation execution device itself.
[0272] According to an embodiment, in operation S650, the hub device 1000 may provide at least a part of the text to the identified device. For example, when it is checked in operation S640 that the function determination model corresponding to the operation execution device is stored in the hub device 1000, the hub device 1000 may provide at least a part of the text to the function determination model corresponding to the operation execution device.
[0273] For example, when it is checked in operation S640 that the function determination model corresponding to the operation execution device is stored in the voice assistant server 2000, the hub device 1000 may send at least a part of the text to the voice assistant server 2000.
[0274] For example, when it is checked in operation S640 that the function determination model corresponding to the operation execution device is stored in the memory of the operation execution device itself, the hub device 1000 may send at least a part of the text to the operation execution device.
[0275] In an embodiment of the present disclosure, the hub device 1000 may provide only a part of the text and not the entire text to the identified device. For example, when the text converted from a voice input is "Play the movie Avengers on the TV", "on the TV" specifies the name of the operation execution device, and thus may be unnecessary information for the TV function determination model 1354. According to an embodiment, the hub device 1000 may parse the text in units of words or phrases by using the first NLU model 1332 (see Figure 2 ), identify words or phrases that specify the name, common name, or installation location of the device, and provide the remaining part of the text except for the words or phrases identified in the entire text to the identified device.
[0276] Figure 7 is a flowchart illustrating a method of providing at least a part of the text to one of the hub device 1000, the voice assistant server 2000, and the operation execution device 4000a according to a voice input of a user, performed by the hub device 1000 according to an embodiment of the present disclosure. Figure 7 Illustrates Figure 6 detailed embodiments of operations S640 and S650 of Figure 7 Operations S710 to S730 of Figure 6 are detailed embodiments of operation S640, and operations S740 to S760 are
[0277] According to an embodiment, operation S710 may be performed after operation S630 of Figure 6 . In operation S710, according to an embodiment, the hub device 1000 may obtain information about the function determination model from at least one device among a plurality of devices that stores the function determination model. In an embodiment of the present disclosure, the hub device 1000 may receive information about the function determination model from at least one device that stores a function determination model for determining functions related to each of the plurality of devices by using the communication interface 1400 (see Figure 2 ). According to an embodiment, the term "information about the function determination model" may refer to information about whether each of the plurality of devices itself stores a function determination model for obtaining operation information about detailed operations performed according to functions and the relationships between the detailed operations. In an embodiment of the present disclosure, when receiving information about the function determination model from a plurality of devices, the hub device 1000 may also obtain information about the storage location of the function determination model of each of the plurality of devices (e.g., device identification information, IP address, or MAC address).
[0278] According to an embodiment, the hub device 1000 may store the received information about the function determination model and the information about the storage location of the function determination model for each device in the database 1360 (see Figure 2 ). In an embodiment of the present disclosure, the hub device 1000 may store the information about whether to store the function determination model and the information about the storage location of the function determination model in the database 1360 in the form of a lookup table according to the name or identification information of the device.
[0279] According to an embodiment, in operation S720, the hub device 1000 may check whether the function determination model corresponding to the determined operation execution device is stored in the memory of the hub device 1000. According to an embodiment, the term "function determination model corresponding to the operation execution device" may be a model for obtaining operation information about detailed operations for performing operations according to the determined function of the operation execution device and the relationships between the detailed operations.
[0280] According to an embodiment, the function determination model corresponding to the operation execution device may be stored in the memory 1300 of the hub device 1000 (see Figure 2 ), may be stored in the memory of the operation execution device itself, or may be stored in the voice assistant server 2000.
[0281] In an embodiment of the present disclosure, the hub device 1000 may obtain the information about the function determination model for each device by searching the lookup table stored in the database 1360. In an embodiment of the present disclosure, the hub device 1000 may check whether the function determination model corresponding to the operation execution device is stored in the memory 1300 of the hub device 1000 based on the obtained information about the function determination model. In another embodiment of the present disclosure, the hub device 1000 may determine whether the function determination model corresponding to the operation execution device is stored in the memory 1300 by accessing the memory 1300 and scanning the data stored in the memory 1300.
[0282] Multiple function determination models of multiple devices may be stored in the memory 1300 of the hub device 1000. In an embodiment of the present disclosure, multiple function determination models corresponding to multiple devices that log in using the same account as the user account and are connected to the hub device 1000 through the network may be stored in the memory 1300 of the hub device 1000.
[0283] According to an embodiment, in operation S730, based on determining that there is no function determination model corresponding to the operation execution device in the memory 1300 of the hub device 1000 in operation S720, the hub device 1000 may check whether the function determination model corresponding to the operation execution device is stored in the memory of the operation execution device based on the information about the function determination model. In an embodiment of the present disclosure, the hub device 1000 may check whether the function determination model corresponding to the operation execution device is stored in the memory of the operation execution device itself by searching the lookup table stored in the database 1360. For example, when the operation execution device determined based on the text is the first device, the hub device 1000 may check the information about whether the function determination model is stored in the first device itself by searching the lookup table.
[0284] In an embodiment of the present disclosure, the function determination model may not be stored in the memory of the first device determined to be the operation execution device. According to an embodiment, in operation S740, when the hub device 1000 determines in operation S730 that the function determination model is not stored in the operation execution device based on the information about the function determination model, the hub device 1000 sends at least a part of the text to the voice assistant server. In an embodiment of the present disclosure, the hub device 1000 may send at least a part of the text to the function determination model corresponding to the operation execution device stored in the voice assistant server. In this case, the hub device 1000 may also send the device identification information of the operation execution device together with at least a part of the text.
[0285] According to an embodiment, in operation S750, when the hub device 1000 checks in operation S720 that the function determination model corresponding to the operation execution device is stored in the memory of the hub device 1000, the hub device 1000 may provide at least a part of the text to the previously stored function determination model. In an embodiment of the present disclosure, the hub device 1000 may select the function determination model corresponding to the operation execution device from among a plurality of previously stored function determination models in the memory 1300 (see Figure 2 ) and may provide at least a part of the text to the selected function determination model. For example, when the operation execution device is the second device, the processor 1200 of the hub device 1000 may select the second function determination model corresponding to the second device from the first function determination model and the second function determination model previously stored in the memory 1300 and may provide at least a part of the text to the second function determination model.
[0286] According to an embodiment, in operation S760, when the hub device 1000 checks in operation S730 that the function determination model corresponding to the operation execution device is stored in the memory of the operation execution device itself, the hub device 1000 may send at least a part of the text to the function determination model stored in the operation execution device.
[0287] According to an embodiment, in operations S740, S750, and S760, the hub device 1000 may send only at least a part of the text, rather than the entire text. In an embodiment of the present disclosure, the processor 1200 of the hub device 1000 may formulate the text in units of words or phrases by using the first NLU model, may identify words or phrases that specify a name, a general name, or the installation location of a device, and may send the remaining part of the text except for the words or phrases identified in the entire text to the function determination model of the operation execution device.
[0288] Figure 8 It is a flowchart illustrating a method of operating a hub device 1000, a voice assistant server 2000, an IoT server 3000, and an operation execution device 4000a according to an embodiment of the present disclosure.
[0289] Figure 8 It is illustrated in Figure 7 the steps depicted in After that, it is a flowchart of the operations of entities in a multi-device system environment including a hub device 1000, a voice assistant server 2000, an IoT server 3000, and an operation execution device 4000a. Figure 7 the steps depicted in It indicates a state where it is checked that the function determination model corresponding to the operation execution device 4000a is not stored in both the memory of the hub device 1000 and the memory of the operation execution device 4000a itself.
[0290] Referring to Figure 8 , the hub device 1000 may include an ASR module 1310, an NLG module 1320, a device determination model 1330, and a function determination device determination module 1340. In Figure 8 the embodiment of , the hub device 1000 may not store the function determination model, or may not store the function determination model corresponding to the operation device.
[0291] According to an embodiment, the voice assistant server 2000 may store multiple function determination models 2342, 2344, and 2346. For example, the function determination model 2342, which is stored in the voice assistant server 2000 as the first function determination model, may be a model for determining the function of an air conditioner and obtaining operation information regarding detailed operations related to the determined function and the relationships between the detailed operations. For example, the function determination model 2344, which is the second function determination model, may be a model for determining the function of a TV and obtaining operation information regarding detailed operations related to the determined function and the relationships between the detailed operations, and the function determination model 2346 may be a model for determining the function of a washing machine and obtaining operation information regarding detailed operations related to the determined function and the relationships between the detailed operations.
[0292] In Figure 8 the embodiment, it may be determined that the operation execution device 4000a is a "TV", and the function determination model 2344 corresponding to the TV may be stored in the voice assistant server 2000.
[0293] In operation S810, the hub device 1000 sends at least a part of the text and the identification information of the operation execution device 4000a to the voice assistant server 2000. The hub device 1000 may send at least a part of the text converted from the voice input to the voice assistant server 2000 without sending the entire text. In the embodiments of the present disclosure, based on determining that the operation execution device 4000a is a "TV", in the text corresponding to "play the movie Avengers on the TV", "on the TV" specifies the name or common name of the operation execution device 4000a, so it may be unnecessary information. In addition, since the hub device 1000 sends the identification information (e.g., device id) of the operation execution device 4000a to the voice assistant server 2000, "on the TV" is unnecessary information. The processor 1200 of the hub device 1000 (see Figure 2 ) may parse the text in units of words or phrases by using the first NLU model 1332, may identify words or phrases specifying the name, common name, or installation location of the device, and may send the remaining part of the text except for the words or phrases identified in the entire text to the voice assistant server 2000.
[0294] In operation S810, while receiving at least a part of the text and the identification information of the operation execution device 4000a, the hub device 1000 may send the user account information of the hub device 1000 and the operation execution device 4000a to the voice assistant server 2000.
[0295] According to an embodiment, in operation S820, the voice assistant server 2000 selects a function determination model corresponding to the operation execution device 4000a. In an embodiment of the present disclosure, the voice assistant server 2000 may identify the operation execution device 4000a based on the identification information of the operation execution device 4000a received from the hub device 1000, and may select a function determination model corresponding to the operation execution device 4000a from among the plurality of function determination models 2342, 2344, and 2346. The term "function determination model corresponding to the operation execution device" refers to a model for obtaining operation information regarding detailed operations for executing the determined function of the operation execution device 4000a and the relationships between the detailed operations. In Figure 8 an embodiment, when the operation execution device 4000a is a TV, the voice assistant server 2000 may select, from among the plurality of function determination models 2342, 2344, and 2346 stored in the memory, a function determination model 2344 for obtaining operation information regarding detailed operations for executing an operation according to the function of the TV, for example, the function of playing a movie, and the relationships between the detailed operations.
[0296] In operation S830, the voice assistant server 2000 interprets the text by using the NLU model of the selected function determination model and determines the intent based on the analysis result. The voice assistant server 2000 may analyze at least a part of the text received from the hub device 1000 by using the NLU model 2344a of the function determination model 2344. The NLU model 2344a, as an AI model trained to interpret text related to a specific device, may be a model trained to determine the intent and parameters related to the operation expected by the user. The NLU model 2344a may be a model trained to determine the function related to the type of a specific device when the input text is received.
[0297] In an embodiment of the present disclosure, the voice assistant server 2000 may parse at least a part of the text in units of words or phrases by using the NLU model 2344a, may infer the meaning of the words extracted from the parsed text by using the linguistic features (e.g., grammatical elements) of the parsed morphemes, words, or phrases, and may obtain the intent and parameters from the text by matching the inferred meaning with predefined intents and parameters. The intent, as information indicating the intent of the user's utterance included in the text, may be used to determine the operation to be executed by the operation execution device 4000a. The parameter refers to variable information for determining the detailed operation of the operation execution device 4000a related to the intent. The parameter may be information corresponding to the intent, and multiple types of parameters may correspond to one intent. For example, when the text is "Play the movie Avengers on the TV", the intent may be "Video content playback", and the parameter may be "movie" or "movie Avengers".
[0298] In an embodiment of the present disclosure, the voice assistant server 2000 may determine an intention from at least a part of the text.
[0299] In operation S840, the voice assistant server 2000 obtains operation information about an operation to be performed by the operation execution device 4000a based on the intention. In an embodiment of the present disclosure, the voice assistant server 2000 plans the operation information to be performed by the operation execution device 4000a based on the intention and parameters by using the action plan management module 2344b of the function determination model 2344. The action plan management module 2344b may interpret the operation to be performed by the operation execution device 4000a based on the intention and parameters. The action plan management module 2344b may select detailed operations related to the interpreted operation from among the operations of the previously stored device, and may plan the execution order of the selected detailed operations. The action plan management module 2344b may obtain operation information about the detailed operations to be performed by the operation execution device 4000a by using the planning result. The term "operation information" may be information related to the detailed operations to be performed by the device, the relationship between the detailed operations, and information related to the execution order of the detailed operations. The operation information may include, but is not limited to, the functions to be performed by the operation execution device 4000a to perform the detailed operations, the execution order of the functions, the input values required to execute the functions, and the output values output as the execution results of the functions.
[0300] In operation S850, the voice assistant server 2000 sends the obtained operation information and the identification information of the operation execution device 4000a to the IoT server 3000.
[0301] In operation S860, the IoT server 3000 obtains a control command based on the identification information of the operation execution device 4000a and the received operation information. The IoT server 3000 may include a database in which control commands and operation information for multiple devices are stored. In an embodiment of the present disclosure, the IoT server 3000 may select, based on the identification information of the operation execution device 4000a, a control command for controlling the detailed operations of the operation execution device 4000a from among the control commands related to the multiple devices previously stored in the database.
[0302] According to an embodiment, the control command, as information readable and executable by the device, may include instructions for sequentially performing detailed operations in the execution order according to the operation information when the device executes the function.
[0303] According to an embodiment, in operation S870, the IoT server 3000 may send a control command to the operation execution device 4000a by using the identification information of the operation execution device 4000a.
[0304] In operation S880, the operation execution device 4000a may execute an operation corresponding to the received control command. For example, the operation execution device 4000a may play a movie based on the control command.
[0305] In an embodiment of the present disclosure, after executing the operation, the operation execution device 4000a may send information about the operation execution result to the IoT server 3000.
[0306] Figure 9 It is a flowchart illustrating a method for an operation hub device 1000 and an operation execution device 4000a according to an embodiment of the present disclosure.
[0307] Figure 9 It is illustrated in Figure 7 the step of a multi-device system including the hub device 1000 and the operation execution device 4000a after Figure 7 the depicted step indicates a state in which it is checked that the function determination model corresponding to the operation execution device 4000a is not stored in the memory of the hub device 1000 but in the memory of the operation execution device 4000a itself. Although Figure 9 the voice assistant server 2000 and the IoT server 3000 are not shown in
[0308] Referring to Figure 9 , the hub device 1000 may include an ASR module 1310, an NLG module 1320, a device determination model 1330, and a function determination device determination module 1340. In Figure 8 the embodiment of
[0309] the operation execution device 4000a may store the function determination model 4032 of the operation execution device 4000a itself. The operation execution device 4000a itself may analyze the text, and the function determination model 4032 for executing an operation based on the analysis result of the text according to the user's intention may be stored in the memory of the operation execution device 4000a. For example, when the operation execution device 4000a is an "air conditioner", the function determination model 4032 stored in the memory of the operation execution device 4000a may be a model for determining the functions of the air conditioner and obtaining operation information about detailed operations related to the determined functions and the relationships between the detailed operations.
[0310] In operation S910, the hub device 1000 sends at least a part of the text and a text transmission notification signal to the function determination model 4032 of the operation execution device 4000a. The hub device 1000 may send at least a part of the text converted from voice input to the operation execution device 4000a without sending the entire text. For example, in the text "Lower the temperature by 2°C in the air conditioner", "in the air conditioner" specifies the name or general name of the operation execution device 4000a, so it may be unnecessary information. In addition, since the operation execution device 4000a itself is an air conditioner, "in the air conditioner" is unnecessary information. The processor 1200 of the hub device 1000 (see Figure 2 ) can parse the text in units of words or phrases by using the first NLU model 1332, can identify words or phrases specifying names, general names, or installation locations of devices, and can send the remaining part of the text except for the words or phrases identified in the entire text to the operation execution device 4000a.
[0311] In an embodiment of the present disclosure, the hub device 1000 may send a text transmission notification signal to the operation execution device 4000a. The text transmission notification signal is a signal notifying that the text is sent to the operation execution device 4000a. When the operation execution device 4000a receives the text transmission notification signal, the ASR operation, the operation of determining the operation of the operation execution device, and the operation of selecting the function determination model may be omitted, and the operation execution device 4000a may directly provide at least a part of the text to the function determination model 4032 and may determine the intention from the text by using the function determination model 4032.
[0312] However, the present disclosure is not limited thereto, and in operation S910, the hub device 1000 may not send a text transmission notification signal to the operation execution device 4000a. In this case, when the operation execution device 4000a receives at least a part of the text from the hub device 1000, the operation execution device 4000a may be preset to provide at least a part of the received text to the function determination model 4032. In an embodiment of the present disclosure, data regarding a policy set to identify the received text and provide the text to the function determination model 4032 may be stored in the memory of the operation execution device 4000a.
[0313] According to an embodiment, in operation S920, the operation execution device 4000a may determine an intention by interpreting text using the NLU model 4034. In an embodiment of the present disclosure, the operation execution device 4000a may analyze at least a part of the text received from the hub device 1000 by using the NLU model 4034 of the function determination model 4032. The NLU model 4034, as an AI model trained to interpret text related to the operation execution device 4000a, may be a model trained to determine an intention and parameters related to an operation of a user intention. The NLU model 4034 may be a model trained to determine a function related to a type of a specific device when an input text is received.
[0314] In an embodiment of the present disclosure, the operation execution device 4000a may parse at least a part of the text in units of words or phrases by using the NLU model 4034, may infer the meaning of a word extracted from the parsed text by using linguistic features (e.g., grammatical elements) of the parsed morphemes, words, or phrases, and may obtain an intention and parameters from the text by matching the inferred meaning with predefined intentions and parameters. The intention, as information indicating the user's utterance intention included in the text, may be used to determine an operation to be performed by the operation execution device 4000a. The parameter refers to variable information for determining a detailed operation of the operation execution device 4000a related to the intention. The parameter may be information corresponding to the intention, and multiple types of parameters may correspond to one intention. For example, when the text is "lower the temperature by 2°C in the air conditioner", the intention may be "set temperature control" or "set temperature decrease", and the parameter may be "2°C".
[0315] In an embodiment of the present disclosure, the operation execution device 4000a may determine an intention only from at least a part of the text.
[0316] According to an embodiment, in operation S930, the operation execution device 4000a obtains operation information about an operation to be performed by the operation execution device 4000a based on an intention. In an embodiment of the present disclosure, the operation execution device 4000a plans the operation information of the operation to be performed by the operation execution device 4000a based on the intention and parameters by using the action plan management module 4036 of the function determination model 4032. The action plan management module 4036 can interpret the operation to be performed by the operation execution device 4000a based on the intention and parameters. The action plan management module 4036 can select detailed operations related to the interpreted operation from the previously stored operations of the device and can plan the execution order of the selected detailed operations. The action plan management module 4036 can obtain operation information about the detailed operations to be performed by the operation execution device 4000a by using the planning result. The term "operation information" may be information related to the detailed operations to be performed by the device, the relationship between the detailed operations, and the execution order of the detailed operations. The operation information may include, but is not limited to, the functions performed by the operation execution device 4000a to execute the detailed operations, the execution order of the functions, the input values required to execute the functions, and the output values output as the execution results of the functions.
[0317] According to an embodiment, in operation S940, the operation execution device 4000a obtains a control command based on the operation information. In an embodiment of the present disclosure, a plurality of control commands corresponding to a plurality of different operation information may be stored in the memory of the operation execution device 4000a. In an embodiment of the present disclosure, the operation execution device 4000a can select a control command for controlling the detailed operation according to the operation information from among the plurality of control commands previously stored in the memory. The control command, as information readable and executable by the operation execution device 4000a, may include instructions for sequentially executing the detailed operations according to the operation information in the execution order when the operation execution device 4000a executes the functions.
[0318] According to an embodiment, in operation S950, the operation execution device 4000a can perform an operation corresponding to the control command according to the control command. For example, the operation execution device 4000a can reduce the set temperature by 2°C based on the control command.
[0319] Figure 10 It is a flowchart illustrating a method of an operation hub device 1000 and an operation execution device 4000a according to an embodiment of the present disclosure.
[0320] Figure 10 It is illustrated in Figure 7 the step of After that, it is a flowchart of the operations of entities in a multi-device system environment including the hub device 1000 and the operation execution device 4000a. Figure 7 The steps shown Indicates a state where it is detected that a function determination model corresponding to the operation execution device 4000a is stored in the memory of the hub device 1000. Although Figure 10 the voice assistant server 2000 and the IoT server 3000 are not shown in
[0321] Reference Figure 10 , according to an embodiment, the hub device 1000 may include an ASR module 1310, an NLG module 1320, a device determination model 1330, a function determination device determination module 1340, a speaker function determination model 1352, and a TV function determination model 1354. In Figure 10 the embodiment of
[0322] According to an embodiment, in operation S1010, the hub device 1000 may select a function determination model corresponding to the operation execution device 4000a from the speaker function determination model 1352 and the TV function determination model 1354 previously stored in the hub device 1000. In the embodiment of the present disclosure, based on determining that the operation execution device 4000a is a TV, the hub device 1000 may select the TV function determination model 1354 corresponding to the TV from the speaker function determination model 1352 and the TV function determination model 1354.
[0323] In operation S1020, the hub device 1000 provides at least a part of the text to the selected function determination model. In the embodiment of the present disclosure, the processor 1200 of the hub device 1000 (see Figure 2)At least a part of the text can be provided to the TV function determination model 1354 stored in the memory. In this case, the processor 1200 can provide at least a part of the text to the TV function determination model 1354 without providing the entire text. For example, based on determining that the operation execution device 4000a is a "TV", in the text "Play the movie Avengers on the TV", "on the TV" specifies the name or common name of the operation execution device 4000a, so it can be unnecessary information. In addition, since the TV function determination model 1354 is a function determination model corresponding to the TV, "on the TV" is unnecessary information for the TV function determination model 1354. The processor 1200 can parse the text in units of words or phrases by using the first NLU model 1332, can identify words or phrases that specify the name, common name, or installation location of the device, and can provide the remaining part of the text except for the words or phrases identified in the entire text to the TV function determination model 1354.
[0324] In operation S1030, the hub device 1000 determines the intention by interpreting at least a part of the text by using the second NLU model of the selected function determination model. In an embodiment of the present disclosure, the processor 1200 of the hub device 1000 (see Figure 2 ) can analyze at least a part of the text by using the second NLU model 1354a of the TV function determination model 1354. The second NLU model 1354a, which is a model dedicated to a specific device, can be an AI model trained to obtain an intention related to the device, where the device corresponds to the operation execution device 4000a determined by the first NLU model 1332 and corresponds to at least a part of the text. In addition, the second NLU model 1354a can be a model trained to determine the operation of the device related to the user's intention by interpreting the text.
[0325] According to an embodiment, in an embodiment of the present disclosure, the processor 1200 can parse the text in units of morphemes, words, or phrases by using the second NLU model 1354a, can identify the meanings of the morphemes, words, or phrases parsed through syntactic and semantic analysis, and can determine the intention and parameters by matching the identified meanings with predefined words. The term "parameter" used herein refers to variable information for determining the detailed operation of the target device related to the intention. Descriptions of the same intention and parameters are not provided here in reference to Figure 7 the same intention and parameters.
[0326] In an embodiment of the present disclosure, the processor 1200 can determine the intention only from at least a part of the text by using the second NLU model 1354a.
[0327] In operation S1040, the hub device 1000 obtains operation information about an operation to be performed by the operation execution device 4000a based on the intention. In an embodiment of the present disclosure, the processor 1200 of the hub device 1000 (see Figure 2 ) can obtain operation information about at least one detailed operation related to the intention and parameters by using the action plan management module 1354b of the TV function determination model 1354. The action plan management module 1354b can manage information about the detailed operations of the operation execution device 4000a and the relationships between the detailed operations. The processor 1200 of the hub device 1000 can plan the detailed operations to be performed by the operation execution device 4000a and the execution order of the detailed operations based on the intention and parameters by using the action plan management module 1354b, and can obtain the operation information.
[0328] The operation information can be information related to the detailed operations to be performed by the operation execution device 4000a and the execution order of the detailed operations. The operation information can include information related to the detailed operations to be performed by the operation execution device 4000a, the relationships between each detailed operation and another detailed operation, and the execution order of the detailed operations. The operation information can include, but is not limited to, the functions to be performed by the operation execution device 4000a to perform a specific operation, the execution order of the functions, the input values required to execute the functions, and the output values output as the execution results of the functions.
[0329] According to an embodiment, in operation S1050, the hub device 1000 generates a control command based on the operation information. The control command refers to an instruction that can be read and executed by the operation execution device 4000a, so that the operation execution device 4000a performs the detailed operations included in the operation information.
[0330] Figure 10 Different from Figure 9 , the operations of determining the intention (S1030), obtaining the operation information (S1040), and generating the control command (S1050) are performed by the hub device 1000. Referring to Figure 9 and Figure 10 , the entity that performs the operations of determining the intention, obtaining the operation information, and generating the control command varies depending on whether the function determination model corresponding to the operation execution device 4000a is stored in the hub device 1000 or in the operation execution device 4000a.
[0331] According to an embodiment, in operation S1060, the hub device 1000 sends a control command to the operation execution device 4000a by using the identification information of the operation execution device 4000a. In the embodiments of the present disclosure, the hub device 1000 can identify the operation execution device 4000a by using the identification information of the operation execution device 4000a among multiple devices connected through a network and pre-registered using the same user account, and can send a control command to the identified operation execution device 4000a.
[0332] According to an embodiment, in operation S1070, the operation execution device 4000a performs an operation according to the received control command.
[0333] Figure 11 is a flowchart illustrating a method of operating a hub device 1000, a voice assistant server 2000, an IoT server 3000, a third-party IoT server 5000, and a third-party device 5100 according to an embodiment of the present disclosure.
[0334] Figure 11 is illustrated in Figure 7 the step of After that, it is a flowchart of the operations of entities in a multi-device system environment including a hub device 1000, a voice assistant server 2000, an IoT server 3000, a third-party IoT server 5000, and a third-party device 5100. Figure 7 the steps shown in represents a state where it is checked that the function determination model corresponding to the operation execution device 4000a is not stored in both the memory of the hub device 1000 and the memory of the operation execution device 4000a itself.
[0335] Referring to Figure 11 , the hub device 1000 may include an ASR module 1310, an NLG module 1320, a device determination model 1330, and a function determination device determination module 1340. In Figure 11 the embodiments of
[0336] The voice assistant server 2000 may store multiple function determination models 2342, 2344, and 2348. The voice assistant server 2000 may store the function determination model 2348 (referred to as a third-party function determination model).
[0337] A third-party device 5100 refers to a device manufactured by a manufacturer, an affiliate, or a company with a technology bundle other than the manufacturer of the hub device 1000. In an embodiment of the present disclosure, when the user of the third-party device 5100 is the same as the user of the hub device 1000, when the user registers the third-party device 5100 by using the user account for logging in to the hub device 1000, the third-party function determination model 2348 may be stored in the voice assistant server 2000. For example, when the user registers the third-party device 5100 by logging in with the user account of the hub device 1000, the third-party IoT server 5000 may permit access rights to the IoT server 3000, and the IoT server 3000 may access the third-party IoT server 5000 and may request the third-party function determination model 2348 for obtaining operation information about detailed operations according to the relationship between multiple functions of the third-party device 5100 and the detailed operations. The open authorization (oAuth) method may be used to register the third-party device 5100 to the IoT server 3000. The third-party IoT server 500 may provide the data of the third-party function determination model 2348 to the IoT server 3000 according to the request, and the IoT server 3000 may provide the third-party function determination model 2348 to the voice assistant server 2000.
[0338] The third-party IoT server 5000 is a server that stores and manages at least one of the device identification information of the third-party device 5100, the device type of the third-party device 5100, the function execution capability information, the location information, and the status information. The third-party IoT server 5000 may obtain, determine, or generate a control command for controlling the third-party device 5100 by using the device information of the third-party device 5100. The third-party IoT server 5000 may send a second control command to the third-party device 5100 that determines to execute an operation based on the operation information. The third-party IoT server 5000 may be operated by the manufacturer of the third-party device 5100, but is not limited thereto.
[0339] In Figure 11 the embodiment of, it may be determined that the operation execution device is the third-party device 5100.
[0340] According to an embodiment, in operation S1110, the hub device 1000 sends at least a part of the text and the identification information of the third-party device 5100 to the voice assistant server 2000. Except that the identification information of the third party 5100 - for example, the device id of the third-party device 5100 - is sent to the voice assistant server 2000, operation S1110 is the same as operation S810 of Figure 8 and thus its repeated description will not be provided.
[0341] According to an embodiment, in operation S1120, the voice assistant server 2000 selects a function determination model corresponding to the third-party device 5100 by using the received identification information of the third-party device 5100. In the embodiment of the present disclosure, the voice assistant server 2000 may identify the third-party device 5100 based on the identification information of the third-party device 5100 received from the hub device 1000, and may select the third-party function determination model 2348 corresponding to the third-party device 5100 from among the plurality of function determination models 2342, 2344, and 2348.
[0342] According to an embodiment, in operation S1130, the voice assistant server 2000 analyzes the text by using the NLU model 2348a of the third-party function determination model 2348 and determines the intent based on the analysis result. Except for using the NLU model 2348a of the third-party function determination model 2348, operation S1130 is the same as Figure 8 operation S830, and thus a repeated description thereof will not be provided here.
[0343] According to an embodiment, in operation S1140, the voice assistant server 2000 obtains operation information about an operation to be performed by the third-party device 5100 based on the intent. Except for obtaining information about the operation to be performed by the third-party device 5100, operation S1140 is the same as Figure 8 operation S840, and thus a repeated description thereof will not be provided here.
[0344] According to an embodiment, in operation S1142, the voice assistant server 2000 sends the identification information of the third-party device 5100 (e.g., the device id of the third-party device 5100) and the operation information to the IoT server 3000.
[0345] According to an embodiment, in operation S1150, the IoT server 3000 obtains a first control command based on the identification information of the third-party device 5100 and the received operation information. The IoT server 3000 may include a database storing operation information about the third-party device 5100 and control commands according to the operation information. The IoT server 3000 may obtain the operation information of the third-party device 5100 from the third-party IoT server 5000 through the registration method of the third-party device 5100, and may store the obtained operation information in the database. In the embodiment of the present disclosure, the IoT server 3000 may identify the third-party device 5100 based on the identification information of the third-party device 5100, and may generate a first control command by using the operation information of the third-party device 5100.
[0346] The first control command may include instructions for sequentially performing detailed operations according to operation information in an execution order when a third-party device 5100 performs a function. In an embodiment of the present disclosure, the first control command may not be read by the third-party device 5100.
[0347] According to an embodiment, in operation S1152, the IoT server 3000 sends identification information of the third-party device 5100 and the first control command to the third-party IoT server 5000.
[0348] According to an embodiment, in operation S1160, the third-party IoT server 5000 converts the first control command into a second control command that is readable and executable by the third-party device 5100. The third-party IoT server 5000 may convert the first control command into a second control command that is readable and executable by the third-party device 5100 by using the received identification information of the third-party device 5100.
[0349] According to an embodiment, in operation S1162, the third-party IoT server 5000 sends the second control command to the third-party device 5100 by using the identification information of the third-party device 5100.
[0350] According to an embodiment, in operation S1170, the third-party device 5100 performs an operation corresponding to the received second control command according to the received second control command.
[0351] Figure 12A is a conceptual diagram illustrating operations of a hub device 1000 and a plurality of devices (e.g., a first device 4100a and a second device 4200a) according to an embodiment of the present disclosure.
[0352] Figure 12A is a block diagram illustrating basic elements for describing operations of a hub device 1000 and a plurality of devices (e.g., a first device 4100a and a second device 4200a). However, the elements of the hub device 1000 and the plurality of devices (4100a and 4200a) are not limited to Figure 12A the elements of.
[0353] Figure 12A The arrow of indicates the movement, transmission, and reception of data including voice signals and text between the hub device 1000 and the first device 4100a. The circled numbers indicate the order of performing operations.
[0354] Reference Figure 12A, a hub device 1000 and multiple devices (e.g., a first device 4100a and a second device 4200a) can be connected to each other by using a wired communication method or a wireless communication method and can perform data communication. In an embodiment of the present disclosure, the hub device 1000 and multiple devices (e.g., a first device 4100a and a second device 4200a) can be directly connected to each other through a communication network, but the present disclosure is not limited thereto. In an embodiment of the present disclosure, although the hub device 1000 and multiple devices (e.g., a first device 4100a and a second device 4200a) can be connected to a voice assistant server 2000 (see Figure 3 ), the hub device 1000 can be connected to multiple devices (e.g., a first device 4100a and a second device 4200a) through the voice assistant server 2000.
[0355] The hub device 1000 and multiple devices (e.g., a first device 4100a and a second device 4200a) can be connected through a LAN, WAN, VAN, mobile radio communication network, satellite communication network, or a combination thereof. Examples of the wireless communication method can include but are not limited to Wi-Fi, Bluetooth, BLE, Zigbee, WFD, UWB, IrDA, and NFC.
[0356] The hub device 1000 is a device that receives a voice signal and controls at least one of multiple devices (e.g., a first device 4100a and a second device 4200a) based on the received voice signal.
[0357] Multiple devices (e.g., a first device 4100a and a second device 4200a) can be devices that log in using the same user account as the hub device 1000 and are previously registered with the IoT server 3000 using the user account of the hub device 1000.
[0358] At least one of multiple devices (e.g., a first device 4100a and a second device 4200a) can be a listening device that receives voice input from a user. In Figure 12A the embodiment, the first device 4100a can be a listening device that receives voice input including a user's speech from the user. The listening device can be, but is not limited to, a device that only receives voice input from the user. In an embodiment of the present disclosure, the listening device can be an operation execution device that receives a control command from the hub device 1000 and performs an operation for a specific function. In an embodiment of the present disclosure, the listening device can receive voice input related to the function performed by the listening device from the user. For example, the first device 4100a can be an air conditioner, and the first device 4100a can receive voice input such as "lower the air conditioner temperature to 20 °C" from the user through a microphone 4140.
[0359] The first device 4100a, as a listening device, can receive a voice signal from a voice input received from a user. In an embodiment of the present disclosure, the first device 4100a can convert the sound received through the microphone 4140 into an acoustic signal, and can obtain a voice signal by removing noise (e.g., non-speech components) from the acoustic signal.
[0360] The first device 4100a can send the voice signal to the hub device 1000 (step 1).
[0361] The hub device 1000 can receive the voice signal from the first device 4100a, and can convert the voice signal into text by performing ASR using the data of the ASR module 1310 previously stored in the memory (step 2).
[0362] The memory 1300 of the hub device 1000 (see Figure 2 ) can include a device determination model 1330 that detects an intention from the text by interpreting the text and determines a device that performs an operation corresponding to the detected intention. The device determination model 1330 can determine an operation execution device from among a plurality of devices registered according to a user account (e.g., the first device 4100a and the second device 4200a). The hub device 1000 can detect an intention from the text by interpreting the text using the data of the first NLU model 1332 included in the device determination model 1330. The first NLU model 1332 is a model for determining an intention by interpreting the text and determining an operation execution device based on the intention. The hub device 1000 can determine that the operation execution device that performs an operation corresponding to the intention is the first device 4100a by using the data of the device determination model 1330 (step 3).
[0363] A function determination model corresponding to the operation execution device determined by the hub device 1000 can be stored in the memory 1300 of the hub device 1000, can be stored in the first device 4100a itself determined as the operation execution device, or can be stored in the memory 2300 of the voice assistant server 2000 (see Figure 3 ). The function determination model corresponding to each device is a model for obtaining operation information regarding detailed operations performed according to the determined functions of the device and the relationships between the detailed operations.
[0364] The hub device 1000 itself can store a function determination model corresponding to at least one of multiple devices (e.g., the first device 4100a and the second device 4200a). For example, when the hub device 1000 is a voice assistant speaker, the hub device 1000 can store a speaker function determination model 1352, which is used to obtain operation information regarding detailed operations for performing the functions of the voice assistant speaker and the relationships between the detailed operations.
[0365] The hub device 1000 can also store a function determination model corresponding to another device. For example, the hub device 1000 can store a refrigerator function determination model 1356, which is used to obtain operation information regarding detailed operations corresponding to the refrigerator and the relationships between the detailed operations. The refrigerator can be a device that was previously registered to the IoT server 3000 using the same user account as the hub device 1000.
[0366] The speaker function determination model 1352 and the refrigerator function determination model 1356 can respectively include a second NLU model 1352a and 1356a and action plan management modules 1352b and 1356b. In Figure 12A an embodiment, the hub device 1000 includes a second NLU model 1352a and an action plan management module 1352b for the speaker, and includes a second NLU module 1356b and an action plan management module 1356b for the refrigerator. Except that the device type is changed from a TV to a refrigerator, the second NLU models 1352a and 1354a and the action plan management modules 1352b and 1354b included in Figure 2 the speaker function determination model 1352 and the TV function determination model 1354 are the same, and thus repeated descriptions will not be given.
[0367] The hub device 1000 can identify a device storing the function determination model 4132 by using the data of the function determination device determination module 1340, and the function determination model 4132 corresponds to the first device 4100a determined to be an operation execution device. In an embodiment of the present disclosure, the hub device 1000 can obtain information about the function determination model from each of a plurality of devices (e.g., the first device 4100a and the second device 4200a). The term "information about the function determination model" refers to information about whether each of a plurality of devices (e.g., the first device 4100a and the second device 4200a) itself stores a function determination model for obtaining operation information about detailed operations performed according to functions and the relationships between the detailed operations. In an embodiment of the present disclosure, when the hub device 1000 receives information about the function determination model from each of a plurality of devices (e.g., the first device 4100a and the second device 4200a), the hub device 1000 can also obtain information about the storage location of the function determination model of each of the plurality of devices (e.g., device identification information, IP address, or MAC address).
[0368] In another embodiment of the present disclosure, the hub device 1000 can identify a device among the hub device 1000, the first device 4100a, the second device 4200a, and the voice assistant server 2000 that stores a function determination model corresponding to the operation execution device by using the database 1360 (see Figure 2 ) including information about the function determination model of the device stored in the memory. In an embodiment of the present disclosure, the hub device 1000 can search the database 1360 according to the device identification information of the first device 4100a determined to be the operation execution device by using the data of the function determination device determination module 1340, and can obtain information about the storage location of the function determination model 4132 corresponding to the first device 4100a based on the search result of the database 1360.
[0369] The hub device 1000 can send at least a part of the text to the first device 4100a identified as storing the function determination model 4132 corresponding to the operation execution device by using the function determination device determination module 1340 (step 4).
[0370] The first device 4100a may receive at least a portion of the text from the hub device 1000 and may interpret at least a portion of the text by using a function determination model 4132 previously stored in the memory 4130. In an embodiment of the present disclosure, the first device 4100a may analyze at least a portion of the text by using an NLU model 4134 included in the function determination model 4132 and may obtain operation information regarding an operation to be performed by the first device 4100a based on the analysis result of at least a portion of the text. The function determination model 4132 may include an action plan management module 4136 configured to manage operation information related to detailed operations of the device so as to generate detailed operations to be performed by the first device 4100a and an execution order of the detailed operations. The action plan management module 4136 may manage operation information regarding the detailed operations of the first device 4100a and the relationships between the detailed operations. The action plan management module 4136 may plan the detailed operations to be performed by the first device 4100a and the execution order of the detailed operations based on the analysis result of at least a portion of the text. The first device 4100a may plan the detailed operations and the execution order of the detailed operations based on the analysis result of at least a portion of the text by the function determination model 4132 and may perform an operation based on the planning result (step 5).
[0371] However, the first device 4100a of the present disclosure is not limited to performing an operation based on at least a portion of the text received from the hub device 1000. In another embodiment of the present disclosure, the first device 4100a may receive a voice input from a user by using a microphone 4140, may convert the received voice input into text by using an ASR model 4138, and may perform an operation corresponding to the text by interpreting the text by using the function determination model 4132. That is, even without the participation of the hub device 1000, the first device 4100a may independently perform an operation.
[0372] Figure 12A An embodiment is illustrated in which the first device 4100a as a listening device for receiving a voice input from a user and an operation execution device for performing an operation related to the voice input through the hub device 1000 are the same. For example, when the first device 4100a receives a voice input from a user saying "lower the air conditioner temperature to 20 °C", the hub device 1000 may receive the user's voice input from the first device 4100a as a listening device and may determine the first device 4100a as an operation execution device related to the voice input. The first device 4100a may receive text related to the voice input from the hub device 1000 and may perform an air conditioner temperature adjustment operation based on the text.
[0373] And Figure 12ADifferently, the listening device and the operation execution device can be different from each other, which will be described with reference to Figure 12B hereinafter.
[0374] Figure 12B FIG. 6 is a conceptual diagram illustrating the operations of a hub device 1000 and multiple devices (e.g., a first device 4100a and a second device 4200a) according to an embodiment of the present disclosure.
[0375] Figure 12B FIG. 7 is a block diagram illustrating basic elements for describing the operations of the hub device 1000 and multiple devices (e.g., the first device 4100a and the second device 4200a). Elements identical to those of the hub device 1000 and the multiple devices (e.g., the first device 4100a and the second device 4200a) related to Figure 12A will not be described repeatedly.
[0376] With reference to Figure 12B , the second device 4200a can be a listening device that receives voice input from a user. In an embodiment of the present disclosure, although the second device 4200a can be a listening device, the second device 4200a may not be an operation execution device that receives a control command from the hub device 1000 or performs an operation for a specific function. In Figure 12B the embodiment of, the second device 4200a, which is a TV, can receive voice input unrelated to the TV, such as "lower the air conditioner temperature to 20°C", through a microphone 4240.
[0377] The second device 4200a can obtain a voice signal from the voice input received through the microphone 4240 and can send the voice signal to the hub device 1000 (step 1).
[0378] The hub device 1000 can receive the voice signal from the second device 4200a and can perform ASR by using data of an ASR module 1310 previously stored in a memory to convert the voice signal into text (step 2).
[0379] The hub device 1000 can determine an intention by interpreting the text by using a first NLU model 1332 and can determine an operation execution device related to the intention by using data of a device determination model 1330 (step 3). In Figure 12B the embodiment of, the hub device 1000 can determine, by using data of the device determination model 1330, that the operation execution device that performs an operation corresponding to the intention is the first device 4100a. The first device 4100a can be a device different from the listening device that receives voice input from a user. For example, the first device 4100a can be an air conditioner.
[0380] The hub device 1000 can identify a device in which a function determination model 4132 corresponding to the first device 4100a determined to be an operation execution device is stored by using data of the function determination device determination module 1340, and can send at least a part of the text to the identified first device 4100a (step 4).
[0381] The first device 4100a can receive at least a part of the text from the hub device 1000, can analyze at least a part of the text by using the function determination model 4132, can plan a detailed operation and an execution order of the detailed operation based on the analysis result, and can execute the operation based on the planned result (step 5).
[0382] Since the function determination model 4232 that can interpret information about the detailed operation related to the function of the second device 4200a and the execution order of the detailed operation is stored in the memory 4230 of the second device 4200a, the second device 4200a may not interpret the text related to the operation of the first device 4100a. In Figure 12B the embodiment, since the second device 4200a is a TV and the function determination model 4232 corresponding to the TV is stored in the memory 4230, the second device 4200a may not interpret a voice input saying "lower the air conditioner temperature to 20 °C". In this case, the second device 4200a as a listening device can send the voice input received from the user to the hub device 1000, and the hub device 1000 can send the text related to the operation to the first device 4100a determined to be the operation execution device. The first device 4100a can perform an air conditioner temperature adjustment operation based on the text received from the hub device 1000.
[0383] Reference Figure 12A and Figure 12BIn an embodiment, instead of sending a voice command related to an operation to be performed to a specific device that performs an operation related to a specific function, a user may send a voice command to any one of a plurality of devices (e.g., a first device 4100a and a second device 4200a) connected via a wired / wireless communication network by using a user account. For example, when a user generates a voice command saying "lower the air conditioner temperature to 20 °C", the user may send the voice command to an air conditioner that performs a function related to "cooling temperature setting", or may send a voice command related to cooling temperature adjustment to a TV that is not related to the air conditioner at all. A listening device that receives a voice input from the user among the plurality of devices (e.g., the first device 4100a and the second device 4200a) may send the voice input to the hub device 1000, and the hub device 1000 may determine an operation execution device corresponding to the user's utterance intention included in the voice input and may control the operation execution device to perform the operation.
[0384] In Figure 12A and Figure 12B In an embodiment, since a user does not need to directly specify a specific device related to a function among a plurality of devices (e.g., the first device 4100a and the second device 4200a) and send a voice command to the specific device, and an operation execution device for performing an operation related to the function is automatically determined, user convenience can be improved. In addition, since the hub device 1000 determines the operation execution device without involving an external server such as the voice assistant server 2000, it may not be necessary to use a network, and thus network usage fees can be reduced.
[0385] Figure 13 is a flowchart illustrating a method in which a hub device 1000 according to an embodiment of the present disclosure determines an operation execution device based on a voice signal received from a listening device and sends text to a device storing a function determination model corresponding to the operation execution device.
[0386] According to an embodiment, in operation S1310, the hub device 1000 receives a voice signal from a listening device. The listening device may be any one of a plurality of devices that are logged in by using the same user account as the user account of the hub device 1000 and are previously registered with the IoT server 3000 by using the same user account as the user account of the hub device 1000. The listening device may be connected to the hub device 1000 via a wired or wireless communication network. The listening device may be, but is not limited to, a device that only receives voice input from a user. In an embodiment of the present disclosure, the listening device may be an operation execution device that receives a control command from the hub device 1000 and performs an operation for a specific function.
[0387] In an embodiment of the present disclosure, the processor 1200 of the hub device 1000 (see Figure 2 ) may receive a voice signal from the listening device through the communication interface 1400 (see Figure 2 ).
[0388] According to an embodiment, in operation S1320, the hub device 1000 converts the received voice signal into text by performing ASR. In an embodiment of the present disclosure, the hub device 1000 may perform ASR by using a predefined model such as AM or LM to convert the voice signal received from the listening device into computer-readable text. When the hub device 1000 receives an acoustic signal without noise removal from the listening device, the hub device 1000 may obtain a voice signal by removing noise from the received acoustic signal and perform ASR on the voice signal.
[0389] According to an embodiment, in operation S1330, the hub device 1000 analyzes the text by using a first NLU model and determines an operation execution device corresponding to the analyzed text by using a device determination model. Except that the listening device is also included in the device candidates that can be determined as the operation execution device, operation S1330 is the same as Figure 6 operation S630 of, so repeated description will not be given.
[0390] According to an embodiment, in operation S1340, the hub device 1000 identifies a device that stores a function determination model corresponding to the operation execution device determined from among the listening device and the operation execution device. The function determination model corresponding to the operation execution device determined by the hub device 1000 may be stored in the memory 1300 of the hub device 1000 (see Figure 2 ), may be stored in the internal memory of the operation execution device itself, or may be stored in the memory 2300 of the voice assistant server 2000 (see Figure 3 ). When the determined operation execution device and the listening device are different, the function determination model corresponding to the operation execution device may be stored in the listening device. The function determination model corresponding to the operation execution device is a model used by the device determined as the operation execution device to obtain operation information about the detailed operations to be performed according to the determined functions of the operation execution device and the relationships between the detailed operations.
[0391] In an embodiment of the present disclosure, the processor 1200 of the hub device 1000 (see Figure 2 ) may, by using the function determination device determination module 1340 (see Figure 2)The program code or data of ) is used to identify a device that stores and operates a function determination model corresponding to an operation execution device. In an embodiment of the present disclosure, the hub device 1000 may obtain storage information of the function determination model from each of a plurality of devices connected to the hub device 1000 through a wired or wireless communication network and logged in using the same user account as the hub device 1000. The term "storage information of the function determination model" refers to information on whether each of the plurality of devices stores the function determination model itself, where the function determination model is used to obtain operation information on detailed operations performed according to functions and the relationships between the detailed operations. In an embodiment of the present disclosure, when the hub device 1000 receives information on the function determination model from a plurality of devices, the hub device 1000 may also obtain information on the storage location of the function determination model of each of the plurality of devices (e.g., device identification information, IP address, or MAC address). When the operation execution device and the listening device are different, the hub device 1000 may receive information on the function determination model from the listening device and may determine whether the function determination model corresponding to the operation execution device is stored in the internal memory of the listening device based on the received information on the function determination model.
[0392] In another embodiment of the present disclosure, the hub device 1000 may identify, by using the database 1360 (see Figure 2 ), a device among the hub device 1000, the operation execution device, and the voice assistant server 2000 that stores a function determination model corresponding to the operation execution device. The database 1360 includes information on the function determination models of the devices stored in the memory 1300 (see Figure 2 ). In an embodiment of the present disclosure, the hub device 1000 may search the database 1360 according to the device identification information of the operation execution device by using the program code or data of the function determination device determination module 1340, and may obtain information on the storage location of the function determination model corresponding to the operation execution device based on the search result of the database 1360. When the operation execution device and the listening device are different, the hub device 1000 may search the database 1360 for the identification information of the listening device and may determine whether the function determination model corresponding to the operation execution device is stored in the internal memory of the listening device based on the search result of the database 1360.
[0393] Reference will be made to Figure 14 to describe operation S1340.
[0394] According to an embodiment, in operation S1350, the hub device 1000 sends at least a part of the text to the identified device. The hub device 1000 may use the communication interface 1400 (see Figure 2)Send at least a portion of the text to the identified device. For example, based on determining that the function determination model corresponding to the operation execution device is stored in the internal memory of the operation execution device, the hub device 1000 can send at least a portion of the text to the operation execution device by using the communication interface 1400. When the listening device and the operation execution device are the same, the hub device 1000 can send at least a portion of the text to the listening device. As another example, based on determining that the function determination model corresponding to the operation execution device is not stored in the listening device and the operation execution device, but is stored in the memory 2300 of the voice assistant server 2000 (see Figure 3 ) therein, the hub device 1000 can send at least a portion of the text to the voice assistant server 2000 by using the communication interface 1400.
[0395] In an embodiment of the present disclosure, the hub device 1000 can control the communication interface to separate the part regarding the name of the operation execution device from the text and send only the remaining part of the text to the identified device. For example, when the device identified as storing the function determination model corresponding to the operation execution device is "air conditioner" and the text is "Lower the set temperature to 20°C in the air conditioner", when the text is sent to the air conditioner, it is not necessary to send "in the air conditioner". In this case, the hub device 1000 can parse the text in units of words or phrases, can identify the words or phrases specifying the name, common name, installation location, etc. of the device, and can provide the remaining part of the text except for the words or phrases identified in the text to the device (e.g., the air conditioner).
[0396] Figure 14 is a flowchart illustrating a method in which the hub device 1000 according to an embodiment of the present disclosure sends text to a device storing a function determination model corresponding to an operation execution device. Figure 14 Illustrates Figure 13 Operations S1340 and S1350 of Figure 14 Operations S1410 to S1450 of Figure 13 are detailed embodiments of operation S1340 of Figure 13 and operations S1460 to S1480 are detailed embodiments of operation S1350 of
[0397] After performing Figure 13 operation S1330 of Figure 13)The device identification information of the operation execution device (e.g., device ID information) and the device identification information of the listening device determined in [reference], and then it is possible to determine whether the listening device and the operation execution device are the same by comparing the obtained device identification information of the operation execution device with the obtained device identification information of the listening device.
[0398] Based on determining that the listening device and the operation execution device are the same (yes) in operation S1410, the hub device 1000 can obtain the storage information of the function determination model from the listening device (S1420). In an embodiment of the present disclosure, the hub device 1000 can receive the storage information of the function determination model from the listening device by using the communication interface 1400 (see Figure 2 )The storage information of the function determination model refers to information on whether the function determination model is stored in the internal memory of the listening device itself, and the function determination model is used to obtain operation information on detailed operations for performing operations according to functions and the relationships between the detailed operations. In an embodiment of the present disclosure, when the hub device 1000 receives information on the function determination model from the listening device, the hub device 1000 can also obtain information on the storage location of the function determination model (e.g., identification information, IP address, or MAC address of the listening device).
[0399] In operation S1430, the hub device 1000 determines whether the function determination model corresponding to the determined operation execution device is stored in the internal memory of the listening device based on the storage information of the function determination model of the listening device. In an embodiment of the present disclosure, the processor 1200 of the hub device 1000 (see Figure 2 ) can determine whether the function determination model corresponding to the operation execution device is stored in the internal memory of the listening device by using the program code or data of the function determination device determination module 1340 (see Figure 2 )In an embodiment of the present disclosure, the processor 1200 can determine whether the function determination model stored in the internal memory of the listening device is the same as the function determination model corresponding to the operation execution device based on the obtained storage information of the function determination model of the listening device.
[0400] In another embodiment of the present disclosure, the processor 1200 can use the database 1360 that stores information on function determination models of multiple devices (see Figure 2) to determine whether a function determination model corresponding to the operation execution device is stored in the internal memory of the listening device. In an embodiment of the present disclosure, the processor 1200 may search the database 1360 according to the device identification information of the listening device by using the program code or data of the function determination device determination module 1340, and may obtain information on whether the function determination model corresponding to the operation execution device is stored in the listening device based on the search result of the database 1360.
[0401] Based on determining in operation S1430 that the function determination model corresponding to the operation execution device is stored in the internal memory of the listening device (yes), the hub device 1000 sends at least a part of the text to the listening device (S1460). In an embodiment of the present disclosure, the processor 1200 may send at least a part of the text to the function determination model previously stored in the internal memory of the listening device by using the communication interface 1400.
[0402] Based on determining in operation S1430 that the function determination model corresponding to the operation execution device is not stored in the internal memory of the listening device (no), the hub device 1000 sends at least a part of the text to the voice assistant server 2000 (S1480). In an embodiment of the present disclosure, the processor 1200 may send at least a part of the text to the function determination model corresponding to the operation execution device among the multiple function determination models 2342, 2344, 2346, and 2348 previously stored in the memory 1300 of the voice assistant server 2000 (see Figure 3 ) by using the communication interface 1400.
[0403] In operations S1460 and S1480, the processor 1200 may send only at least a part of the text, without sending the entire text, which is the same as described in reference Figure 13 and thus will not be repeated here.
[0404] Based on determining in operation S1410 that the listening device and the operation execution device are different from each other (no), the hub device 1000 obtains the storage information of the function determination model from the operation execution device (S1440). In an embodiment of the present disclosure, the hub device 1000 may receive the storage information of the function determination model from the operation execution device by using the communication interface 1400 (see Figure 2 ). In an embodiment of the present disclosure, when the hub device 1000 receives information about the function determination model from the operation execution device, the hub device 1000 may also obtain information on the storage location of the function determination model (e.g., the identification information, IP address, or MAC address of the operation execution device).
[0405] In operation S1450, the hub device 1000 determines whether a function determination model is stored in the internal memory of the operation execution device. In an embodiment of the present disclosure, the processor 1200 may determine whether a function determination model corresponding to the operation execution device is stored in the internal memory of the operation execution device by using the program code or data of the function determination device determination module 1340. In another embodiment of the present disclosure, the processor 1200 may determine whether a function determination model corresponding to the operation execution device is stored in the internal memory of the operation execution device by using the database 1360 that stores information about function determination models of multiple devices. In an embodiment of the present disclosure, the processor 1200 may search the database 1360 according to the device identification information of the operation execution device by using the program code or data of the function determination device determination module 1340, and may obtain information about whether the function determination model corresponding to the operation execution device is stored in the internal memory of the operation execution device itself based on the search result of the database 1360.
[0406] Based on determining in operation S1450 that the function determination model corresponding to the operation execution device is stored in the internal memory of the operation execution device (yes), the hub device 1000 sends at least a part of the text to the operation execution device (operation S1470). In an embodiment of the present disclosure, the processor 1200 may send at least a part of the text to the function determination model previously stored in the internal memory of the operation execution device by using the communication interface 1400.
[0407] Based on determining in operation S1450 that the function determination model corresponding to the operation execution device is not stored in the operation execution device (no), the hub device 1000 sends at least a part of the text to the voice assistant server 2000 (S1480).
[0408] Figure 15 FIG. is a flowchart illustrating a method of operating the hub device 1000, the voice assistant server 2000, and the listening device 4000b according to an embodiment of the present disclosure. Figure 15 It is illustrated that the listening device 4000b is determined as the operation execution device by the hub device 1000 via natural language interpretation.
[0409] Reference Figure 15 , the hub device 1000 may include an ASR module 1310, an NLG module 1320, a device determination model 1330, a function determination device determination module 1340, and multiple function determination models (e.g., 1352, 1354, and 1356). However, the present disclosure is not limited thereto, and the hub device 1000 may not store a function determination model, or may not store a function determination model corresponding to the operation device.
[0410] The voice assistant server 2000 can store multiple function determination models 2342, 2344, and 2346. For example, the function determination model 2342, which is the first function determination model stored in the voice assistant server 2000, can be a model for determining the functions of an air conditioner and obtaining operation information regarding detailed operations related to the determined functions and the relationships between the detailed operations. For example, the function determination model 2344, which is the second function determination model, can be a model for determining the functions of a TV and obtaining operation information regarding detailed operations related to the determined functions and the relationships between the detailed operations, and the function determination model 2346, which is the third function determination model, can be a model for determining the functions of a washing machine and obtaining operation information regarding detailed operations related to the determined functions and the relationships between the detailed operations.
[0411] The listening device 4000b is a device that receives voice input from a user. However, the present disclosure is not limited thereto, and in an embodiment of the present disclosure, the listening device 4000b can be an operation execution device that receives a control command from the hub device 1000 and performs operations for a specific function. In Figure 15 the embodiment, the listening device 4000b can include a function determination model 4032b, an ASR module 4038b, and a microphone 4040b. The function determination model 4032b can include an NLU model 4034b and an action plan management module 4036b.
[0412] In operation S1510, the listening device 4000b receives voice input from a user. In an embodiment of the present disclosure, the listening device 4000b can receive voice input related to the functions performed by the listening device 4000b from the user through the microphone 4040b. For example, when the listening device 4000b is an air conditioner, the listening device 4000b can receive voice input such as "lower the air conditioner temperature to 20°C" from the user through the microphone 4040b.
[0413] In an embodiment of the present disclosure, the listening device 4000b can obtain a voice signal from the voice input received from the user. In an embodiment of the present disclosure, the listening device 4000b can convert the sound received through the microphone 4040b into an acoustic signal and can obtain a voice signal by removing noise (e.g., non-speech components) from the acoustic signal.
[0414] In operation S1520, the listening device 4000b determines whether a device determination model is stored in the memory of the listening device 4000b. In an embodiment of the present disclosure, the listening device 4000b may determine whether program code or data corresponding to the device determination model is stored by scanning the memory. However, the present disclosure is not limited thereto, and the listening device 4000b may obtain device specification information based on device identification information (e.g., device id information), and may determine whether the device determination model is stored in the internal memory of the listening device 4000b by using the device specification information.
[0415] Based on determining that the device determination model is not stored in the memory of the listening device 4000b (No), the listening device 4000b sends the voice signal to the hub device 1000 (S1522).
[0416] Based on determining that the device determination mode is stored in the memory of the listening device 4000b (Yes), the listening device 4000b converts the voice signal into text by performing ASR (S1530). The listening device 4000b may convert the voice signal into computer-readable text by performing ASR using the ASR module 4038b. When the listening device 4000b receives an acoustic signal without noise removal, the listening device 4000b may obtain a voice signal by removing noise from the received acoustic signal, and may perform ASR on the voice signal.
[0417] According to an embodiment, in operation S1540, the listening device 4000b determines the listening device 4000b as an operation execution device by using the device determination model to interpret the text. In an embodiment of the present disclosure, the device determination model included in the listening device 4000b may include an NLU model, and the listening device 4000b may determine the listening device 4000b as an operation execution device by using the NLU model to interpret the text. For example, when the text converted from the voice signal is "lower the air conditioner temperature to 20 °C", the listening device 4000b may determine the listening device 4000b as an air conditioner and as an operation execution device by using the NLU model to interpret the text.
[0418] In operation S1550, the listening device 4000b provides text to the function determination model 4032. For example, the function determination model 4032, which is the function determination model for an air conditioner, can be a model for obtaining operation information regarding detailed operations to be performed according to the functions of the air conditioner and the relationships between the detailed operations. The function determination model 4032 can include an NLU model 4034b, which is configured to obtain operation information related to the operations to be performed by the listening device 4000b based on the analysis results of at least a part of the text. The function determination model 4032b can include an action plan management module 4036b, which manages operation information related to the detailed operations of the device to generate the detailed operations to be performed by the listening device 4000b and the execution order of the detailed operations. The action plan management module 4036b can plan the detailed operations to be performed by the listening device 4000b and the execution order of the detailed operations based on the analysis results of at least a part of the text.
[0419] In an embodiment of the present disclosure, the listening device 4000b can provide text representing "lower the air conditioner temperature to 20 °C" to the function determination model 4032b for the air conditioner, can obtain operation information to be performed by the air conditioner, such as information about the operation of "lowering the set temperature", by interpreting the text using the NLU model 4034b, and can plan the detailed operations for performing the operation of "lowering the set temperature" and the execution order of the detailed operations by using the action plan management module 4036b.
[0420] In operation S1522, the hub device 1000 receives a voice signal from the listening device 4000b.
[0421] According to an embodiment, in operation S1560, the hub device 1000 converts the voice signal into text by performing ASR. The detailed method for the hub device 1000 to perform ASR is the same as the method described in Figure 6 operation S620, and thus no detailed description will be given.
[0422] According to an embodiment, in operation S1570, the hub device 1000 interprets the text by using the first NLU model 1332 and determines the listening device 4000b as the operation determination device related to the text by using the device determination model 1330. For example, the hub device 1000 can interpret the text saying "lower the air conditioner temperature to 20 °C" by using the first NLU model 1332 to obtain the intention corresponding to "lowering the set temperature of the air conditioner", and can determine the listening device 4000b, which is an air conditioner, as the operation execution device based on the obtained intention by using the device determination model 1330. The detailed method for the hub device 1000 to determine the operation execution device is the same as Figure 6The method described in operation S630 is the same and will not be described in detail.
[0423] According to an embodiment, in operation S1580, the hub device 1000 determines whether a function determination model is stored in the listening device 4000b. In the embodiment of the present disclosure, the processor 1200 of the hub device 1000 (see Figure 2 ) determines whether the function determination model corresponding to the listening device 4000b is stored in the internal memory of the listening device 4000b based on the storage information of the function determination model obtained from the listening device 4000b. In another embodiment of the present disclosure, the processor 1200 may determine whether the function determination model corresponding to the listening device 4000b is stored in the internal memory of the listening device 4000b by using the database 1360 (see Figure 2 ) that stores information about the function determination models of multiple devices. The detailed method for determining whether the function determination model corresponding to the listening device 4000b is stored in the internal memory of the listening device 4000b is the same as Figure 14 the method described in operation S1430 and thus repeated description will not be given.
[0424] Based on determining that the function determination model is stored in the listening device 4000b (yes) in operation S1580, the hub device 1000 sends the text to the function determination model of the listening device 4000b (S1592).
[0425] Based on determining that the function determination model is not stored in the listening device 4000b (no) in operation S1580, the hub device 1000 sends the text to the function determination model corresponding to the listening device 4000b of the voice assistant server 2000 (S1594). When the listening device 4000b is an air conditioner and it is determined that the function determination model for the air conditioner is not stored in the internal memory of the listening device 4000b, the hub device 1000 may send the text to the function determination model 2342, which is the function determination model for the air conditioner among the multiple function determination models 2342, 2344, m, and 2346 previously stored in the memory 2300 of the voice assistant server 2000 (see Figure 3 ).
[0426] In an embodiment of the present disclosure, operations S1530 to S1550 performed by the listening device 4000b and operations S1560 to S1594 performed by the hub device 1000 may be performed in parallel by separate entities. Additionally, operations S1530 to S1550 performed by the listening device 4000b may be performed independently and are not related to operations S1560 to S1594 performed by the hub device 1000. That is, although the listening device 4000b may be determined by the hub device 1000 as an operation execution device and may receive at least a part of the text, the present disclosure is not limited thereto. In an embodiment of the present disclosure, the listening device 4000b may determine the operation execution device by interpreting the user's voice input and may provide at least a part of the text to the function determination model 4032b.
[0427] Figure 16 FIG. is a flowchart illustrating a method of operating a hub device 1000, a voice assistant server 2000, an operation execution device 4000a, and a listening device 4000b according to an embodiment of the present disclosure. Figure 16 It is illustrated that the listening device 4000b and the operation execution device 4000a are different from each other. For example, the listening device 4000b may be a TV, and the operation execution device 4000a may be an air conditioner.
[0428] Reference Figure 16 , the hub device 1000 may include an ASR module 1310, an NLG module 1320, a device determination model 1330, a function determination device determination module 1340, and a plurality of function determination models (e.g., 1352, 1354, and 1356).
[0429] The voice assistant server 2000 may store a plurality of function determination models 2342, 2344, and 2346.
[0430] Figure 16 The hub device 1000 and the voice assistant server 2000 of Figure 15 are the same as the hub device 1000 and the voice assistant server 2000 of
[0431] are the same as Figure 15 are different from Figure 16 The listening device 4000b of
[0432] The operation execution device 4000a itself can store the function determination model 4032. The operation execution device 4000a itself can analyze text, and the function determination model 4032 for performing operations based on the analysis result of the text according to the user's intention can be stored in the memory of the operation execution device 4000a. For example, when the operation execution device 4000a is an "air conditioner", the function determination model 4032 stored in the memory of the operation execution device 4000a can be a model for determining the functions of the air conditioner and obtaining operation information on the determined functions and the relationships between the detailed operations.
[0433] According to an embodiment, in operation S1610, the listening device 4000b receives a voice input from the user. Operation S1610 is the same as Figure 15 operation S1510, so repeated description will not be given.
[0434] According to an embodiment, in operation S1620, the listening device 4000b determines, from among a plurality of devices previously registered using the user account, the device including the device determination model 1330 as the hub device 1000. In an embodiment of the present disclosure, the listening device 4000b can obtain device determination model information from a plurality of devices connected via a wired or wireless communication network and logged in using the same user account, and can determine, based on the obtained device determination model information, the device storing the device determination model 1330 as the hub device 1000. The device determination model information can include at least one of information on whether there is a device determination model for each of the plurality of devices, identification information of the device / server in which the device determination model is stored, and information on the IP address or MAC address of the stored device / server.
[0435] In operation S1630, the listening device 4000b sends the voice signal to the hub device 1000.
[0436] According to an embodiment, in operation S1640, the hub device 1000 converts the voice signal into text by performing ASR.
[0437] According to an embodiment, in operation S1650, the hub device 1000 interprets the text by using the first NLU model 1332 and determines the operation execution device related to the text by using the device determination model 1330.
[0438] Operations S1640 and S1650 are the same as Figure 6 operations S620 and S630, respectively, so repeated description will not be given.
[0439] According to an embodiment, in operation S1660, the hub device 1000 determines whether a function determination model is stored in the operation execution device. In an embodiment of the present disclosure, the processor 1200 of the hub device 1000 (see Figure 2 ) determines whether the function determination model 4032 corresponding to the operation execution device 4000a is stored in the internal memory of the operation execution device 4000a based on the storage information of the function determination model 4032 obtained from the operation execution device 4000a. In another embodiment of the present disclosure, the processor 1200 may determine whether the function determination model corresponding to the operation execution device 4000a is stored in the internal memory of the operation execution device 4000a by using the database 1360 (see Figure 2 ) that stores information about the function determination models of multiple devices. The detailed method for determining whether the function determination model corresponding to the operation execution device 4000a is stored in the internal memory of the operation execution device 4000a is the same as the method described in Figure 14 's operation S1450, so repeated description will not be given.
[0440] Based on determining that the function determination model 4032 is stored in the operation execution device 4000a (yes) in operation S1660, the hub device 1000 sends text to the operation execution device 4000a (S1672). For example, when the operation execution device 4000a is an air conditioner and it is determined that the function determination model 4032 stored in the memory of the operation execution device 4000a is a model for obtaining operation information of the air conditioner, the hub device 1000 may send at least a part of the text to the function determination model 4032 stored in the operation execution device 4000a.
[0441] Based on determining that the function determination model is not stored in the operation execution device 4000a (no) in operation S1660, the hub device 1000 sends text to the voice assistant server 2000 (S1674). For example, when the operation execution device 4000a is an air conditioner but it is determined that the function determination model 4032 stored in the memory of the operation execution device 4000a is a model for obtaining operation information of a TV, the hub device 1000 may send at least a part of the text to the function determination model 2342 for obtaining operation information of the air conditioner among the multiple function determination models 2342, 2344, and 2346 stored in the voice assistant server 2000.
[0442] Figure 17 FIG. is a diagram illustrating an example in which the operation execution device 4000a updates a function determination model according to an embodiment of the present disclosure.
[0443] Refer to Figure 17, the voice assistant server 2000 may include a device determination model 2330 and multiple function determination models 2342, 2344, and 2346. The operation execution device 4000a may include a function determination model 4032, and the function determination model 4032 may include an NLU model 4034 and an action plan management module 4036.
[0444] The device determination model 2330 included in the voice assistant server 2000 may be a model configured to determine the latest version of the operation execution device 4000a by interpreting a voice input for the latest function or the latest device. The multiple function determination models 2342, 2344, and 2346 included in the voice assistant server 2000 may be the latest version of models configured to interpret voice inputs related to the latest functions of, for example, an air conditioner, a refrigerator, and a TV and generate operation information as an analysis result.
[0445] The function determination model 4032 included in the operation execution device 4000a is configured to interpret text converted from a voice input and generate operation information for performing a specific function. The function determination model 4032 may include an NLU model 4034 trained to interpret only text regarding a limited number of functions and an action plan management module 4036 trained to generate operation information from the text interpreted by the NLU model 4034. In an embodiment of the present disclosure, the function determination model 4032 included in the operation execution device 4000a may not be the latest version, and thus may not be able to interpret text regarding the latest function and may not be able to generate operation information regarding the latest function. When receiving text regarding the latest function, since the function determination model 4032 may not be able to interpret the text and may not be able to determine the function even when using the NLU model 4034, the function determination model 4032 may not be able to generate operation information even when using the action plan management module 4036. In this case, the operation execution device 4000a cannot determine the function.
[0446] According to an embodiment, in operation S1710, the operation execution device 4000a sends device information of the function determination model and update request information to the voice assistant server 2000. The update request information of the function determination model may be information for requesting synchronization of the version of the function determination model 4032 stored in the memory of the operation execution device 4000a with the version of the function determination model corresponding to the operation execution device 4000a among the multiple function determination models 2342, 2344, and 2346 stored in the voice assistant server 2000.
[0447] In an embodiment of the present disclosure, the operation execution device 4000a may periodically send update request information of the function determination model to the voice assistant server 2000 at a specific time interval, specific date, etc. However, the present disclosure is not limited thereto, and when updating an application or firmware, the operation execution device 4000a may send update request information of the function determination model to the voice assistant server 2000. In another embodiment of the present disclosure, when the operation execution device 4000a receives a control command from the IoT server 3000, the operation execution device 4000a may send update request information of the function determination model to the voice assistant server 2000.
[0448] According to an embodiment, in operation S1720, the voice assistant server 2000 sends update data of the function determination model.
[0449] According to an embodiment, in operation S1730, the operation execution device 4000a updates the function determination model to the latest version by using the update data. In an embodiment of the present disclosure, the operation execution device 4000a may overwrite and update the previously stored function determination model 4032 by using data of the latest version of the function determination model received from the voice assistant server 2000. The operation execution device 4000a may perform training by updating the function determination model 4032 to interpret text of the latest function regarding addition, modification, or deletion, and generate operation information according to the interpretation result.
[0450] According to an embodiment, in operation S1740, the voice assistant server 2000 updates the device determination model 2330. In an embodiment of the present disclosure, when the voice assistant server 2000 interprets text related to the latest function and determines a device for performing an operation on the latest function, the operation execution device 4000a may be updated such that the device determination model 2330 is also included in the device candidates.
[0451] According Figure 17 to an embodiment, since the function determination model 4032 stored in the internal memory of the operation execution device 4000a itself is synchronized with the latest function determination model of the voice assistant server 2000, when the user speaks to perform the latest function, even without accessing the voice assistant server 2000 through a communication network, the operation can be performed by the function determination model 4032 of the operation execution device 4000a, thereby reducing network usage fees and improving server operation efficiency.
[0452] Figure 18 is a flowchart illustrating a method of an operation hub device 1000, a voice assistant server 2000, an IoT server 3000, and an operation execution device 4000a according to an embodiment of the present disclosure. Figure 18The figure shows that when a control command is received from the IoT server 3000, the operation execution device 4000a updates the function determination model 4032 to the latest version.
[0453] Reference Figure 18 , the hub device 1000 may include an ASR module 1310, an NLG module 1320, a device determination model 1330, a function determination device determination module 1340, and multiple function determination models (e.g., 1352, 1354, and 1356).
[0454] The operation execution device 4000a may include a function determination model 4032, and the function determination model 4032 may include an NLU model 4034 and an action plan management module 4036.
[0455] Figure 18 The hub device 1000 and the operation execution device 4000a of Figure 16 are the same as the hub device 1000 and the operation execution device 4000a of
[0456] and thus will not be described repeatedly. Figure 3 The voice assistant server 2000 may include a device determination model 2330 and multiple function determination models 2342, 2344, and 2346. The device determination model 2330 is the same as
[0457] the device determination model 2330 of
[0458] and thus will not be described repeatedly.
[0459] In operation S1810, the IoT server 3000 sends a control command to the operation execution device 4000a.
[0460] Based on determining that the operation information corresponding to the control command is not stored in the function determination model 4032 in operation S1820 (No), the operation execution device 4000a sends update request information of the function determination model 4032 and device information to the voice assistant server 2000 (S1840). The device information may include at least one of device identification information (e.g., device id information) of the operation execution device 4000a, storage information of the function determination model 4032 of the operation execution device 4000a, and version information of the function determination model 4032. The version information of the function determination model 4032 may include information about the attributes, types, or quantities of functions that can be recognized by using the function determination model 4032.
[0461] In operation S1842, the voice assistant server 2000 sends update data of the function determination model to the operation execution device 4000a.
[0462] In operation S1850, the operation execution device 4000a updates the function determination model 4032 to the latest version by using the update data. Operation S1850 is the same as Figure 17 operation S1730, so repeated descriptions will not be given.
[0463] In operation S1860, the operation execution device 4000a sends update information of the function determination model 4032 to the voice assistant server 2000. In an embodiment of the present disclosure, the operation execution device 4000a may send at least one of information about the version of the function determination model 4032 after the update and information about the attributes, types, or quantities of functions that can be recognized by using the function determination model 4032 after the update to the voice assistant server 2000.
[0464] In operation S1870, the voice assistant server 2000 updates the device determination model 2330 based on the update information of the function determination model of the operation execution device. In an embodiment of the present disclosure, the voice assistant server 2000 may update the device determination model 2330 based on at least one of the updated version information of the function determination model 4032 and information about the attributes, types, and quantities of functions that can be recognized by using the updated function determination model 4032 received from the operation execution device 4000a, so that the operation execution device 4000a is included in the device candidate list for determining the device for executing the latest function.
[0465] In operation S1880, the voice assistant server 2000 sends update data of the device determination model 2330 to the hub device 1000.
[0466] In operation S1890, the hub device 1000 updates the device determination model 1330 based on the received updated data. The device determination model 1330 can be synchronized with the device determination model 2330 of the voice assistant server 2000 by being updated. Therefore, when a voice input of a user related to the latest function is received through the hub device 1000, since the hub device 1000 itself can determine the operation execution device 4000a by using the device determination model 1330 stored in the internal memory 1300 (see Figure 2 ), without accessing the voice assistant server 2000 through the communication network, network usage fees can be reduced, processing time can be reduced, and thus the response speed can be improved.
[0467] Figure 19 FIG. is a flowchart illustrating a method of operating the hub device 1000, the voice assistant server 2000, the IoT server 3000, and the new device 4000c.
[0468] Referring to Figure 19 , the hub device 1000 may include an ASR module 1310, an NLG module 1320, a device determination model 1330, and a function determination device determination module 1340.
[0469] The voice assistant server 2000 may include a device determination model 2330 and multiple function determination models 2342, 2344, and 2346. The device determination model 2330 is the same as the Figure 3 device determination model 2330, and thus duplicate descriptions will not be given.
[0470] The new device 4000c is a target device that logs in using the same user account as that of the hub device 1000 and is expected to be registered in the user account of the voice assistant server 2000. The new device 4000c can be connected to the hub device 1000, the voice assistant server 2000, and the IoT server 3000 through a wired or wireless communication network.
[0471] In the Figure 19 embodiment, the new device 4000c itself can analyze text, and a function determination model 4032c for performing an operation based on the analysis result of the text according to the user's intention can be stored in the memory of the new device 4000c. For example, when the new device 4000c is an "air purifier", the function determination model 4032c stored in the memory of the new device 4000c can be a model for determining the functions of the air conditioner and obtaining operation information on the determined functions and the relationships between the detailed operations. The function determination model 4032c may include an NLU model 4034c and an action plan management module 4036c.
[0472] However, the present disclosure is not limited thereto, and the new device 4000c may not include the function determination model 4032c. In another embodiment of the present disclosure, the new device 4000c may include a device determination model.
[0473] In operation S1910, the new device 4000c obtains user account information through login. The user account information includes a user ID and a password. The obtained user account information may be the same as the user account information of the hub device 1000.
[0474] In operation S1920, the hub device 1000 provides the device identification information and the determination capability information of the device determination model 1330 to the IoT server 3000. The term "device determination capability information" refers to information about the capability of the hub device 1000 to interpret text by using the first NLU model 1332 included in the device determination model 1330 and to determine a device for performing an operation based on the interpretation result of the text by using the device determination model 1330. The capability information may include information about device candidates used by the device determination model 1330 to determine an operation execution device. For example, the device determination model 1330 of the hub device 1000 may be a model trained to only interpret text related to an air conditioner, a TV, and a refrigerator and to determine only one of the air conditioner, the TV, and the refrigerator as an operation execution device based on the interpretation result. In this case, the candidate devices may be "air conditioner, TV, and refrigerator".
[0475] In operation S1930, the new device 4000c provides the user account information, the identification information of the new device 4000c (e.g., device ID information), and the storage information of the device determination model and the function determination model to the IoT server 3000. The term "storage information of the device determination model" refers to information about whether the new device 4000c itself stores a device determination model for determining an operation execution device in an internal memory. The term "storage information of the function determination model" refers to information about whether the new device 4000c stores a function determination model for obtaining operation information about detailed operations for performing an operation according to a function and the relationship between the detailed operations in an internal memory. The storage information of the function determination model may include information about whether not only the function determination model corresponding to the new device 4000c but also the function determination models corresponding to devices other than the new device 4000c are stored.
[0476] In operation S1940, the IoT server 3000 adds the new device 4000c to the device list corresponding to the user account. The IoT server 3000 may store the device identification information and storage information of the device determination model and the function determination model of the new device 4000c in the device list registered according to each user account. Operation S1940 may be an operation to register the new device 4000c. When operation S1940 is executed, the new device 4000c is a registered device.
[0477] In operation S1950, the IoT server 3000 provides the user account, the device list, and the storage information of the device determination model and the function determination model of each device candidate included in the device list to the voice assistant server 2000.
[0478] In operation S1960, the voice assistant server 2000 determines the device that stores the device determination model among the multiple devices registered in the user account as the hub device 1000. In an embodiment of the present disclosure, the voice assistant server 2000 may identify the device that stores the device determination model based on the storage information of the device determination model of each of the multiple devices registered according to the user account obtained from the IoT server 3000, and may determine the identified device as the hub device 1000.
[0479] In operation S1970, the voice assistant server 2000 sends the device identification information and storage information of the device determination model and the function determination model of the new device 4000c to the hub device 1000.
[0480] In operation S1980, the hub device 1000 updates the device determination model 1330 to add the new device 4000c to the device candidates that can be determined as operation execution devices by the device determination model 1330. By the update, the device determination model 1330 can interpret the text related to the new device 4000c by using the first NLU model 1332, and the device determination model 1330 can determine the new device 4000c as an operation execution device according to the interpretation result.
[0481] For example, when the new device 4000c is an air purifier, before the device determination model 1330 is updated, the device determination model 1330 may not be able to interpret text such as "operate the fine dust purification mode", may not be able to determine the operation execution device, and thus may output a "failure" message. When the new device 4000c is added to the device candidates by updating the device determination model 1330, the device determination model 1330 can interpret the text saying "operate the fine dust purification mode" by using the updated first NLU model 1332, and can determine the air purifier as the new device 4000c as the operation execution device based on the interpretation result of the text.
[0482] Figure 20 FIG. is a diagram illustrating a network environment including a hub device 1000, a plurality of devices 4000, and a voice assistant server 2000.
[0483] Reference Figure 20 , the hub device 1000, the plurality of devices 4000, the voice assistant server 2000, and the IoT server 3000 may be connected to each other by using a wired communication or a wireless communication method and may perform communication. In an embodiment of the present disclosure, the hub device 1000 and the plurality of devices 4000 may be directly connected to each other through a communication network, but the present disclosure is not limited thereto.
[0484] The hub device 1000 and the plurality of devices 4000 may be connected to the voice assistant server 2000, and the hub device 1000 may be connected to the plurality of devices 4000 through a server. In addition, the hub device 1000 and the plurality of devices 4000 may be connected to the IoT server 3000. In another embodiment of the present disclosure, the hub device 1000 and the plurality of devices 4000 may both be connected to the voice assistant server 2000 through a communication network and may be connected to the IoT server 30000 through the voice assistant server 2000. In another embodiment of the present disclosure, the hub device 1000 may be connected to the plurality of devices 4000, and the hub device 1000 may be connected to the plurality of devices 4000 through one or more nearby access points. In addition, the hub device 1000 may be connected to the plurality of devices 4000 in a state where the hub device 1000 is connected to the voice assistant server 2000 or the IoT server 3000.
[0485] The hub device 1000, the plurality of devices 4000, the voice assistant server 2000, and the IoT server 3000 may be connected through a LAN, a WAN, a VAN, a mobile radio communication network, a satellite communication network, or a combination thereof. Examples of wireless communication methods may include but are not limited to Wi-Fi, Bluetooth, BLE, Zigbee, WFD, UWB, IrDA, and NFC.
[0486] In an embodiment of the present disclosure, the hub device 1000 may receive a voice input from a user. At least one of the plurality of devices 4000 may be a target device that receives a control command from the voice assistant server 2000 and / or the IoT server 3000 and performs a specific operation. At least one of the plurality of devices 4000 may be controlled to perform a specific operation based on the voice input from the user received by the hub device 1000. In an embodiment of the present disclosure, at least one of the plurality of devices 4000 may receive a control command from the hub device 1000 without receiving a control command from the voice assistant server 2000 and / or the IoT server 3000.
[0487] The hub device 1000 can receive voice input (e.g., speech) from a user. In an embodiment of the present disclosure, the hub device 1000 can include an ASR model. In an embodiment of the present disclosure, the hub device 1000 can include an ASR model with limited functions. For example, the hub device 1000 can include an ASR model having a function of detecting a specified voice input (e.g., a wake-up input such as "Hi, Bixby" or "OK, Google") or a function of preprocessing a voice signal obtained from a partial voice input. Although the hub device 1000 is Figure 20 an AI speaker in, the present disclosure is not limited thereto. In an embodiment of the present disclosure, one of the plurality of devices 4000 can be the hub device 1000. In addition, the hub device 1000 can include a first NLU model, a second NLU model, and a natural language generation model. In this case, the hub device 1000 can receive the user's voice input through a microphone, or can receive the user's voice input from at least one of the plurality of devices 4000. When receiving the user's voice input, the hub device 1000 can process the user's voice input by using the ASR model, the first NLU model, the second NLU model, and the natural language generation model, and can provide a response to the user's voice input.
[0488] The hub device 1000 can determine the type of the target device for performing the operation of the user's intention based on the received voice signal. The hub device 1000 can receive the voice signal as an analog signal and can convert the voice part into computer-readable text by performing ASR. The hub device 1000 can interpret the text by using the first NLU model and can determine the target device based on the interpretation result. The hub device 1000 can determine at least one of the plurality of devices 4000 as the target device. The hub device 1000 can select a second NLU model corresponding to the determined target device from among the plurality of stored second NLU models. The hub device 1000 can determine the operation to be performed by the target device requested by the user by using the selected second NLU model. Based on determining that there is no second NLU model corresponding to the determined target device among the plurality of stored second NLU models, the hub device 1000 can send at least a part of the text to at least one of the plurality of devices 4000 and the voice assistant server 2000. The hub device 1000 sends the information about the determined operation to the target device so that the determined target device performs the determined operation.
[0489] The hub device 1000 may receive information of multiple devices 4000 from the IoT server 3000. The hub device 1000 may determine a target device by using the information of the received multiple devices 4000. In addition, the hub device 1000 may control the target device to perform a determined operation by using the IoT server 3000 as a relay server for sending information about the determined operation.
[0490] The hub device 1000 may receive a voice input of a user through a microphone and may send the received voice input to the voice assistant server 2000. In an embodiment of the present disclosure, the hub device 1000 may obtain a voice signal from the received voice input and may send the voice signal to the voice assistant server 2000.
[0491] In Figure 20 the embodiment, the multiple devices 4000 include, but are not limited to, a first device 4100 as an air conditioner, a second device 4200 as a TV, a third device 4300 as a washing machine, and a fourth device 4400 as a refrigerator. For example, the multiple devices 4000 may include at least one of a smart phone, a tablet PC, a tablet PC, a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop PC, a netbook computer, a workstation, a server, a PDA, a PMP, an MP3 player, a mobile medical device, a camera, and a wearable device. In an embodiment of the present disclosure, the multiple devices 4000 may be household appliances. The household appliances may include at least one of the following items: a TV, a digital video disc (DVD) player, an audio device, a refrigerator, an air conditioner, a vacuum cleaner, an oven, a microwave oven, a washing machine, an air purifier, a set-top box, a home automation control panel, a security control panel, a game console, an electronic key, a camera, and an electronic photo frame.
[0492] The voice assistant server 2000 may determine the type of a target device for performing an operation for a user intent based on the received voice signal. The voice assistant server 2000 may receive the voice signal as an analog signal from the hub device 1000 and may convert the voice part into computer-readable text by performing ASR. The voice assistant server 2000 may interpret the text by using a first NLU model and may determine the target device based on the interpretation result. In addition, the voice assistant server 2000 may receive at least a part of the text and information about the target device determined by the hub device 1000 from the hub device 1000. In this case, the hub device 1000 converts the user voice signal into text by using the ASR model and the first NLU model of the hub device 1000 and determines the target device by interpreting the text. In addition, the hub device 1000 sends at least a part of the text and information about the determined target device to the voice assistant server 2000.
[0493] The voice assistant server 2000 can determine an operation to be performed by a target device requested by a user by using a second NLU model corresponding to the determined target device. The voice assistant server 2000 can receive information on a plurality of devices 4000 from the IoT server 3000. The voice assistant server 2000 can determine a target device by using the received information on the plurality of devices 4000. In addition, the voice assistant server 2000 can control the target device to perform the determined operation by using the IoT server 3000 as a relay server for sending information on the determined operation. The IoT server 3000 can store information on a plurality of devices 4000 that are connected through a network and previously registered. In an embodiment of the present disclosure, the IoT server 3000 can store at least one of identification information (e.g., device ID information) of the plurality of devices 4000, a device type of each of the plurality of devices 4000, and information on a function execution ability of each of the plurality of devices 4000.
[0494] In an embodiment of the present disclosure, the IoT server 3000 can store status information on power-on / off or an operation being performed for each of the plurality of devices 4000. The IoT server 3000 can send a control command for performing the determined operation to the target device among the plurality of devices 4000. The IoT server 3000 can receive information on the determined target device and information on the determined operation from the voice assistant server 2000, and can send a control command to the target device based on the received information.
[0495] Figure 21A and Figure 21B is a diagram illustrating a voice assistant model 200 that can be executed by the hub device 1000 and the voice assistant server 2000 according to an embodiment of the present disclosure.
[0496] Reference Figure 21A and Figure 21B , the voice assistant model 200 is implemented as software. The voice assistant model 200 can be configured to determine a user's intention from a voice input of the user and control a target device related to the user's intention. When a device controlled by the voice assistant model 200 is added, the voice assistant model 200 can include a first assistant model 200a configured to update an existing model to a new model by learning or the like, and a second assistant model 200b configured to add a model corresponding to the added device to the existing model.
[0497] The first assistant model 200a is a model that determines a target device related to a user's intention by analyzing the user's voice input. The first assistant model 200a may include an ASR model 202, an NLG model 204, a first NLU model 300a, and a device determination model 310. In an embodiment of the present disclosure, the device determination model 310 may include the first NLU model 300a. In another embodiment of the present disclosure, the device determination model 310 and the first NLU model 300a may be configured as separate elements.
[0498] The device determination model 310 is a model for performing an operation of determining a target device by using the analysis result of the first NLU model 300a. The device determination model 310 may include a plurality of detailed models, and one of the plurality of detailed models may be the first NLU model 300a. The first NLU model 300a or the device determination model 310 may be an AI model.
[0499] When a device controlled by the voice assistant model 200 is added, the first assistant model 200a may update at least the device determination model 310 and the first NLU model 300a by learning. Learning may refer to learning using both the training data for training the existing device determination model and the first NLU model and the additional training data related to the added device. In addition, learning may refer to updating the device determination model and the first NLU model by using only the additional training data related to the added device.
[0500] The second assistant model 200b, which is a model dedicated to a specific device, is a model that determines an operation to be performed by a target device corresponding to a user's voice input from among a plurality of operations that can be performed by the specific device. In Figure 21A this case, the second assistant model 200b may include a plurality of second NLU models 300b, an NLG model 206, and an action plan management module 210. The plurality of second NLU models 300b may respectively correspond to a plurality of different devices. The second NLU model, the NLG model, and the action plan management module may be models implemented by a rule-based system. In an embodiment of the present disclosure, the second NLU model, the NLG model, and the action plan management model may be AI models. The plurality of second NLU models may be elements of a plurality of function determination models.
[0501] When adding a device controlled by the voice assistant model 200, the second assistant model 200b may be configured to add a second NLU model corresponding to the added device. That is, in addition to the existing multiple second NLU models 300b, the second assistant model 200b may further include a second NLU model corresponding to the added device. In this case, the second assistant model 200b may be configured to select, from among the multiple second NLU models including the added second NLU model, a second NLU model corresponding to the determined target device by using the information about the target device determined by the first assistant model 200a.
[0502] Reference Figure 21B , the second assistant model 200b may include multiple action plan management models and multiple NLG models. In Figure 21B , the multiple second NLU models included in the second assistant model 200b may respectively correspond to Figure 21A the second NLU model 300b, each of the multiple NLG models included in the second assistant model 200b may correspond to Figure 21A the NLG model 206, and each of the multiple action plan management models included in the second assistant model 200b may correspond to Figure 21A the action plan management module 210.
[0503] In Figure 21B , the multiple action plan management models may be configured to respectively correspond to the multiple second NLU models. In addition, the multiple NLG models may be configured to respectively correspond to the multiple second NLU models. In another embodiment of the present disclosure, one NLG model may be configured to correspond to the multiple second NLU models, and one action plan management model may be configured to correspond to the multiple second NLU models.
[0504] In Figure 21B , when adding a device controlled by the voice assistant model 200, the second assistant model 200b may be configured to add a second NLU model, an NLG model, and an action plan management model corresponding to the added device.
[0505] In Figure 21BIn [context], when a device controlled by the voice assistant model 200 is added, the first NLU model 300a can be configured to be updated to a new model through learning or the like. In addition, when the device determines that the model 310 includes the first NLU model 300a, the device determines that the model 310 can be configured such that when a device controlled by the voice assistant model 200 is added, the existing model can be completely updated to a new model through learning or the like. The first NLU model 300a or the device determination model 310 can be an AI model. Learning can refer to learning using both the training data for training the existing device determination model and the first NLU model and the additional training data related to the added device. In addition, learning can refer to updating the device determination model and the first NLU model by only using the additional training data related to the added device.
[0506] In Figure 21B [context], when a device controlled by the voice assistant model 200 is added, the second assistant model 200b can be updated by adding the second NLU model, the NLG model, and the action plan management model corresponding to the added device to the existing model. The second NLU model, the NLG model, and the action plan management model can be models implemented through a rule-based system.
[0507] In Figure 21B In an embodiment of [context], the second NLU model, the NLG model, and the action plan management model can be AI models. The second NLU model, the NLG model, and the action plan management model can all be managed as one device according to the corresponding device. In this case, the second assistant model 200b can include a plurality of second assistant models 200b-1, 200b-2, and 200b-3 corresponding to a plurality of devices respectively. For example, the second NLU model corresponding to the TV, the NLG model corresponding to the TV, and the action plan management model corresponding to the TV can be managed as the second assistant model 200b-1. In addition, the second NLU model corresponding to the speaker, the NLG model corresponding to the speaker, and the action plan management model corresponding to the speaker can be managed as the second assistant model 200b-2 corresponding to the speaker. In addition, the second NLU model corresponding to the refrigerator, the NLG model corresponding to the refrigerator, and the action plan management model corresponding to the refrigerator can be managed as the second assistant model 200b-3 corresponding to the refrigerator.
[0508] When adding a device controlled by the voice assistant model 200, the second assistant model 200b can be configured to add a second assistant model corresponding to the added device. That is, in addition to the existing multiple second assistant models 200b-1 to 200b-3, the second assistant model 200b can also include a second assistant model corresponding to the added device. In this case, the second assistant model 200b can be configured to select, from among multiple second assistant models including the second assistant model corresponding to the added device, a second assistant model corresponding to the determined target device by using the information about the target device determined by the first assistant model 200a.
[0509] The program executed by the hub device 1000, the voice assistant server 2000, and the multiple devices 4000 according to the present disclosure can be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. The program can be executed by any system capable of executing computer-readable instructions.
[0510] Software can include a computer program, code, instructions, or a combination of one or more of them, and can configure a processing device to operate as needed or command the processing device individually or jointly.
[0511] Software can be implemented in a computer program including instructions stored in a computer-readable storage medium. The computer-readable storage medium can include, for example, a magnetic storage medium (e.g., ROM, RAM, floppy disk, hard disk, etc.) and an optical reading medium (e.g., compact disc (CD)-ROM, DVD, etc.). The computer-readable recording medium can be distributed in computer systems connected to a network and can store and execute computer-readable code in a distributed manner. The medium can be computer-readable, can be stored in a memory, and can be executed by a processor.
[0512] The computer-readable storage medium can be provided in the form of a non-transitory storage medium. Here, "non-transitory" means that the storage medium does not include a signal and is tangible, but does not distinguish whether the data is stored semi-permanently or temporarily on the storage medium.
[0513] In addition, the program according to an embodiment of the present disclosure can be provided in a computer program product. The computer program product is a product that can be purchased between a seller and a buyer.
[0514] The computer program product can include a software program and a computer-readable storage medium in which the software program is stored. For example, the computer program product can include a product provided by a device or an electronic market (e.g., GooglePlay TMProducts of the type of software programs that are electronically distributed by a manufacturer (e.g., in a store, an App store, etc.) (such as downloadable applications). For electronic distribution, at least a part of the software program can be stored in a storage medium or generated temporarily. In this case, the storage medium can be the storage medium of the manufacturer's server, the server of an electronic marketplace, or the broadcast server that temporarily stores the software program.
[0515] A computer program product can include the storage medium of a server or the storage medium of a device in a system including a server and a device. Alternatively, when there is a third de...
Claims
1. A method for controlling a device performed by a hub device, the method comprises: Receiving a voice signal from a listener device; Converting the received voice signal into text by performing automatic speech recognition (ASR); Analyzing the text by using a first natural language understanding (NLU) model, and determining an operation execution device corresponding to the analyzed text by using a device determination model; Identifying a device that stores a function determination model corresponding to the determined operation execution device from the determined operation execution device and the listener device; And Providing at least a part of the text to the identified device.
2. The method according to claim 1, wherein, The device determination model includes a first NLU model, and the first NLU model is configured to analyze the text and determine an operation execution device based on the analysis result of the text.
3. The method according to claim 1, wherein, The function determination model includes a second NLU model, and the second NLU model is configured to analyze at least a part of the text and obtain operation information about an operation to be performed by the determined operation execution device based on the analysis result of at least a part of the text.
4. The method according to claim 1, further comprising determining whether the determined operation execution device is the same as the listener device.
5. The method according to claim 4, wherein, When it is determined that the operation execution device is the same as the listener device, identifying the device that stores the function determination model includes: Obtaining function determination model information about whether the listener device stores the function determination model in an internal memory; and Based on the obtained function determination model information, determining the listener device as the device that stores the function determination model.
6. The method according to claim 4, wherein, When it is determined that the operation execution device is a device different from the listener device, identifying the device that stores the function determination model includes: Obtaining function determination model information about whether the determined operation execution device stores the function determination model in an internal memory; and Based on the obtained function determination model information, determining whether the operation execution device is the device that stores the function determination model.
7. The method according to claim 1, further comprises: Receiving update data of the device determination model from a voice assistant server; And Updating the device determination model by using the received update data.
8. The method according to claim 7, wherein, The update data includes data for updating the device determination model based on update information of a function determination model included in at least one of an operation execution device or a listener device to determine an updated function from the text and determine an operation execution device corresponding to the updated function.
9. The method according to claim 1, further comprises: Receiving device information of a new device from a voice assistant server, where the device information of the new device includes at least one of device identification information of the new device, storage information of a device determination model, or storage information of a function determination model; And Adding the new device to a device candidate that can be determined as an operation execution device by the device determination model by using the received device information of the new device, and updating the device determination model.
10. A hub device for controlling a device, the hub device comprising: A communication interface configured to perform data communication with at least one of a voice assistant server or a plurality of devices including a listener device; A voice signal receiver configured to receive a voice signal from the listener device; A memory configured to store a program including one or more instructions; and A processor configured to execute one or more instructions of the program stored in the memory, wherein the processor is further configured to: Convert the received voice signal into text by performing automatic speech recognition (ASR), Analyze the text by using a first natural language understanding (NLU) model, and determine an operation execution device corresponding to the analyzed text by using a device determination model, Identify a device storing a function determination model corresponding to the determined operation execution device from the determined operation execution device and the listener device, and Send at least a part of the text to the identified device by using the communication interface.
11. The hub device according to claim 10, wherein, The device determination model includes a first NLU model configured to analyze the text and determine the operation execution device based on the analysis result of the text.
12. The hub device according to claim 10, wherein, The function determination model includes a second NLU model configured to analyze at least a part of the text and obtain operation information about an operation to be performed by the determined operation execution device based on the analysis result of at least a part of the text.
13. The hub device according to claim 10, wherein, The processor is further configured to determine whether the determined operation execution device is the same as the listener device.
14. The hub device according to claim 13, wherein, When it is determined that the operation execution device is the same as the listener device, the processor is further configured to obtain function determination model information about whether the listener device stores the function determination model in an internal memory, and determine the listener device as the device storing the function determination model based on the obtained function determination model information.
15. The hub device according to claim 13, wherein, When it is determined that the operation execution device is a device different from the listener device, the processor is further configured to obtain function determination model information about whether the determined operation execution device stores the function determination model in an internal memory, and determine whether the operation execution device is the device storing the function determination model based on the obtained function determination model information.
16. The hub device according to claim 10, wherein, The processor is further configured to: Receive update data of the device determination model from the voice assistant server by using the communication interface, and Update the device determination model by using the received update data.
17. The hub device according to claim 16, wherein, The updated data includes data for updating a device determination model to determine an updated function from text and determine an operation execution device corresponding to the updated function based on update information for determining a model based on functions included in at least one of an operation execution device or a listener device.
18. The hub device according to claim 10, wherein, the processor is further configured to: receive device information of a new device from a voice assistant server by using a communication interface, the device information of the new device including at least one of device identification information of the new device, storage information of a device determination model, or storage information of a function determination model, and add the new device to a device candidate that can be determined as an operation execution device by the device determination model by using the received device information of the new device, and update the device determination model.
Citation Information
Patent Citations
Cap assembly for a drug delivery device, kit for assembling a cap, method for assembling a cap and drug delivery device comprising a cap assembly.
KR1020190123310A
Device and method for distributed resource allocation in wireless multi-hop network
KR1020200027217A