System and method for registring devices for voice assistant services
By combining or deleting the functions of pre-registered devices, registering new devices with relevant discourse data and action data, and generating or updating the voice assistant model, the functional response problem of the new device in the voice assistant service is solved, and flexible and efficient device control is achieved.
Patent Information
- Application Number
- CN202510162922.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-10
- Filing Date
- 2020-07-29
- Publication Date
- 2025-06-06
AI Technical Summary
When providing voice assistant services, new devices are added to the home network environment, and it is difficult to respond to user voice input by considering the functionality of the new device, especially if the new device is not a device pre-registered with voice assistant services.
Register a new device by combining or deleting the functions of the pre-registered device, register a new device using discourse data related to the functions of the pre-registered device, obtain discourse data and action data related to the functions of the new device, and generate or update the voice assistant model assigned to the new device.
It realizes the functions of effectively reflecting the new device in the voice assistant service, ensuring that the new device can be correctly understood and controlled, and improving the flexibility and adaptability of voice assistant service.
Smart Images

Figure CN120106076A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a system and method for registering a new device for a voice assistant service. Background Art
[0002] With the development of multimedia and network technology, various services have been provided to users through devices. In particular, with the development of voice recognition technology, users can speak their voices (e.g., words) into voice assistant devices and receive response messages as replies to the voice input through service providing agents.
[0003] When the voice assistant service is to understand the intent contained in the voice input from the user, artificial intelligence (AI) technology may be used to decipher the correct intent of the voice input from the user, and rule-based natural language understanding (NLU) may also be used.
[0004] However, when providing a voice assistant service, when a new device is added to a home network environment including a plurality of devices, it is difficult to provide device control in response to a user's voice input by considering the function of the new device. In particular, even when the new device is not a device pre-registered with the voice assistant service, it is necessary to effectively reflect the function of the new device in the voice assistant service. Summary of the invention
[0005] Technical Solution
[0006] A system and method are provided for registering a new device using functionality of a pre-registered device for a voice assistant service.
[0007] According to an aspect of the present disclosure, a system and method for registering a function of a new device by combining or deleting a function of at least one pre-registered device are provided.
[0008] According to an aspect of the present disclosure, a system and method for registering a new device using utterance data related to functions of a pre-registered device are provided.
[0009] According to an aspect of the present disclosure, there is provided a system and method for acquiring utterance data and action data related to a function of a new device using utterance data and action data of a pre-registered device.
[0010] According to one aspect of the present disclosure, a system and method are provided for generating and updating a voice assistant model designated for a new device using utterance data and action data related to functions of the new device.
[0011] Additional aspects will be set forth in part in the description which follows and, in part, will be apparent from the description, or may be learned by practice of the presented embodiments of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which:
[0013] Figure 1 is a conceptual diagram of a system for providing a voice assistant service according to an embodiment of the present disclosure;
[0014] Figure 2 A mechanism for registering a new device with a voice assistant server based on the capabilities of pre-registered devices according to an embodiment of the present disclosure is presented;
[0015] Figure 3 is a flowchart showing a method for registering a new device by a voice assistant server according to an embodiment of the present disclosure;
[0016] Figure 4 is a flow chart showing a method in which a server compares the functionality of a pre-registered device with the functionality of a new device according to an embodiment of the present disclosure;
[0017] Figure 5A A comparison between the functionality of a pre-registered device and the functionality of a new device according to an embodiment of the present disclosure is shown;
[0018] Figure 5B A comparison between a set of functions of a pre-registered device and functions of a new device according to an embodiment of the present disclosure is shown;
[0019] Figure 5C Demonstrating a comparison of the functions and combinations of function groups of pre-registered devices with the functions of a new device according to an embodiment of the present disclosure;
[0020] Figure 5D Demonstrating comparison of the functions and combinations of function groups of a plurality of pre-registered devices with the functions of a new device according to an embodiment of the present disclosure;
[0021] Figure 5E Deleting some functions of a pre-registered device and comparing the remaining functions with the functions of a new device according to an embodiment of the present disclosure is demonstrated;
[0022] Figure 6 is a flowchart showing a method in which a voice assistant server generates utterance data and action data related to a function different from a function of a pre-registered device among functions of a new device according to an embodiment of the present disclosure;
[0023] Fig. 7A A query output from a voice assistant server for generating speech data and action data related to a function of a new device according to an embodiment of the present disclosure is shown;
[0024] Figure 7B A query output for recommending utterance sentences to generate utterance data and action data related to a function of a new device according to an embodiment of the present disclosure is shown;
[0025] Figure 8 is a flowchart showing a method for a voice assistant server to expand speech data according to an embodiment of the present disclosure;
[0026] Fig.9A The method of generating similar utterance data from utterance data according to an embodiment of the present disclosure is shown;
[0027] Fig. 9B Representative utterance sentences and proximate utterance data mapped to action data according to an embodiment of the present disclosure are shown;
[0028] Fig. 10A Generating proximate utterance data from utterance data according to another embodiment of the present disclosure is shown;
[0029] Fig. 10B Representative utterance sentences and proximate utterance data mapped to action data according to an embodiment of the present disclosure are shown;
[0030] Fig.11A Presenting speech data according to an embodiment of the present disclosure;
[0031] Fig. 11B Speech data according to another embodiment of the present disclosure is shown;
[0032] Fig.12 shows specifications of a device according to another embodiment of the present disclosure;
[0033] Fig.13 is a block diagram of a voice assistant server according to an embodiment of the present disclosure;
[0034] Fig.14 is a block diagram of a voice assistant server according to another embodiment of the present disclosure;
[0035] Fig.15 is a conceptual diagram showing an action plan management model according to an embodiment of the present disclosure;
[0036] Fig.16 A capsule database stored in an action plan management model according to an embodiment of the present disclosure is shown;
[0037] Fig.17 is a block diagram of an Internet of Things (IoT) cloud server according to an embodiment of the present disclosure; and
[0038] Fig.18 is a block diagram of a client device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0039] According to a first aspect of the present disclosure, there is provided a method for registering a new device for a voice assistant service, which is performed by a server, the method comprising: obtaining a first technical specification indicating a first function of a pre-registered device and a second technical specification indicating a second function of the new device; comparing the first function of the pre-registered device with the second function of the new device based on the first technical specification and the second technical specification; identifying the first function of the pre-registered device that matches the second function of the new device as a matching function based on the comparison; obtaining pre-registered speech data related to the matching function; generating action data for the new device based on the matching function and the pre-registered speech data; and storing the pre-registered speech data and action data associated with the new device, wherein the action data includes data related to a series of detailed operations of the new device corresponding to the pre-registered speech data.
[0040] According to a second aspect of the present disclosure, a server for registering a new device for a voice assistant service is provided, the server comprising: a communication interface; a memory storing a program comprising one or more instructions; and a processor configured to execute one or more instructions of the program stored in the memory to: obtain a first technical specification indicating a first function of a pre-registered device and a second technical specification indicating a second function of the new device, compare the first function of the pre-registered device with the second function of the new device based on the first technical specification and the second technical specification; identify the first function of the pre-registered device that matches the second function of the new device as a matching function; obtain pre-registered speech data related to the matching function; generate action data for the new device based on the matching function and the pre-registered speech data; and store the pre-registered speech data and action data associated with the new device in a database, and wherein the action data includes data related to a series of detailed operations of the new device corresponding to the pre-registered speech data.
[0041] According to a third aspect of the present disclosure, there is provided a computer-readable recording medium having thereon a program for a computer to execute the method of the first aspect of the present disclosure and the operation of the second aspect of the present disclosure.
[0042] Invention Mode
[0043] Embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, embodiments of the present disclosure may be implemented in a variety of different forms and are not limited to those discussed herein. In the accompanying drawings, parts not related to the description are omitted for clarity, and the same numerals refer to the same elements throughout the specification.
[0044] When A is referred to as being “connected” to B, it means being “directly connected” to B or “electrically connected” to B, and possibly including C interposed between A and B. Unless otherwise stated, the terms “include” (or including) or “comprise” (or comprising) are inclusive or open-ended, and do not exclude additional, unrecited elements or method steps.
[0045] Throughout this disclosure, the expression “at least one of a, b, or c” indicates only a; only b; only c, both a and b; both a and c; both b and c; all of a, b, and c, or variations thereof.
[0046] According to the embodiments of the present disclosure, functions related to artificial intelligence (AI) are implemented by a processor and a memory storing computer-readable instructions executed by the processor. One or more processors may be present. One or more processors may include a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), a digital signal processor (DSP), etc., a dedicated graphics processor, such as a graphics processing unit (GP), a visual processing unit (VPU), etc., or a dedicated AI processor, such as a neural processing unit (NPU). One or more processors may control the processing of input data according to predefined operating rules or an AI model stored in a memory. When one or more processors are dedicated AI processors, they can be designed as a hardware structure specifically for processing a specific AI model.
[0047] Predefined operation rules or AI models can be constructed by learning. Specifically, the predefined operation rules or AI models constructed by learning refer to predefined operation rules or AI models for performing the desired functions (or objects) constructed when the basic AI model is trained by processing training data through a learning algorithm. Such learning can be performed by the device itself that performs AI according to the present disclosure, or by a separate server and / or system. Examples of learning algorithms may include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited thereto.
[0048] The AI model may include multiple neural network layers. Each of the multiple neural network layers may have multiple weight values, and the neural network calculation is performed by the operation according to the operation result of the previous layer and the multiple weight values. The multiple weight values assigned to the multiple neural network layers can be optimized by the learning results of the AI model. For example, multiple weight values can be updated during the learning process to reduce or minimize the loss value or cost value obtained by the AI model. The artificial neural network may include, for example, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN) or a deep Q network, but is not limited thereto.
[0049] The present disclosure will now be described with reference to the accompanying drawings.
[0050] Figure 1 is a conceptual diagram of a system for providing a voice assistant service according to an embodiment of the present disclosure.
[0051] refer to Figure 1 In an embodiment of the present disclosure, a system for providing a voice assistant service may include a client device 1000, at least one additional device 2000, a voice assistant server 3000, and an Internet of Things (IoT) cloud server 4000. At least one device 2000 may be a device pre-registered in the voice assistant server 3000 or the IoT cloud server 4000 for the voice assistant service.
[0052] The client device 1000 may receive voice input (e.g., speech) from a user. In an embodiment of the present disclosure, the client device 1000 may include a voice recognition module. In an embodiment of the present disclosure, the client device 1000 may include a voice recognition module with limited voice processing capabilities. For example, the client device 1000 may include a voice recognition module with a function of detecting a specified voice input (e.g., a wake-up input such as "Hi, Bixby.", "Okay, Google.", etc.) or a function of preprocessing a voice signal obtained from certain voice inputs. The client device 1000 may be an AI speaker, but is not limited thereto. In an embodiment of the present disclosure, some of the at least one device 2000 may also be implemented as an additional client device 1000.
[0053] At least one device 2000 may be a target controlled device that performs a specific operation in response to a control command from the voice assistant server 3000 and / or the IoT cloud server 4000. The at least one device 2000 may be controlled to perform a specific operation based on a user's voice input received by the client device 1000. In an embodiment of the present disclosure, at least some of the at least one device 2000 may receive a control command from the client device 1000 without receiving any control command from the voice assistant server 3000 and / or the IoT cloud server 4000. Therefore, the at least one device 2000 may be controlled by one or more of the client device 1000, the voice assistant server 3000, and the IoT cloud server 4000 based on the user's voice.
[0054] The client device 1000 may receive the user's voice input through a microphone and forward the voice input to the voice assistant server 3000. In an embodiment of the present disclosure, the client device 1000 may obtain a voice signal from the received voice input and forward the voice signal to the voice assistant server 3000.
[0055] The voice assistant server 3000 can receive a user's voice input from the client device 1000, select a target device to perform an operation from at least one device 2000 according to the user's intention by interpreting the user's voice input, and provide information about the target device or the operation to be performed by the target device to the IoT cloud server 4000 or at least one target device 2000 to be controlled.
[0056] The IoT cloud server 4000 may register and manage information about at least one device 2000 for the voice assistant service, and provide the device information of at least one device 2000 for the voice assistant service to the voice assistant server 3000. The device information of at least one device 2000 may be information related to a device for providing a voice assistant service, and includes, for example, identification (ID) information (device ID information), capability information, location information, and status information of the device. In addition, the IoT cloud server 4000 may receive information about a target device and an operation to be performed by the target device from the voice assistant server 3000, and provide control information for controlling the operation of at least one device 2000 to at least one target device 2000 to be controlled.
[0057] When a new device 2900 is added to, for example, a home network for voice assistant service, the voice assistant server 3000 may generate utterance data and action data of the new device 2900 using pre-registered functions, utterance data, and operations corresponding to the utterance data of at least one device 2000. The voice assistant server 3000 may generate or update a voice assistant model to be used for the new device 2900 using the utterance data and action data of the new device 2900.
[0058] The speech data is related to the speech representing the user's speech issued by the user to obtain the voice assistant service. The speech data can be used to interpret the user's intention related to the operation of the device 2000. The speech data may include at least one of a speech sentence in text form or a speech parameter in the form of an output value of a natural language understanding (NLU) model. The speech parameter may be data output from the NLU model, which may include intentions and parameters. The intention may be information determined by interpreting the text using the NLU model, and may indicate the intention of the user's speech. For example, the intention may be information indicating the user's intention or the user's request for the device operation to be performed by at least one device 2000 to be controlled. The intention may include information indicating the intention of the user's speech (hereinafter referred to as the intention information) and a numerical value corresponding to the information indicating the user's intention. The numerical value may indicate the probability that the text is related to the information indicating a specific intention or intention. When the result of interpreting the text using the NLU model is that there are multiple pieces of information indicating the user's intention, one of the multiple pieces of intention information with the highest numerical value may be determined as the intention. In addition, the parameter may be variable information for determining the detailed operation of the device related to the intention. The parameter may be information related to the intention, and there may be multiple types of parameters corresponding to the intention. The parameter may include variable information for determining operation information of the device and a numerical value indicating the probability that the text is related to the variable information. As a result of interpreting the text using the NLU model, multiple variable information indicating the parameter may be obtained. In this case, one of the multiple variable information having the highest numerical value corresponding to the variable information may be determined as the parameter.
[0059] Action data may be data corresponding to specific speech data and related to a series of detailed operations of at least one device 2000. For example, the action data may include information corresponding to the specific speech data and related to the detailed operations to be performed by at least one device 2000, the correlation between each detailed operation and another detailed operation, and the execution order of the detailed operations. The correlation between a detailed operation and another detailed operation may include information about another detailed operation to be performed before performing a detailed operation to perform the one detailed operation. For example, when the operation to be performed is "play music", "power on" may be another detailed operation to be performed before the "play music" operation. Action data may include, for example, functions to be performed by the target device to perform specific operations, the execution order of functions, input values required to perform functions, and output values output as a result of performing functions.
[0060] When a new device 2900 is identified, the voice assistant server 3000 may acquire function information of the new device 2900, and determine preregistered utterance data that can be used in relation to the function of the new device 2900 by comparing the function information of at least one preregistered device 2000 with the function information of the new device 2900. In addition, the voice assistant server 3000 may edit the preregistered utterance data and the corresponding function, and generate action data using the edited utterance data and data related to the function.
[0061] At least one device 2000 may include a smart phone, a tablet personal computer (tablet PC), a personal computer (PC), a smart TV (smart TV), a personal digital assistant, a laptop, a media player, a micro server, a global positioning system (GPS), an electronic book (e-book) reader, a digital broadcast terminal, a navigation, an information kiosk, an MP3 player, a digital camera, and a mobile or non-mobile computing device, but is not limited thereto. In addition, at least one device 2000 may be a household appliance such as a table lamp, an air conditioner, a television, a robot cleaner, a washing machine, a scale, a refrigerator, a set-top box, a home automation control panel, a security control panel, a game console, an electronic key, a camcorder, or an electronic photo frame equipped with communication and data processing functions. In addition, at least one device 2000 may be a wearable device such as a watch, glasses, a hairband, a ring, each of which has a communication function and a data processing function. However, it is not limited thereto, and at least one device 2000 may include any type of device capable of sending or receiving data from a voice assistant server 3000 and / or an IoT cloud server 4000 through a network such as a home network, a wired or wireless network, or a cellular data network.
[0062] The network may include a local area network (LAN), a wide area network (WAN), a value-added network (VAN), a mobile radio communication network, a satellite communication network, and any combination thereof. Figure 1 The network constituent entities shown in the figure perform a comprehensive data communication network in which smooth communication is performed between each other, and the network includes wired Internet, wireless Internet, and mobile wireless communication networks. Wireless communication may include any of various wireless communication protocols and technologies, including wireless LAN (Wi-fi), Bluetooth, low-power Bluetooth, Zigbee, Wi-fi Direct (WFD), ultra-wideband (UWB), infrared data association (IrDA), near field communication (NFC), etc., but is not limited thereto.
[0063] Figure 2 A mechanism for registering a new device with a voice assistant server based on the functionality of a pre-registered device according to an embodiment of the present disclosure is presented.
[0064] refer to Figure 2 , when a new device 2900 (e.g., air conditioner B) is identified, the voice assistant server 3000 may obtain function and operation information of air conditioner B, which indicates the function of air conditioner B and the operation of air conditioner B. Some functions and operations may include temperature setting, fan speed setting, operation scheduling, humidity setting, etc. The voice assistant server 3000 may compare the function information of air conditioner B with the function information of the pre-registered device 2000, air conditioner A, and dehumidifier A.
[0065] The voice assistant server 3000 may compare the functions of the air conditioner B with the functions of the air conditioner A and the functions of the dehumidifier A, and identify, among the functions of the air conditioner B, functions that are the same as or similar to the functions of the air conditioner A and the dehumidifier A. For example, the voice assistant server 3000 may identify that “power on / off”, “cooling mode on / off”, “dehumidification mode on / off”, “heating up / down”, and “humidity control” among the functions of the air conditioner B correspond to the functions of the air conditioner A and the dehumidifier A.
[0066] The voice assistant server 3000 may determine the speech data corresponding to at least one of the identified functions and generate action data corresponding to the determined speech data. For example, the voice assistant server 3000 may generate or edit speech data corresponding to at least one of the functions of the air conditioner B using the speech data "turn on the power" corresponding to the "power on" of the air conditioner A, the speech data "lower the temperature" corresponding to the "cooling mode on, lower the temperature" of the air conditioner A, the speech data "increase the temperature" corresponding to the "cooling mode on, raise the temperature" of the air conditioner A, the speech data "turn on the power" corresponding to the "power on" of the dehumidifier A, and the speech data "lower the humidity" corresponding to the "dehumidification" of the dehumidifier A.
[0067] Figure 3is a flowchart showing a method for registering a new device performed by the voice assistant server 3000 according to an embodiment of the present disclosure. Fig.12 The specifications of a device according to an embodiment of the present disclosure are presented.
[0068] In operation S300, the voice assistant server 3000 acquires function information about functions of the new device 2900 and the pre-registered device 2000. When the new device 2900 is added to the system for the voice assistant service, the voice assistant server 3000 may identify functions supported by the new device 2900 from the technical specifications of the new device 2900 acquired through the new device 2900 or an external source such as the manufacturer of the new device 2900 or a database storing the technical specifications of the device. Fig.12 , the voice assistant server 3000 can identify the functions supported by the new device 2900 from a specification that includes the device's identification number, model or part number, the name of the executable function, a description of the executable function, and information about factors required to perform the function.
[0069] The voice assistant server 3000 may identify the function of the pre-registered device 2000 from the technical specifications of the device 2000. The voice assistant server 3000 may identify the function of the device 2000 from the technical specifications stored in the database (DB) of the voice assistant server 3000. Alternatively, the voice assistant server 3000 may receive the specification of the device 2000 stored in the database of the IoT cloud server 4000 and identify the function of the device 2000 from the technical specifications. Similar to the technical specifications of the new device 2900, the technical specifications of the pre-registered device 2000 may include information about the identification number, model or part number of the device, the name of the executable function, a description of the executable function, and information about factors required to execute the function.
[0070] In operation S310, the voice assistant server 3000 may determine whether the function of the new device 2900 is the same as or similar to the function of the pre-registered device 2000. The voice assistant server 3000 may identify any function of the new device 2900 that is the same as or similar to the function of the pre-registered device 2000 by comparing the function of the pre-registered device 2000 with the function of the new device 2900.
[0071] The voice assistant server 3000 may identify a function name indicated as supported by the new device 2900 from the technical specifications of the new device 2900, and determine whether the identified name is the same as or similar to a function name supported by the pre-registered device 2000. In this case, the voice assistant server 3000 may store information about names and synonyms indicating certain functions in association, and determine whether the function of the pre-registered device 2000 and the function of the new device 2900 are the same as or similar to each other based on the stored information about the synonyms.
[0072] In addition, the voice assistant server 3000 can determine whether the functions are the same or similar to each other by referring to the utterance data. The voice assistant server 3000 can use the utterance data related to the function of the pre-registered device 2000 to determine whether the function of the new device 2900 is the same or similar to the function of the pre-registered device 2000. In this case, the voice assistant server 3000 can determine whether the function of the new device 2900 is the same or similar to the function of the pre-registered device 2000 based on the meaning of the words included in the utterance data.
[0073] The voice assistant server 3000 may determine whether a single function of the new device 2900 is the same as or similar to a single function of the pre-registered device 2000. A single function may refer to functions such as "power on", "power off", "heat up", and "temperature down". The voice assistant server 3000 may determine whether a set of functions of the new device 2900 is the same as or similar to a set of functions of the pre-registered device 2000. A set of functions may refer to a set of functional combinations of a single function, such as "power on + temperature up", "temperature down + dehumidification".
[0074] When it is determined in operation S310 that the function of the pre-registered device 2000 and the function of the new device 2900 are identical or similar to each other, the voice assistant server 3000 may acquire pre-registered utterance data related to the identical or similar function in operation S320.
[0075] The voice assistant server 3000 may extract utterance data corresponding to a function determined to be the same as or similar to a function of the new device 2900 among functions of the pre-registered device 2000 from the database.
[0076] The voice assistant server 3000 may extract utterance data corresponding to a function group determined to be the same as or similar to the function group of the new device 2900 among the function groups of the pre-registered device 2000 from the database.
[0077] In this case, utterance data corresponding to the function of the pre-registered device 2000 and utterance data corresponding to the function group of the pre-registered device 2000 may be stored in the database before the new device 2900 is installed or set to the network.
[0078] At the same time, the voice assistant server 3000 may edit a function and a group of functions determined to be the same or similar, and generate speech data corresponding to the edited functions. The voice assistant server 3000 may combine functions determined to be the same or similar and generate speech data corresponding to the combined functions. In addition, the voice assistant server 3000 may combine a function and a group of functions determined to be the same or similar, and generate speech data corresponding to the combined functions. In addition, the voice assistant server 3000 may delete some functions in a group of functions determined to be the same or similar, and generate speech data corresponding to a group of functions from which some functions are deleted.
[0079] The voice assistant server 3000 can expand the utterance data. The voice assistant server 3000 can generate similar utterance data having the same meaning but different expressions from the extracted or generated utterance data by modifying the expression of the extracted or generated utterance data.
[0080] In operation S330, the voice assistant server 3000 may generate action data for the new device 2900 based on the same or similar functions and speech data. The action data may be data indicating the detailed operation of the device and the execution order of the detailed operation. The action data may include, for example, an identification value of the detailed operation, the execution order of the detailed operation, and a control command for executing the detailed operation, but is not limited thereto.
[0081] For example, when the function corresponding to the utterance data is a single function, the voice assistant server 3000 may generate action data including detailed operations representing the single function. In another example, when the function corresponding to the utterance data is a group of functions, the voice assistant server 3000 may generate detailed operations representing the functions in the group and the execution order of the detailed operations.
[0082] In operation S340 , the voice assistant server 3000 may generate or update a voice assistant model related to the new device 2900 using the utterance data and the action data.
[0083] The voice assistant server 3000 may generate or update a voice assistant model related to the new device 2900 using utterance data corresponding to the function of the pre-registered device 2000 related to the function of the new device 2900, newly generated utterance data related to the function of the new device 2900, and extended utterance data and action data. The voice assistant server 3000 may accumulate and store utterance data and action data related to the new device 2900. In addition, the voice assistant server 3000 may generate or update a conceptual action network (CAN), which is a package type database included in the action plan management model.
[0084] The voice assistant model associated with the new device 2900 is associated with the new device 2900 as a model for the voice assistant service, which determines the operation to be performed by the target device corresponding to the user's voice input. The voice assistant model associated with the new device 2900 may include, for example, an NLU model, a natural language generation (NLG) model, and an action plan management model. The NLU model associated with the new device 2900 is an AI model for interpreting the user's input voice in consideration of the function of the new device 2900, and the NLG model associated with the new device 2900 is an AI model for generating a natural language for a conversation with the user in consideration of the function of the new device 2900. In addition, the action plan management model associated with the new device 2900 is a model for planning the operation information performed by the new device 2900 in consideration of the function of the new device 2900. The action plan management model can select the detailed operation to be performed by the new device 2900 based on the voice issued by the interpreted user, and plan the execution order of the selected detailed operation. The action plan management model can use the planning results to obtain the operation information about the detailed operation to be performed by the new device 2900. The operation information may be information related to the detailed operation to be performed by the device, the association between the detailed operations, and the execution order of the detailed operations. The operation information may include, for example, the functions performed by the new device 2900 to perform the detailed operation, the execution order of the functions, the input values required to perform the functions, and the output values output as a result of performing the functions.
[0085] When a voice assistant model for a new device 2900 already exists, the voice assistant server 3000 may update the voice assistant model.
[0086] The voice assistant server 3000 may generate or update a voice assistant model associated with the new device 2900 using speech data corresponding to the functions of the pre-registered device 2000 associated with the functions of the new device 2900, newly generated speech data associated with the functions of the new device 2900, and expanded speech data and action data.
[0087] The action plan management model can manage information about multiple detailed operations and information about the relationship between the multiple detailed operations. The correlation between each of the multiple detailed operations and another detailed operation can include information about another detailed operation to be performed before performing one detailed operation to perform the one detailed operation.
[0088] In an embodiment of the present disclosure, the action plan management model may include a CAN, an encapsulation type database indicating the operations of the device and the correlation between the operations. The CAN may include functions to be performed by the device to perform specific operations, the execution order of the functions, the input values required to perform the functions, and the output values output as a result of performing the functions, and may be implemented in an ontology diagram including knowledge triples indicating concepts and relationships between concepts.
[0089] When it is determined in operation S310 that the functions of the pre-registered device 2000 and the functions of the new device 2900 are not identical or similar to each other, the voice assistant server 3000 may request speech data and action data for functions different from the functions of the pre-registered device 2000 in operation S350. The voice assistant server 3000 may register different functions of the new device 2900 and output a query message to the user to generate and edit speech data related to the different functions. The query message may be provided to the user via the client device 1000, the new device 2900, or the developer's device. The voice assistant server 3000 receives a user's response to the query message from the client device 1000, the new device 2900, or the developer's device. The voice assistant server 3000 may provide a software development kit (SDK) tool for registering the functions of the new device 2900 with the client device 1000, the new device 2900, or the developer's device. In addition, the voice assistant server 3000 may provide a list of functions different from the functions of the pre-registered device 2000 in the functions of the new device 2900 to the user's device 2000 or the developer's device. The voice assistant server 3000 may provide recommended utterance data related to at least some different functions to the user's device 2000 or the developer's device.
[0090] In operation S360, the voice assistant server 3000 may acquire speech data and action data. The voice assistant server 3000 may interpret the response to the query using an NLU model. The voice assistant server 3000 may interpret the user's response or the developer's response using an NLU model trained to register functions and generate speech data. The voice assistant server 3000 may generate speech data related to the functions of the new device 2900 based on the interpreted response. The voice assistant server 3000 may generate speech data related to the functions of the new device 2900 using the interpreted user response or the interpreted developer's response, and recommend speech data. The voice assistant server 3000 may select some functions of the new device 2900 and generate speech data related to each selected function. In addition, the voice assistant server 3000 may select some functions of the new device 2900 and generate speech data related to the combination of the selected functions. In addition, the voice assistant server 3000 may generate synonymous speech data having the same meaning but different expressions from the generated speech data. The voice assistant server 3000 may generate action data using the generated speech data. The voice assistant server 3000 may identify functions of the new device 2900 related to the generated utterance data, and determine an execution order of the identified functions to generate action data corresponding to the generated utterance data.
[0091] Figure 4 is a flowchart illustrating a method performed by a server to compare functions of a pre-registered device with functions of a new device according to an embodiment of the present disclosure.
[0092] In operation S400, the voice assistant server 3000 may compare the functions of the pre-registered device 2000 with the functions of the new device 2900. The voice assistant server 3000 may compare the function names supported by the new device 2900 with the function names supported by the pre-registered device 2000. In this case, the voice assistant server 3000 may store information about names and synonyms indicating certain functions, and compare the functions of the pre-registered device 2000 with the functions of the new device 2900 based on the stored information about the synonyms.
[0093] In addition, the voice assistant server 3000 may refer to the utterance data stored in the IoT cloud server 4000 to compare the functions of the pre-registered device 2000 and the functions of the new device 2900. The voice assistant server 3000 may use the utterance data related to the functions of the pre-registered device 2000 to determine whether the functions of the new device 2900 are the same as or similar to the functions of the pre-registered device 2000. In this case, the voice assistant server 3000 may determine whether the functions of the new device 2900 are the same as or similar to the functions of the pre-registered device 2000 based on the meaning of the words included in the utterance data.
[0094] In operation S405, the voice assistant server 3000 may determine whether there is any function of the new device 2900 that does not correspond to the function of the pre-registered device 2900. The voice assistant server 3000 may determine whether all functions of the new device 2900 correspond to at least one function of the pre-registered device 2000. For example, the voice assistant server 3000 may identify a function corresponding to the function of the new device 2900 among the functions of the first device 2100, the second device 2200, and the third device 2300.
[0095] When the function name of the new device 2900 is the same as the function name of the pre-registered device 2000, the voice assistant server 3000 may determine that the function of the pre-registered device 2000 corresponds to the function of the new device 2900.
[0096] In addition, when it is determined that the function name of the new device 2900 is similar to the function name of the pre-registered device 2000 and the function of the pre-registered device 2000 and the function of the new device 2900 have the same effect of controlling the device, the voice assistant server 3000 can determine that the function of the pre-registered device 2000 corresponds to the function of the new device 2900.
[0097] In operation S405, when it is determined that the function of the new device 2900 corresponds to the function of the pre-registered device 2000, the voice assistant server 3000 may perform operations S320 to S340. The voice assistant server 3000 may generate the utterance data and action data related to the function of the new device 2900 using the utterance data and action data related to the function of the pre-registered device 2000 corresponding to the function of the new device 2900, and generate or update the voice assistant model for providing the voice assistant service related to the new device 2900.
[0098] In operation S405, when it is determined that at least one function of the new device 2900 does not correspond to the function of the pre-registered device 2000, the voice assistant server 3000 may combine the functions of the pre-registered device 2000 in operation S410.
[0099] The voice assistant server 3000 may combine a single function of at least one device 2000. For example, the voice assistant server 3000 may combine a first function of the first device 2100 and a second function of the first device 2100. In another example, the voice assistant server 3000 may combine a first function of the first device 2100 and a third function of the second device 2200.
[0100] The voice assistant server 3000 may combine a set of functions of at least one device 2000. For example, the voice assistant server 3000 may combine a first set of functions of the first device 2100 and a second set of functions of the first device 2100. In another example, the voice assistant server 3000 may combine a first set of functions of the first device 2100 and a third set of functions of the second device 2200.
[0101] The voice assistant server 3000 may combine a single function and a group of functions of at least one device 2000. For example, the voice assistant server 3000 may combine a first function of the first device 2100 and a first group of functions of the first device 2100. In another example, the voice assistant server 3000 may combine the first function of the first device 2100 and a third group of functions of the second device 2200.
[0102] The voice assistant server 3000 may combine the functions of the pre-registered device 2000 according to the utterance data corresponding to the functions of the pre-registered device 2000. For example, the voice assistant server 3000 may extract the first utterance data corresponding to the first function of the first device 2100 and the second utterance data corresponding to the second function of the first device 2100 from the database, and determine to combine the first function and the second function based on the meaning of the first utterance data and the second utterance data. In another example, the voice assistant server 3000 may extract the first utterance data corresponding to the first function of the first device 2100 and the third utterance data corresponding to the third function of the second device 2200 from the database, and determine to combine the first function and the third function based on the meaning of the first utterance data and the third utterance data.
[0103] In operation S415, the voice assistant server 3000 may compare the combined function with the function of the new device 2900. The voice assistant server 3000 may compare the name of the combined function with the name of the function supported by the pre-registered device 2000. In addition, the voice assistant server 3000 may compare the combined function with the function of the new device 2900 with reference to the utterance data stored in the IoT cloud server 4000.
[0104] In operation S420, the voice assistant server 3000 may determine whether any function of the new device 2900 does not correspond to the function of the pre-registered device 2900. When the name of the combined function is the same as the function name of the pre-registered device 2000, the voice assistant server 3000 may determine that the combined function corresponds to the function of the new device 2900.
[0105] When it is determined that the name of the combined function is similar to the function name of the pre-registered device 2000 and the combined function has the same purpose as the function of the new device 2900, the voice assistant server 3000 may determine that the combined function corresponds to the function of the new device 2900.
[0106] In operation S420, when it is determined that the function of the new device 2900 corresponds to the function of the pre-registered device 2000, the voice assistant server 3000 may perform operations S320 to S340. The voice assistant server 3000 may generate the utterance data and action data related to the function of the new device 2900 using the utterance data and action data related to the function of the pre-registered device 2000 corresponding to the function of the new device 2900 and the utterance data and action data related to the combined function, and generate or update the voice assistant model for providing the voice assistant service related to the new device 2900.
[0107] When it is determined in operation S420 that the functions of the new device 2900 correspond to the functions of the pre-registered device 2000, the voice assistant server 3000 may delete some functions of the device 2000.
[0108] The voice assistant server 3000 may delete some single functions of at least one device 2000. The voice assistant server 3000 may delete any single function of the device 2000 that is determined to be unsupported by the new device 2900.
[0109] The voice assistant server 3000 may delete some functions of a set of functions of at least one device 2000. The voice assistant server 3000 may delete any function of a set of functions of the device 2000 that is determined not to be supported by the new device 2900.
[0110] The voice assistant server 3000 may delete some of the function groups of at least one device 2000. The voice assistant server 3000 may delete any set of functions of the device 2000 that is determined to be unsupported by the new device 2900.
[0111] In operation S430, the voice assistant server 3000 may compare the functions of the device 2000 remaining after the deletion with the functions of the new device 2900. The voice assistant server 3000 may compare the names of the remaining functions after the deletion with the names of the functions supported by the pre-registered device 2000. In addition, the voice assistant server 3000 may refer to the utterance data stored in the IoT cloud server 4000 to compare the remaining functions after the deletion with the functions of the new device 2900.
[0112] In operation S435, the voice assistant server 3000 may determine whether any function of the new device 2900 does not correspond to the function of the pre-registered device 2900. When the name of the remaining function after deletion is the same as the function name of the pre-registered device 2000, the voice assistant server 3000 may determine that the remaining function of the device 2000 after deletion corresponds to the function of the new device 2900.
[0113] In addition, when it is determined that the name of the remaining function after deletion is similar to the name of the function of the pre-registered device 2000 and the remaining function after deletion has the same purpose as the function of the new device 2900, the voice assistant server 3000 can determine that the remaining function after deletion corresponds to the function of the new device 2900.
[0114] In operation S435, when it is determined that the function of the new device 2900 corresponds to the function of the pre-registered device 2000, the voice assistant server 3000 may perform operations S320 to S340. The voice assistant server 3000 may generate the utterance data and action data related to the function of the new device 2900 using the utterance data and action data related to the function of the pre-registered device 2000 corresponding to the function of the new device 2900, the utterance data and action data related to the combined function, and the utterance data and action data related to the remaining function after deletion, and generate or update the voice assistant model for providing the voice assistant service related to the new device 2900.
[0115] In operation S435, when it is determined that any one function of the new device 2900 does not correspond to the function of the pre-registered device 2000, the voice assistant server 3000 may perform operation S350.
[0116] Although operations S400, S410, S415, S425, and S430 are Figure 4 , but the order is not limited thereto. For example, before comparing the functions of the new device 2900 with the functions of the pre-registered device 2000 as in operation S400, the functions of the pre-registered device 2000 may be combined as in operation S410 or some functions may be deleted as in S425 to establish a database. In this case, by using the database, the functions of the new device 2900 may be compared with the functions of the pre-registered device 2000, the combined functions of the pre-registered device 2000, and the remaining functions after the deletion in operations S400, S415, and S430, and it may be determined whether any function of the new device 2900 does not correspond to the function of the pre-registered device 2000.
[0117] Figure 5A A comparison between the functions of a pre-registered device and the functions of a new device according to an embodiment of the present disclosure is shown.
[0118] refer to Figure 5A , the voice assistant server 3000 can compare the functions of the pre-registered air conditioner A, the functions of the pre-registered dehumidifier A, and the functions of the new air conditioner B.
[0119] For example, the functions supported by the new air conditioner B may include power on / off, cooling mode on / off, dehumidification mode on / off, temperature setting, heating / cooling, humidity setting, humidification / dehumidification, AI mode on / off, etc. In addition, for example, the functions supported by the pre-registered air conditioner A may include power on / off, cooling mode on / off, temperature setting, heating / cooling, etc. In addition, for example, the functions supported by the pre-registered dehumidifier A may include power on / off, humidity setting, humidification / dehumidification, etc.
[0120] The voice assistant server 3000 may determine that “power on / off”, “cooling mode on / off”, “temperature setting”, “heating up / down”, “humidity setting”, and “humidification / dehumidification” among the functions of the new air conditioner B correspond to the functions of the air conditioner A and the dehumidifier A. The voice assistant server 3000 may use the utterance data related to the functions of the air conditioner A and the utterance data related to the functions of the dehumidifier A to determine whether these functions correspond to each other.
[0121] The voice assistant server 3000 may acquire the utterance data "power on / off", "cooling mode on / off", "temperature setting", and "heating up / cooling down" corresponding to each function provided by the pre-registered air conditioner A, and acquire the utterance data "power on / off", "humidity setting", and "humidification / dehumidification" corresponding to each function provided by the pre-registered dehumidifier A. In addition, the voice assistant server 3000 may generate action data of the new air conditioner B using the matched functions and the acquired utterance data.
[0122] Figure 5B A comparison between a set of capabilities of a pre-registered device and capabilities of a new device according to an embodiment of the present disclosure is presented.
[0123] refer to Figure 5B , the voice assistant server 3000 can recognize that a set of functions "cooling mode on + heating" of the pre-registered air conditioner A matches the "cooling mode on / off" and "heating up / cooling" of the new air conditioner B.
[0124] The voice assistant server 3000 may acquire the utterance data “increase temperature” corresponding to a set of functions “cooling mode on + heating up” of the pre-registered air conditioner A. In addition, the voice assistant server 3000 may generate action data for executing the function of “cooling mode on” followed by the function of “heating up” using the functions “cooling mode on / off” and “heating up / cooling down” of the new air conditioner B and the acquired utterance data “increase temperature”.
[0125] Figure 5C It is demonstrated that the combination of functions and groups of functions of pre-registered devices are compared with the functions of a new device according to an embodiment of the present disclosure.
[0126] refer to Figure 5C , the voice assistant server 3000 can recognize that the combination of the function "power on" of the pre-registered air conditioner A and a set of functions "cooling mode on + cooling" of the pre-registered air conditioner A matches the "power on / off", "cooling mode on / off" and "heating up / cooling down" of the new air conditioner B.
[0127] The voice assistant server 3000 may obtain the utterance data "turn on the power" corresponding to the pre-registered function "power on" of the air conditioner A, and the utterance data "lower the temperature" corresponding to the pre-registered set of functions "cooling mode on + cooling" of the air conditioner A. In addition, the voice assistant server 3000 may edit the acquired utterance data. For example, the voice assistant server 3000 may generate utterance data indicating "turn on the air conditioner" and "lower the temperature" from the utterance data "turn on the power" and "lower the temperature".
[0128] In addition, the voice assistant server 3000 can use the "power on / off", "cooling mode on / off" and "temperature increase / cooling" functions of the new air conditioner B and the generated speech data "turn on the air conditioner and lower the temperature" to execute the "power on" function, then execute the "cooling mode on" function, and then execute the "temperature reduction" function.
[0129] Figure 5D Comparing the functions and combinations of function groups of multiple pre-registered devices with the functions of a new device according to an embodiment of the present disclosure is demonstrated.
[0130] refer to Figure 5D , the voice assistant server 3000 can recognize the combination of the following items: i) the pre-registered function "power on" of air conditioner A; ii) a set of functions "cooling mode on + cooling" of pre-registered air conditioner A; and iii) a set of functions "power on + dehumidification" of dehumidifier A matches the "power on / off", "cooling mode on / off", "heating up / cooling down", and "humidification / dehumidification" of the new air conditioner B.
[0131] The voice assistant server 3000 may acquire: i) the speech data "turn on the power" corresponding to the pre-registered function "power on" of the air conditioner A; ii) the speech data "lower the temperature" corresponding to the pre-registered set of functions "cooling mode on + temperature reduction" of the air conditioner A; and iii) the speech data "lower the humidity" corresponding to the pre-registered set of functions "power on + dehumidification" of the dehumidifier A. In addition, the voice assistant server 3000 may edit the acquired speech data. For example, the voice assistant server 3000 may generate speech data indicating "turn on the air conditioner and reduce the temperature and humidity" from the speech data "turn on the power", "reduce the temperature" and "reduce the humidity".
[0132] In addition, the voice assistant server 3000 can use the functions of "power on / off", "cooling mode on / off", "heating up / cooling down", "dehumidification mode on / off" and "humidification / dehumidification" of the new air conditioner B, and use the generated speech data "turn on the air conditioner and lower the temperature and humidity" to execute the function "power on", and then execute the functions "cooling mode on", "cooling down", "dehumidification mode on" and "dehumidification" in a specific order.
[0133] Figure 5E Deleting some functions of a pre-registered device and comparing the remaining functions with the functions of a new device according to an embodiment of the present disclosure is demonstrated.
[0134] refer to Figure 5E , the voice assistant server 3000 can delete "Set temperature to 26 degrees + Check temperature" from a set of functions of "Set temperature to 26 degrees + Check temperature + AI mode on" of the pre-registered air conditioner A, and obtain the remaining function "AI mode on". In addition, the voice assistant server 3000 can identify the combination of the following: i) the remaining function "AI mode on"; and ii) the function "Power on" of the pre-registered air conditioner A matches the "Power on / off" and "AI mode on / off" of the new air conditioner B.
[0135] The voice assistant server 3000 may acquire the utterance data "Turn on the power" corresponding to the function of "power on" of the pre-registered air conditioner A. In addition, the voice assistant server 3000 may extract the utterance data "Turn on the AI function" corresponding to the remaining function "AI mode on" from the utterance data "Turn on the AI function at a temperature of 26 degrees" corresponding to a set of functions of the pre-registered air conditioner A "Set the temperature to 26 degrees + Check the temperature + AI mode on." In addition, the voice assistant server 3000 may generate utterance data indicating "Turn on the power and then turn on the AI function" from "Turn on the power" and "Turn on the AI function".
[0136] In addition, the voice assistant server 3000 can use the "power on / off" and "AI mode on / off" functions of the new air conditioner B and the generated speech data "turn on the power and then turn on the AI function" to generate action data for executing the "power on" function and then executing the "AI mode on" function.
[0137] Figure 6 is a flowchart showing a method performed by a voice assistant server according to an embodiment of the present disclosure for generating speech data and action data related to a function of a new device that is different from a function of a pre-registered device.
[0138] In operation S600, the voice assistant server 3000 uses the NLG model to output a query for registering additional functions and generating or editing speech data. The voice assistant server 3000 can provide a graphical user interface (GUI) for registering functions of a new device 2900 and generating speech data to the user's device 2000 or the developer's device. The developer's device can be installed with a specific SDK for registering a new device, and receive the GUI from the voice assistant server 3000 through the SDK.
[0139] The voice assistant server 3000 may provide a guide text or guide voice data to the user, the user's device 2000, or the developer's device to register the functions of the new device 2900 and generate utterance data. The voice assistant server 3000 may generate a query for registering additional functions and generating utterance data using an NLG model trained to register functions and generate utterance data.
[0140] In addition, the voice assistant server 3000 may provide the user's device 2000 or the developer's device with a list of functions that are different from the functions of the pre-registered device 2000 among the functions of the new device 2900. The voice assistant server 3000 may provide the user's device 2000 or the developer's device with recommended utterance data related to at least some of the different functions.
[0141] In operation S610, the voice assistant server 3000 may interpret the response to the query using the NLU model. The voice assistant server 3000 may receive the user's response to the query from the user's device 2000, or receive the developer's response to the query from the developer's device. The voice assistant server 3000 may interpret the user's response or the developer's response using the NLU model trained to register functions and generate utterance data.
[0142] In addition, the voice assistant server 3000 may receive a user's response input from the user's device 2000 through a GUI provided for the user's device 2000, or receive a developer's response input from the developer's device through a GUI provided for the developer's device.
[0143] In operation 620, the voice assistant server 3000 generates speech data related to the functions of the new device 2900 based on the interpreted responses. The voice assistant server 3000 may generate speech data related to the functions of the new device 2900 using the interpreted user responses or the interpreted developer responses, and recommend the generated speech data. The voice assistant server 3000 may select some functions of the new device 2900 and generate speech data related to each selected function. In addition, the voice assistant server 3000 may select some functions of the new device 2900 and generate speech data related to a combination of the selected functions.
[0144] The voice assistant server 3000 may generate utterance data related to the function of the new device 2900 using an NLG model for generating utterance data based on the ID and attributes of the function of the new device 2900. For example, the voice assistant server 3000 may input data representing the ID and attributes of the function of the new device into the NLG model for generating utterance data and obtain utterance data output from the NLG model, but is not limited thereto.
[0145] The voice assistant server 3000 may select at least some of the utterance data generated based on the responses input by the user through the GUI and the responses input by the developer through the GUI. In addition, the voice assistant server 3000 may generate synonymous utterance data having the same meaning but different expressions from the generated utterance data.
[0146] In operation S630, the voice assistant server 3000 may generate action data using the generated utterance data. The voice assistant server 3000 may identify the function of the new device 2900 associated with the generated utterance data, and determine the execution order of the identified function to generate action data corresponding to the generated utterance data. The generated action data may match the utterance data and the proximate utterance data.
[0147] Fig. 7A A query output from a voice assistant server for generating speech data and action data related to a function of a new device according to an embodiment of the present disclosure is shown.
[0148] refer to Fig. 7A , the voice assistant server 3000 may provide a query for receiving a speech sentence related to a new function of the new device 2900, which is an air conditioner, to the user's device 2000 or the developer's device. For example, a text or query voice "say a speech sentence related to the automatic drying function" may be output from the user's device 2000 or the developer's device.
[0149] In another example, the voice assistant server 3000 may receive an utterance sentence “remove the odor of the air conditioner” input from the user's device 2000 or the developer's device.
[0150] The voice assistant server 3000 may modify "remove the odor of the air conditioner" to "remove the odor of the air conditioner". The voice assistant server 3000 may generate action data "current operation off + drying function on" corresponding to the modified utterance sentence "remove the odor of the air conditioner".
[0151] Figure 7B A query output for recommending utterance sentences to generate utterance data and action data related to functions of a new device according to an embodiment of the present disclosure is shown.
[0152] refer to Figure 7B , the voice assistant server 3000 may provide a text or voice for notifying a new function of the new device 2900 as an air conditioner to the user's device 2000 or the developer's device. For example, the voice assistant server 3000 may allow the output of the text or voice "the automatic drying function is a new function" from the user's device 2000 or the developer's device. In addition, the voice assistant server 3000 may generate a recommended utterance sentence related to the new function, namely the automatic drying function, and provide the text or voice representing the recommended utterance sentence to the user's device 2000 or the developer's device. For example, the voice assistant server 3000 may allow the output of the query "Should we register the utterance sentence "Execute the drying function when the air conditioner is turned off?" from the user's device 2000 or the developer's device.
[0153] In addition, the voice assistant server 3000 may receive an input of selecting a recommended utterance sentence from the user's device 2000 or the developer's device.
[0154] The voice assistant server 3000 may generate action data "Check if power off input + drying function on + air conditioner power off" corresponding to the recommended speech sentence "Execute drying function when air conditioner is turned off"
[0155] Figure 8 is a flowchart showing a method for expanding utterance data performed by the voice assistant server 3000 according to an embodiment of the present disclosure.
[0156] In operation S800, the voice assistant server 3000 may obtain synonymous utterance data related to the generated utterance data by inputting the generated utterance data into the AI model. The voice assistant server 3000 may obtain synonymous utterance data output from the AI model by inputting the generated utterance data into the AI model trained to generate utterance data similar to the utterance data. The AI model may be, for example, a model trained using an utterance sentence and a set of synonymous utterance data as learning data.
[0157] For example, Fig.9A As shown, when the speech sentence "remove the smell of the air conditioner" is input into the AI model, synonymous speech data such as "remove the odor of the air conditioner.", "The air conditioner stinks.", "It smells moldy.", "remove the moldy smell", etc. can be output from the AI model. The speech sentence "remove the odor of the air conditioner" input into the AI model can be set as a representative speech sentence. The representative speech sentence can be set by considering, but not exclusively considering, the user's usage frequency, grammatical accuracy, etc.
[0158] In addition, for example, Fig. 10A As shown, when the speech sentence "Execute the drying function when the air conditioner is turned off" is input into the AI model, synonymous speech data such as "Execute the drying function when you turn off the air conditioner.", "Deodorize when you turn off the air conditioner.", "Deodorize after turning off the air conditioner.", etc. can be output from the AI model. In addition, "Execute the drying function when the air conditioner is turned off" can be set as a representative speech sentence. Alternatively, one of the synonymous speech data output from the AI model, "Execute the drying function when you turn off the air conditioner", can be set as a representative speech sentence. The representative speech sentence can be set in consideration of, but not exclusively, the user's frequency of use, grammatical accuracy, etc.
[0159] In operation S810, the voice assistant server 3000 may map action data to utterance data and proximate utterance data. The voice assistant server 3000 may map action data corresponding to utterance data input into the AI model to proximate utterance data output from the AI model.
[0160] For example, Fig. 9B As shown, the representative speech sentence "Remove the smell of the air conditioner" and the synonymous speech data "Remove the odor of the air conditioner.", "The air conditioner stinks.", "It smells moldy.", "Remove the moldy smell" etc. can be mapped to the action data "Current operation is off --> Drying function is on".
[0161] In addition, for example, Fig. 10BAs shown, the representative speech sentence "Perform the drying function when the air conditioner is turned off" and the proximate speech data "Remove the odor of the air conditioner.", "Deodorize when you turn off the air conditioner.", "Deodorize after turning off the air conditioner." can be mapped to the action data "Check the reception of the power off input --> Drying function on --> Air conditioner power off."
[0162] Fig.11A Speech data according to an embodiment of the present disclosure is presented.
[0163] refer to Fig.11A , the utterance data may have a text format. For example, a representative utterance sentence "Turn on the TV for me" and proximate utterance data "Please turn on the TV.", "Turn on the TV." and "The TV is on" may be utterance data.
[0164] Fig. 11B Speech data according to another embodiment of the present disclosure is presented.
[0165] refer to Fig. 11B , the speech data may include speech parameters and speech sentences. The speech parameters may be the output values of the NLU model including intents and parameters. For example, the speech parameters included in the speech data may include the intent (action, function or command) "power on" and the parameter (object or device) "TV". Further, the speech sentences included in the speech data may include texts such as "Turn on the TV for me.", "Please turn on the TV.", "Turn on the TV." and "The TV is on". Although in Fig. 11B The utterance data is illustrated as including the utterance parameters and the utterance sentences, but is not limited thereto, and the utterance data may include only the utterance parameters.
[0166] Fig.13 is a block diagram of a voice assistant server according to an embodiment of the present disclosure.
[0167] refer to Fig.13 , the voice assistant server 3000 may include a communication interface 3100, a processor 3200, and a storage device 3300. The storage device 3300 may include a first voice assistant model 3310, at least one second voice assistant model 3320, an SDK interface module 3330, and a database (DB) 3340.
[0168] The communication interface 3100 communicates to send and receive data to and from the client device 1000, the device 2000, and the IoT cloud server 4000. For example, the communication interface 3100 may include one or more network hardware and software components for wired or wireless communication with the client device 1000, the device 2000, and the IoT cloud server 4000.
[0169] The processor 3200 controls the overall operation of the voice assistant server 3000. For example, the processor 3200 may control the functions of the voice assistant server 3000 by loading a program stored in the storage device 3300 into the memory of the voice assistant server 3000 and executing the loaded program.
[0170] The storage device 3300 may store a program for controlling the processor 3200, and store data related to the functions of the new device 2900. The storage device 3300 may include at least one type of storage medium, including a flash memory, a hard disk, a multimedia card micro memory, a card-type memory (e.g., a secure digital (SD) or extreme digital (XD) memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk.
[0171] The programs stored in the storage device 3300 may be classified into a plurality of modules according to functions, for example, into a first voice assistant model 3310, at least one second voice assistant model 3320, an SDK interface module 3330, and the like.
[0172] The first voice assistant model 3310 is a model for determining a target device related to the user's intention by analyzing the user's voice input. The first voice assistant model 3310 may include an automatic speech recognition (ASR) model 3311, a first NLU model 3312, a first NLG model 3313, a device determination module 3314, a function comparison module 3315, a speech data acquisition module 3316, an action data generation module 3317, and a model updater 3318.
[0173] The ASR model 3311 converts the speech signal into text by performing ASR. The ASR model 3311 may utilize a predefined model such as an acoustic model (AM) or a language model (LM) to perform ASR that converts the speech signal into computer-readable text. When an acoustic signal with noise is received from the client device 110, the ASR model 3311 may obtain the speech signal by eliminating the noise from the received acoustic signal and perform ASR on the optimized speech signal.
[0174] The first NLU model 3312 analyzes the text and determines a first intent related to the user's intent based on the analysis result. The first NLU model 3312 may be a model trained to interpret the text and obtain a first intent corresponding to the text. The intent may be information indicating the intent of the user's utterance contained in the text.
[0175] The device determination model 3314 may use the first NLU model 3312 to perform syntactic analysis or semantic analysis to determine the user's first intent from the converted text. In an embodiment of the present disclosure, the device determination model 3314 may use the first model 3312 to parse the converted text into units of morphemes, words, or phrases, and use the language features (e.g., syntactic elements) of the parsed morphemes, words, or phrases to infer the meaning of the words extracted from the parsed text. The device determination model 3314 may determine the first intent corresponding to the inferred meaning of the word by comparing the inferred meaning of the word with the predefined intent provided from the first NLU model 3312. The device determination model 3314 may determine the type of the target device based on the first intent. In an embodiment of the present disclosure, the device determination model 3314 may determine the type of the target device by using the first intent obtained using the first NLU model 3312. The device determination model 3314 provides the parsed text and the target device information to the second voice assistant model 3320. In an embodiment of the present disclosure, the device determination model 3314 may provide the identification information (e.g., device ID) of the determined target device together with the parsed text to the second voice assistant model 3320.
[0176] The first NLG model 3313 may register a function of the new device 2900 and generate a query message for generating or editing utterance data.
[0177] The function comparison module 3315 may compare the functions of the pre-registered device 2000 with the functions of the new device 2900. The function comparison module 3315 may determine whether the functions of the pre-registered device 2000 are the same as or similar to the functions of the new device 2900. The function comparison module 3315 may identify any functions of the new device 2900 that are the same as or similar to the functions of the pre-registered device 2000.
[0178] The function comparison module 3315 may identify a function name indicated as supported by the new device 2900 from the technical specifications of the new device 2900, and determine whether the identified name is identical or similar to a function name supported by the pre-registered device 2000. In this case, the database 3340 may store information about names and synonyms indicating certain functions, and determine whether the function of the pre-registered device 2000 and the function of the new device 2900 are identical or similar to each other based on the stored information about the synonyms.
[0179] In addition, the function comparison module 3315 may determine whether the functions are identical or similar to each other by referring to the utterance data stored in the database 3340. The function comparison module 3315 may determine whether the function of the new device 2900 is identical or similar to the function of the preregistered device 2000 using the utterance data related to the function of the preregistered device 2000. In this case, the function comparison module 3315 may interpret the utterance data using the first NLU model, and determine whether the function of the new device 2900 is identical or similar to the function of the preregistered device 2000 based on the meaning of the words included in the utterance data.
[0180] Function comparison module 3315 may determine whether a single function of pre-registered device 2000 is the same or similar to a single function of new device 2900. Function comparison module 3315 may determine whether a set of functions of pre-registered device 2000 is the same or similar to a set of functions of new device 2900.
[0181] The utterance data acquisition module 3316 may acquire utterance data related to the function of the new device 2900. The utterance data acquisition module 3316 may extract utterance data corresponding to functions of the pre-registered device 2000 that are determined to be the same as or similar to the functions of the new device 2900 from the utterance data database 3341.
[0182] The utterance data acquisition module 3316 may extract, from the utterance data database 3341, utterance data corresponding to a function group determined to be the same as or similar to the function of the new device 2900 among the function groups of the pre-registered device 2000. In this case, the utterance data corresponding to the function of the pre-registered device 2000 and the utterance data corresponding to the function group of the pre-registered device 2000 may be pre-stored in the utterance data database 3341.
[0183] The speech data acquisition module 3316 can edit the functions or function groups determined to be the same or similar, and generate speech data corresponding to the edited functions. The speech data acquisition module 3316 can combine the functions determined to be the same or similar and generate speech data corresponding to the combined functions. In addition, the speech data acquisition module 3316 can combine the functions and function groups determined to be the same or similar, and generate speech data corresponding to the combined functions. In addition, the speech data acquisition module 3316 can delete some functions in the function group determined to be the same or similar, and generate speech data corresponding to the function group from which some functions are deleted.
[0184] The speech data acquisition module 3316 can expand the speech data. The speech data acquisition module 3316 can generate synonymous speech data with the same meaning but different expressions from the extracted or generated speech data by modifying the expression mode of the extracted or generated speech data.
[0185] The speech data acquisition module 3316 can use the first NLG model 3313 to output a query for registering additional functions, generating or editing speech data. The speech data acquisition module 3316 can provide a guide text or guide voice data to the user, the user's device 2000, or the developer's device for registering the functions of the new device 2900 and generating speech data. The speech data acquisition module 3316 can provide the user's device 2000 or the developer's device with a list of functions that are different from the functions of the pre-registered device 2000 in the functions of the new device 2900. The speech data acquisition module 3316 can provide the user's device 2000 or the developer's device with recommended speech data related to at least some different functions.
[0186] The speech data acquisition module 3316 can use the first NLU model 3312 to interpret the response to the query. The speech data acquisition module 3316 can generate speech data related to the function of the new device 2900 based on the interpreted response. The speech data acquisition module 3316 can use the interpreted user response or the interpreted developer's response to generate speech data related to the function of the new device 2900, and recommend the generated speech data. The speech data acquisition module 3316 can select some functions of the new device 2900 and generate speech data related to each selected function. The speech data acquisition module 3316 can select some functions of the new device 2900 and generate speech data related to the combination of the selected functions. The speech data acquisition module 3316 can use the first NLG model 3313 to generate speech data related to the function of the new device 2900 based on the identifier and attributes of the function of the new device 2900.
[0187] The action data generation module 3317 can generate action data for the new device 2900 based on the same or similar functions and speech data. For example, when the function corresponding to the speech data is a single function, the action data generation module 3317 can generate action data including detailed operations representing the single function. In another example, when the function corresponding to the speech data is a function group, the action data generation module 3317 can generate detailed operations representing the functions in the group and the execution order of the detailed operations. The action data generation module 3317 can generate action data using speech data generated in connection with the new functions of the new device 2900. The action data generation module 3317 can identify the new functions of the new device 2900 related to the speech data, and determine the execution order of the identified functions to generate action data corresponding to the generated speech data. The generated action data can be mapped to the speech data and synonymous speech data.
[0188] The model updater 3318 may generate or update the second voice assistant model 3320 associated with the new device 2900 using the utterance data and the action data. The model updater 3318 may generate or update the second voice assistant model 3320 associated with the new device 2900 using the utterance data associated with the function of the new device 2900 corresponding to the function of the pre-registered device 2000, the newly generated utterance data associated with the function of the new device 2900, and the extended utterance data and action data. The model updater 3318 may accumulate the utterance data and action data associated with the new device 2900 and store the results in the utterance data database 3341 and the action data database 3342. In addition, the model updater 3318 may generate or update a CAN, which is a package type database included in the action plan management model 3323.
[0189] The second voice assistant model 3320 is device-specific and can determine the operation to be performed by the target device corresponding to the user's voice input. The second voice assistant model 3320 may include a second NLU model 3321, a second NLG model 3322, and an action plan management model 3323. The voice assistant server 3000 may include a second voice assistant model 3320 for each device type.
[0190] The second NLU model 3321 is an NLU model for a specific device to analyze the text and determine a second intent related to the user's intent based on the analysis result. The second NLU model 3321 can interpret the user's input voice in consideration of the function of the device. The second NLU model 3321 can be a model trained to interpret the text and obtain the second intent corresponding to the text.
[0191] The second NLG model 3322 may be an NLG model for a specific device for generating a query message required to provide a voice assistant service to a user. The second NLU model 3322 may generate a natural language for a conversation with a user taking into account the specific functions of the device.
[0192] The action plan management model 3323 is a model for a specific device, which is used to determine the operation to be performed by the target device corresponding to the user's voice input. The action plan management model 3323 can plan the operation information to be performed by the new device 2900 in consideration of the specific functions of the new device 2900.
[0193] The action plan management model 3323 can select the detailed operation to be performed by the new device 2900 based on the interpreted speech emitted by the user, and plan the execution order of the selected detailed operation. The action plan management model 3323 can use the plan result to obtain the operation information about the detailed operation to be performed by the new device 2900. The operation information can be information related to the detailed operation to be performed by the device, the association between the detailed operations, and the execution order of the detailed operations. The operation information can include, for example, the function executed by the new device 2900 to perform the detailed operation, the execution order of the function, the input value required to perform the function, and the output value output as a result of performing the function.
[0194] The action plan management model 3323 can manage information about multiple detailed operations and information about the relationship between the multiple detailed operations. The correlation between each of the multiple detailed operations and another detailed operation can include information about another detailed operation to be performed before performing one detailed operation to perform the one detailed operation.
[0195] The action plan management model 3323 may include a CAN, a database in a package format indicating the operations of the device and the dependencies between the operations. The CAN may include the functions to be performed by the device to perform specific operations, the execution order of the functions, the input values required to perform the functions, and the output values output as a result of performing the functions, and may be implemented in an ontology diagram including knowledge triples indicating concepts and the relationships between concepts.
[0196] The SDK interface module 3330 may send data to or receive data from the client device 1000 or the developer's device through the communication interface 3100. The client device 1000 or the developer's device may have a specific SDK for registering a new device installed therein, and receive a GUI from the voice assistant server 3000 through the SDK. The processor 3200 may provide a GUI for registering a new device 2900 and generating speech data to the user's device 2000 or the developer's device through the SDK interface module 3330. The processor 3200 may receive a user's response input through the GUI provided for the user's device 2000 from the user's device 2000 through the SDK interface module 3330, or receive a developer's response input through the GUI provided for the developer's device from the developer's device through the SDK interface module 3330. The SDK interface module 3330 may send data to or receive data from the IoT cloud server 4000 through the communication interface 3100.
[0197] The database 3340 may store various information of the voice assistant service. The database 3340 may include a speech data database 3341 and an action data database 3342.
[0198] The utterance data database 3341 may store utterance data related to functions of the client device 1000 , the device 2000 , and the new device 2900 .
[0199] The action data database 3342 may store action data related to functions of the client device 1000, the device 2000, and the new device 2900. The utterance data stored in the utterance data database and the action data stored in the action data database 3342 may be mapped to each other.
[0200] Fig.14 is a block diagram of a voice assistant server according to another embodiment of the present disclosure.
[0201] refer to Fig.14 , the voice assistant server 3000 may include a second voice assistant model 3320. In this case, the second voice assistant model 3320 may include a plurality of second NLU models 3324, 3325, and 3326. The plurality of second NLU models 3324, 3325, and 3326 may be NLU models dedicated to each of various types of devices.
[0202] Fig.15 is a conceptual diagram showing an action plan management model according to an embodiment of the present disclosure.
[0203] refer to Fig.15 , the action planning base station 3323 may include a speaker CAN 212 , a mobile CAN 214 and a TV CAN 216 .
[0204] The speaker CAN 212 may include information on detailed operations of speaker control, media playback, weather, and TV control, and may include an action plan storing a concept corresponding to each detailed operation in a package format.
[0205] The mobile CAN 214 may include information on detailed operations of a social networking service (SNS), mobile control, a map, and question and answer (Q&A), and may include an action plan storing a concept corresponding to each detailed operation in a package format.
[0206] The TV CAN 216 may include information about detailed operations of shopping, media playback, education, and television playback, and may include an action plan storing a concept corresponding to each detailed operation in a package format. In an embodiment of the present disclosure, a plurality of packages included in each of the speaker CAN 212, the mobile CAN 214, and the TV CAN 216 may be stored in a function registry, which is a constituent element in the action plan management model 3323.
[0207] In an embodiment of the present disclosure, when the voice assistant server 3000 determines a detailed operation corresponding to a second intent and parameter determined by interpreting a text converted from a voice input using a second NLU model, the action plan management model 3323 may include a policy registry. The policy registry may include reference information for determining an action plan when multiple action plans are associated with the text. In an embodiment of the present disclosure, the action plan management model 3323 may include a subsequent registry in which information about subsequent operations is stored to suggest subsequent operations to the user in a specified situation. Subsequent operations may include, for example, subsequent utterances.
[0208] In an embodiment of the present disclosure, the action plan management model 3323 may include a layout registry in which layout information output by the target device is stored.
[0209] In an embodiment of the present disclosure, the action plan management model 3323 may include a vocabulary registry, in which vocabulary information included in the encapsulation information is stored. In an embodiment of the present disclosure, the action plan management model 3323 may include a conversation registry, in which information about conversations or interactions with users is stored.
[0210] Fig.16 An encapsulated database stored in an action plan management model according to an embodiment of the present disclosure is presented.
[0211] refer to Fig.16 The encapsulation database stores detailed operations and associated information about concepts corresponding to the detailed operations. The encapsulation database may be implemented in the form of CAN. The encapsulation database may store a plurality of encapsulations 230, 240, and 250. The encapsulation database may store detailed operations for performing operations related to the user's voice input, input parameters required for the detailed operations, and result values output in the form of CAN.
[0212] The encapsulation database may store information related to the operation of each device. Fig.16 As shown, the encapsulation device may store a plurality of encapsulations 230, 240, and 250 related to operations performed by a specific device (e.g., TV). In an embodiment of the present disclosure, a package (e.g., package A 230) may correspond to an application. A package may include at least one detailed operation and at least one concept for performing a specified function. For example, package A 230 may include a detailed operation 231a and a concept 231b corresponding to the detailed operation 231a, and package B 240 may include a plurality of detailed operations 241a, 242a, and 243a and a plurality of concepts 241b, 242b, and 243b corresponding to the detailed operations 241a, 242a, and 243a, respectively.
[0213] The action plan management model 210 can use the encapsulation stored in the encapsulation database to generate an action plan for performing an operation related to the user's voice input. For example, the action plan management model 210 can use the encapsulation stored in the encapsulation database to generate an action plan. For example, the action plan management model 210 can generate an action plan 260 related to the operation to be performed by the device by using the detailed operation 231a and concept 231b of encapsulation A 230, multiple detailed operations 241a, 242a and 243a of encapsulation B 240, some of multiple concepts 241b, 242b and 243b (i.e., operations 241a and 241b and concepts 241b and 243b) and detailed operations 251a and concept 251b of encapsulation C 250. Of course, any combination of operations and concepts can be selected from the operations and concepts of encapsulation (230, 240, 250) to generate an action plan 260.
[0214] Fig.17 is a block diagram of an IoT cloud server according to an embodiment of the present disclosure.
[0215] refer to Fig.17 , the IoT cloud server 4000 may include a communication interface 4100, a processor 4200, and a storage device 4300. The storage device 4300 may include an SDK interface module 4310, a function comparison module 4320, a device registration module 4330, and a database 4340. The database 4340 may include a device function database 4341 and an action data database 4342.
[0216] The communication interface 4100 communicates with the client device 1000, the device 2000, and the voice assistant server 3000. The communication interface 4100 may include one or more network hardware and software components for wired or wireless communication with the client device 1000, the device 2000, and the voice assistant server 3000.
[0217] The processor 4200 controls the overall operation of the IoT cloud server 4000. For example, the processor 4200 may control the functions of the IoT cloud server 4000 by loading a program stored in the storage device 4300 into the memory and executing the program loaded into the memory.
[0218] The storage device 4300 may store a program that provides control to the processor 4200, and store data related to the functions of the device 2000. The storage device 4300 may include at least one type of storage medium, including a flash memory, a hard disk, a multimedia card micro memory, a card-type memory (e.g., an SD or XD memory), a RAM, a SRAM, a ROM, an EEPROM, a PROM, a magnetic memory, a magnetic disk, and an optical disk.
[0219] The programs stored in the storage device 4300 may be divided into a plurality of modules according to functions, such as an SDK interface module 4310 , a function comparison module 4320 , a device registration module 4330 , and the like.
[0220] The SDK interface module 4310 may transmit or receive data to or from the voice assistant server 3000 through the communication interface 4100. The processor 4200 may provide the function information of the device 2000 to the voice assistant server 3000 through the SDK interface module 4310.
[0221] When the function comparison module 4320 is included in the IoT cloud server 4000, the function comparison module 4320 can be used as the aforementioned function comparison model 3315 of the voice assistant server 3000 according to a client-server or cloud-based model.
[0222] In this case, the function comparison module 4320 may compare the functions of the pre-registered device 2000 with the functions of the new device 2900. The function comparison module 4320 may determine whether the functions of the pre-registered device 2000 are the same as or similar to the functions of the new device 2900. The function comparison module 4320 may identify any functions of the new device 2900 that are the same as or similar to the functions of the pre-registered device 2000.
[0223] The function comparison module 4320 may identify a function name indicated as supported by the new device 2900 from the technical specifications of the new device 2900, and determine whether the identified name is identical or similar to a function name supported by the pre-registered device 2000. In this case, the database 4340 may store information about names and synonyms indicating certain functions, and determine whether the function of the pre-registered device 2000 and the function of the new device 2900 are identical or similar to each other based on the stored information about the synonyms.
[0224] In addition, the function comparison module 4320 may determine whether the functions are identical or similar to each other by referring to the utterance data stored in the database 4340. The function comparison module 3315 may determine whether the functions of the new device 2900 are identical or similar to the functions of the preregistered device 2000 using the utterance data related to the functions of the preregistered device 2000. The function comparison module 4320 may determine whether a single function of the preregistered device 2000 is identical or similar to a single function of the new device 2900. The function comparison module 4320 may determine whether a group of functions of the preregistered device 2000 is identical or similar to a group of functions of the new device 2900.
[0225] The device registration module 4330 can register a device for the voice assistant service. When a new device 2900 is identified, the device registration module 4330 can receive information about the functions of the new device 2900 from the voice assistant server 3000 and register the received information in the database 4330. The information about the functions of the new device 2900 may include, for example, functions supported by the new device 2900, action data related to the functions, etc., but is not limited thereto.
[0226] The database 4340 may store device information required for the voice assistant service. The database 4340 may include a device function database 4341 and an action data database 4342. The device action data database 4340 may store function information of the client device 1000, the device 2000, and the new device 2900. The function information may include information about the ID value of the device function and the name and attribute of the function, but is not limited thereto. The action data database 4342 may store action data related to the functions of the client device 1000, the device 2000, and the new device 2900.
[0227] Fig.18 is a block diagram of a client device according to an embodiment of the present disclosure.
[0228] refer to Fig.18 In an embodiment of the present disclosure, the client device 1000 may include an input module 1100 , an output module 1200 , a processor 1300 , a memory 1400 , and a communication interface 1500 . The memory 1400 may include an SDK module 1420 .
[0229] The device 2000 may operate as the client device 1000, or the new device 2900 may operate as the client device 1000 after pre-registration. The device 2000 or the new device 2900 may also include the following: Fig.18 Parts shown.
[0230] The input module 1100 refers to hardware and / or software that allows a user to input data to control the client device 1000. For example, the input module 1100 may include a keypad, a dome switch, a touch pad (capacitive, resistive, infrared detection type, surface acoustic wave type, integral strain gauge type, piezoelectric effect type), a scroll wheel, a scroll wheel switch, a graphical user interface (GUI) displayed on a display, an audio user interface provided to the user by audio, etc., but is not limited thereto.
[0231] The input module 1100 may receive user input to register a new device 2900 .
[0232] The output module 1200 may output an audio signal, a video signal, or a vibration signal, and the output module 1210 may include at least one of a display, a sound output, or a vibration motor. The input module 1100 and the output module 1200 may be combined into an input / output interface, such as a touch screen display for receiving user input and displaying output information to the user. Of course, any combination of software and hardware components may be provided to perform input / output functions between the client device 1000 and the user.
[0233] The processor 1300 controls the overall operation of the client device 1000. For example, the processor 1300 may execute a program loaded from the memory 1400 to generally control the user input module 1100, the output module 1200, the memory 1400, and the communication interface 1500.
[0234] The processor 1300 may request an input from the user to register a function of the new device 2900. The processor 1300 may perform an operation of registering the new device 2900 with the voice assistant server 300 by controlling the SDK module 1420.
[0235] The processor 1300 may receive a query message for generating and editing utterance data related to the functions of the new device 2900 from the voice assistant server 3000 and output the query message. The processor 1300 may provide the user with a list of functions different from the functions of the pre-registered device 1300 among the functions of the new device 2900. The processor 1300 may provide the user with recommended utterance data related to at least some functions of the new device 2900 via the output module 1200.
[0236] The processor 1300 may receive a user's response to the query message via the input module 1100. The processor 1300 may provide the voice assistant server 3000 with the user's response for generating utterance data and action data related to the function of the new device 2900.
[0237] The communication interface 1500 may include one or more hardware and / or software communication components that allow communication with the voice assistant server 3000, the IoT cloud server 4000, the device 2000, and the new device 2900. For example, the communication interface 1500 may include a short-range communication module (infrared, WiFi, etc.), a mobile communication module (4G, 5G, etc.), and a broadcast receiver.
[0238] The short-range communication module may include a Bluetooth communication module, a Bluetooth low energy (BLE) communication module, an NFC module, a wireless LAN (WLAN) (e.g., Wi-Fi), a communication module, a Zigbee communication module, an IrDA communication module, a WFD communication module, a UWB communication module, an Ant+ communication module, etc., but is not limited thereto.
[0239] The mobile communication module transmits an RF signal to at least one of the following or receives an RF signal from at least one of the following in a mobile communication network: a base station, an external terminal or a server. The RF signal may include a voice call signal, a video call signal or different types of data related to the transmission / reception of text / multimedia messages.
[0240] The broadcast receiver receives a broadcast signal and / or broadcast related information from the outside on a broadcast channel. The broadcast channel may include a satellite channel or a terrestrial channel. According to an embodiment, the client device 1000 may not include a broadcast receiver.
[0241] The memory 1400 may store programs for processing and control of the processor 1300 , and store data input to or output from the device 1000 .
[0242] Storage device 1400 may include at least one type of storage medium, including flash memory, hard disk, multimedia card micro memory, card-type memory (e.g., SD or XD memory), RAM, SRAM, ROM, EEPROM, PROM, magnetic storage, magnetic disk and optical disk.
[0243] The programs stored in the memory 1400 may be classified into a plurality of modules according to functions, such as an SDK module 1420 , a UI module, a touch screen module, a notification module, and the like.
[0244] The SDK module 1420 may be executed by the processor 1300 to perform operations required to register the new device 2900. The SDK module 1420 may be downloaded from the voice assistant server 300 and installed in the client device 1000. The SDK module 1420 may output a GUI for registering the new device 2900 on the screen of the client device 1000. When the client device 1000 does not include any display module, the SDK module 1420 may allow the client device 1000 to output a voice message for registering the new device 2900. The SDK module 1420 may allow the client device 1000 to receive a response from the user and provide the response to the voice assistant server 3000.
[0245] Embodiments of the present disclosure may be implemented in the form of a computer-readable recording medium including computer-executable instructions (such as a program module executed by a computer). Computer-readable recording media may be any available medium accessible to a computer, including volatile media, non-volatile media, removable media, and non-removable media. Computer-readable recording media may also include computer storage media and communication media. Volatile, non-volatile, removable and non-removable media may be implemented by any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may include other data of modulated data signals, such as computer-readable instructions, data structures, or program modules.
[0246] In the specification, the term "module" may refer to a hardware component such as a processor or a circuit system and / or a software component executed by a hardware component such as a processor.
[0247] Several embodiments of the present disclosure have been described above, but those of ordinary skill in the art will understand and appreciate that various modifications may be made without departing from the scope of the present disclosure. Therefore, it will be apparent to those of ordinary skill in the art that the present disclosure is not limited to the embodiments of the present disclosure described, but may include not only the attached claims but also equivalents. For example, an element described in singular form may be implemented as distributed, and an element described in distributed form may be implemented as a combination.
[0248] The scope of the present disclosure is defined by the appended claims, and it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents.
Claims
1. A method for registering a new device for a voice assistant service performed by a server, the method include: obtaining a first technical specification indicating a first function of a pre-registered device and a second technical specification indicating a second function of a new device; comparing a first function of the pre-registered device with a second function of the new device based on the first technical specification and the second technical specification; Based on the comparison result, identifying the second function of the new device that is different from the first function of the pre-registered device as an additional function; Use a natural language generation (NLG) model to output a query for registering additional features; interpreting a response to the query using a natural language understanding (NLU) model; generating, based on the interpreted responses, utterance data related to the additional functionality of the new device; generating first action data for the new device using the generated speech data; as well as storing the speech data and the first action data in association with the new device, The first action data includes data related to a series of detailed operations of the new device corresponding to the speech data.
2. The method according to claim 1, in, Outputting the query includes providing a user with a guide text or guide voice data for registering the additional function of the new device to a client device.
3. The method according to claim 2, further comprising: include: The client device is provided with a list of additional functions of the new device that are different from the functions of the pre-registered device.
4. The method according to claim 2, in, Outputting the query includes providing a query to the client device for receiving an utterance sentence related to the additional function of the new device.
5. The method according to claim 2, in, Outputting the query includes providing a query to the client device for informing the client device of the additional functionality of the new device.
6. The method according to claim 1, in, Generating the speech data includes: obtaining similar speech data related to the generated speech data by inputting the generated speech data into an artificial intelligence model.
7. The method according to claim 1, further comprising: include: Based on the comparison result, identifying a third function of the pre-registered device that matches a fourth function of the new device as a matching function; Acquiring pre-registered speech data related to the matching function; generating second action data for the new device based on the matching function and the pre-registered speech data; as well as The pre-registered utterance data and the second motion data are stored in association with the new device.
8. A server for registering a new device for a voice assistant service, the server include: Communication interface; a memory storing a program including one or more instructions; as well as A processor, wherein the processor is configured to execute the one or more instructions of the program stored in the memory to control the server to perform the following operations: obtaining a first technical specification indicating a first function of a pre-registered device and a second technical specification indicating a second function of a new device; comparing a first function of the pre-registered device with a second function of the new device based on the first technical specification and the second technical specification; Based on the comparison result, identifying the second function of the new device that is different from the first function of the pre-registered device as an additional function; Use a natural language generation (NLG) model to output a query for registering additional features; interpreting a response to the query using a natural language understanding (NLU) model; generating, based on the interpreted responses, utterance data related to the additional functionality of the new device; generating first action data for the new device using the generated speech data; as well as storing the speech data and the first action data in association with the new device, The first action data includes data related to a series of detailed operations of the new device corresponding to the speech data.
9. The server according to claim 8, in, The processor executing the one or more instructions is further configured to provide a user with guidance text or guidance voice data for registering the additional function of the new device to a client device.
10. The server according to claim 9, in, The processor executing the one or more instructions is further configured to provide the client device with a list of additional functions of the new device that are different from the functions of the pre-registered device.
11. The server according to claim 9, in, The processor executing the one or more instructions is further configured to provide a query to the client device for receiving an utterance sentence related to the additional functionality of the new device.
12. The server according to claim 9, in, The processor executing the one or more instructions is further configured to provide a query to the client device to inform the client device of the additional functionality of the new device.
13. The server according to claim 8, in, The processor that executes the one or more instructions is also configured to: obtain similar speech data related to the generated speech data by inputting the generated speech data into an artificial intelligence model.
14. The server according to claim 8, in, The processor executing the one or more instructions is further configured to: Based on the comparison result, identifying a third function of the pre-registered device that matches a fourth function of the new device as a matching function; Acquiring pre-registered speech data related to the matching function; generating second action data for the new device based on the matching function and the pre-registered speech data; as well as The pre-registered utterance data and the second motion data are stored in association with the new device.
15. A computer-readable recording medium having a program recorded thereon, causing a computer to execute the method according to any one of claims 1 to 7.