Speech processing method and apparatus therefor
The voice processing method and device address the challenge of operating multiple devices by converting and completing voice commands, enhancing recognition accuracy and efficiency.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2019-10-01
- Publication Date
- 2026-07-29
AI Technical Summary
Existing technologies fail to determine which electronic device to operate when voice commands included in a user's spoken voice can be processed by multiple devices, leading to inefficiencies and inaccuracies in voice recognition.
A voice processing method and device that converts user spoken voice into text, performs grammatical and semantic analysis, identifies domains and intents, determines completeness of the voice command, generates query voices to complete incomplete commands, and operates the appropriate electronic device.
Improves speech recognition performance by completing incomplete voice commands into complete ones, enabling rapid and accurate operation of the intended device while optimizing processor resources.
Smart Images

Figure 112019100657171-PAT00001_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a voice processing method and a voice processing device, and more specifically, to a voice processing method and a voice processing device in which, when a plurality of electronic devices can process a voice command included in a user's spoken voice, a response spoken voice corresponding to a query spoken voice fed back to the user is received, and one of the electronic devices to be operated is determined and operated. Background Technology
[0002] With the advancement of technology, various services applying speech recognition technology are being introduced in many fields recently. Speech recognition technology can be defined as a series of processes that understand human speech and convert it into text information that a computer can handle, and speech recognition services utilizing this technology may include a series of processes that recognize a user's voice and provide appropriate corresponding services.
[0003] Prior Art 1 discloses a voice conversation device having a function for conversing with a conversation partner, comprising: a voice recognition means for voice recognition of a conversation partner’s speech; a conversation control means for controlling a conversation with a conversation partner according to the recognition result of the voice recognition means; an image recognition means for image recognition of a conversation partner’s face; and a tracking control means for tracking the presence of a conversation partner according to either or both of the recognition result of the image recognition means and the recognition result of the voice recognition means, wherein the conversation control means controls the conversation to continue in accordance with the tracking by the tracking control means.
[0004] Prior art 2 discloses a method, system, and non-transient computer-readable recording medium for providing conversation services using an autonomous behavioral robot that provides a conversation service with a user based on at least one of a person attribute regarding the user and the reliability of the person attribute.
[0005] The aforementioned prior art 1 and prior art 2 disclose content that enables smooth voice conversation with a conversation partner, but there is a problem in that it is impossible to determine which electronic device to operate when voice commands included in the user's spoken voice can be processed by multiple electronic devices.
[0006] The aforementioned background technology is technical information that the inventor possessed for the derivation of the present invention or acquired during the process of deriving the present invention, and it cannot be considered as prior art disclosed to the general public prior to the filing of the present invention. Prior art literature
[0007] Prior Art 1: Korean Registered Patent Publication No. 10-1057705 (Aug. 11, 2011) Prior Art 2: Korean Registered Patent Publication No. 10-1985793 (May 29, 2019) The problem to be solved
[0008] One objective of the present invention is to solve the problem of the prior art in which it is impossible to determine which electronic device to operate when voice commands included in a user's spoken voice can be processed by multiple electronic devices.
[0009] One objective of the present invention is to receive a response speech voice corresponding to a query speech voice fed back to the user when an incomplete speech voice is received from a user, and to complete it into a complete speech voice.
[0010] One objective of the present invention is to receive an incomplete spoken voice from a user, receive a response spoken voice corresponding to a query spoken voice fed back to the user, complete it into a complete spoken voice, and determine and operate one of a plurality of electronic devices.
[0011] One objective of the present invention is to solve the problem of the prior art, where it is impossible to determine which electronic device to operate when voice commands included in a user's spoken voice can be processed by multiple electronic devices, while using optimal process resources. means of solving the problem
[0012] A voice processing method according to one embodiment of the present invention may include the step of receiving a response voice corresponding to a query voice feedbacked to the user, determining one of the electronic devices to be operated, and operating the device, when a plurality of electronic devices can process a voice command included in a user's voice speech.
[0013] Specifically, a voice processing method according to one embodiment of the present invention may include: a step of converting a user’s spoken voice containing a voice command into user spoken text; a step of performing grammatical analysis or semantic analysis on the user spoken text to search for a domain to which the user spoken text belongs and an intent possessed by the user spoken text, and searching for one or more named entities as a result of named entity recognition included in the user spoken text; a step of determining whether the user’s spoken voice is a complete spoken voice or an incomplete spoken voice in response to the search results for the domain, intent, and named entities; a step of generating a query spoken voice requesting information to complete the incomplete spoken voice into a complete spoken voice and providing feedback to the user if the user’s spoken voice is an incomplete spoken voice; and a step of receiving a user’s response spoken voice corresponding to the query spoken voice to complete the complete spoken voice.
[0014] Through the voice processing method according to the present embodiment, when an incomplete spoken voice is received from a user, a response spoken voice corresponding to the query spoken voice fed back to the user is received to complete it into a complete spoken voice, and one of a plurality of electronic devices is selected and operated to perform rapid and accurate voice recognition processing.
[0015] Additionally, the conversion step may include the step of converting a user’s spoken voice, which includes a voice command instructing the operation of at least one of a plurality of electronic devices, into user spoken text.
[0016] Additionally, the searching step may include a domain specifying the type of electronic device that the user intends to operate, an intention indicating how to operate the electronic device, and a step of searching for an entity name including a noun or number having a unique meaning appearing in the user's utterance text.
[0017] Additionally, the determining step may include the step of determining a required slot from the user utterance text, comprising a first slot related to a domain, a second slot related to an intention, and a third slot related to a quantity entity name; the step of determining the user's utterance as a complete utterance if the required slot in the user utterance text is filled with entity names; and the step of determining the user's utterance as an incomplete utterance if there is a required slot in the user utterance text where entity names are missing.
[0018] Additionally, the feedback step may include a step of determining a slot in the user utterance text where an entity name is missing, based on a pre-established state table containing a slot in the user command text that is required according to the domain and intent and a query text to be requested corresponding to the required slot; a step of generating a query text corresponding to the missing slot based on the state table; and a step of converting the generated query text into a query utterance voice and providing feedback to the user.
[0019] In addition, in the voice processing method according to the present embodiment, among the slots required in the user utterance text, the first slot is related to a domain, among the slots required in the user utterance text, the second slot is related to an intention, and among the slots required in the user utterance text, the third slot is related to a quantity entity name, and the step of determining the slot where the entity name is missing may include the step of determining the slot where the entity name is missing in the order of the first slot, the second slot, and the third slot among the slots required in the user utterance text.
[0020] Additionally, the step of generating query text may include: generating query text related to a domain when an entity name is missing in the first slot; generating query text related to an intent when an entity name is missing in the second slot after the entity name is filled in the first slot; and generating query text related to a slot where a quantity entity name is missing when an entity name is missing in the third slot after the entity name is filled in the second slot.
[0021] Additionally, the completing step may include a step of filling in the required slots where the entity name is missing with the entity name through feedback of the query utterance voice and the repeated reception of the user's response utterance voice, and a step of completing the user's utterance voice into a complete utterance voice when the required slots where the entity name is missing are all filled with the entity name.
[0022] In addition, the voice processing method according to the present embodiment may further include the step of determining and operating one of a plurality of electronic devices in response to a command included in the complete voice of a user when the user's voice is completed as a complete voice.
[0023] A voice processing device according to one embodiment of the present invention is a voice processing device comprising one or more processors, wherein the one or more processors convert a user’s spoken voice containing a voice command into user spoken text, perform grammatical analysis or semantic analysis on the user spoken text to search for a domain to which the user spoken text belongs and an intent having the user spoken text, search for one or more named entities as a result of named entity recognition included in the user spoken text, determine whether the user’s spoken voice is a complete spoken voice or an incomplete spoken voice in response to the search results for the domain, intent, and named entities, and if the user’s spoken voice is an incomplete spoken voice, generate a query spoken voice requesting information to complete the incomplete spoken voice into a complete spoken voice and provide feedback to the user, and receive a user’s response spoken voice corresponding to the query spoken voice to complete the complete spoken voice.
[0024] Additionally, one or more processors may be configured to convert a user’s spoken voice, which includes a voice command instructing the operation of at least one of a plurality of electronic devices, into user spoken text when converted into user spoken text.
[0025] Additionally, one or more processors may be configured to search for an entity name that includes a domain, an intent, a domain specifying the type of electronic device the user intends to operate during entity search, an intent indicating how to operate the electronic device, and a noun or number having a unique meaning appearing in the user's utterance text.
[0026] Additionally, one or more processors may be configured to determine, when determining whether a speech is complete or incomplete, a required slot including a first slot related to a domain, a second slot related to an intention, and a third slot related to a quantity entity name from the user speech text, and to determine the user's speech as complete speech if the required slot in the user speech text is filled with entity names, and to determine the user's speech as incomplete speech if there is a required slot in the user speech text where entity names are missing.
[0027] Additionally, one or more processors may be configured to determine, upon feedback to the user, slots where entity names are missing among the slots required in the user utterance text based on a pre-established state table containing slots required in the user command text according to the domain and intent and query text to be requested corresponding to the required slots, generate query text corresponding to the missing slots based on the state table, and convert the generated query text into query utterance speech to provide feedback to the user.
[0028] In addition, in the voice processing device according to the present embodiment, among the slots required in the user utterance text, the first slot is related to a domain, among the slots required in the user utterance text, the second slot is related to an intention, and among the slots required in the user utterance text, the third slot is related to a quantity entity name, and one or more processors may be configured to determine the slots with missing entity names in the order of the first slot, the second slot, and the third slot among the slots required in the user utterance text when determining the slots with missing entity names.
[0029] Additionally, one or more processors may be configured to generate a query text related to a domain when the first slot is missing an entity name when generating the query text, generate a query text related to an intent when the second slot is missing an entity name after the first slot is filled with an entity name, and generate a query text related to a slot where the quantity entity name is missing when the third slot is missing an entity name after the second slot is filled with an entity name.
[0030] Additionally, one or more processors may be configured to fill in the required slots where the entity name is missing with the entity name through feedback of the query utterance and the repeated reception of the user's response utterance when the complete utterance is completed, and to complete the user's utterance as a complete utterance when all required slots where the entity name is missing are filled with the entity name.
[0031] In addition, one or more processors may be further configured to determine and operate one of a plurality of electronic devices in response to a command included in the complete speech voice when the user's speech voice is completed as a complete speech voice.
[0032] In addition to this, other methods for implementing the present invention, other systems, and computer-readable recording media storing a computer program for executing said methods may be further provided.
[0033] Other aspects, features, and advantages other than those described above will become clear from the following drawings, claims, and detailed description of the invention. Effects of the invention
[0034] According to the present invention, when an incomplete speech voice is received from a user, speech recognition performance can be improved by receiving a response speech voice corresponding to the query speech voice fed back to the user and completing it into a complete speech voice.
[0035] In addition, when an incomplete speech voice is received from a user, a response speech voice corresponding to the query speech voice fed back to the user is received to complete it into a complete speech voice, and by determining and operating one of a plurality of electronic devices, rapid and accurate speech recognition processing can be performed.
[0036] In addition, although the voice processing device itself is a mass-produced, uniform product, users perceive the voice processing device as a personalized device, so it can produce the effect of a customized product for the user.
[0037] In addition, user satisfaction can be enhanced by providing various services through voice recognition processing, and rapid and accurate voice recognition processing can be performed.
[0038] In addition, if multiple electronic devices can process voice commands included in the user's spoken voice using only optimal processor resources, the power efficiency of the voice processing device can be improved by determining and operating one of the electronic devices.
[0039] The effects of the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description below. Brief explanation of the drawing
[0040] FIG. 1 is an exemplary diagram of a voice processing environment including an electronic device including a voice processing device according to one embodiment of the present invention, a server, and a network connecting the same. FIG. 2 is a schematic block diagram of a voice processing device according to one embodiment of the present invention. FIG. 3 is a schematic block diagram of an information processing unit according to one embodiment of the voice processing device of FIG. 2. FIG. 4 is an illustrative diagram explaining an embodiment of completing a user's incomplete speech voice into a complete speech voice in FIG. 3. Figure 5 is an example diagram illustrating the query text to be fed back through state table matching for the user's incomplete speech voice in Figure 3. FIG. 6 is a flowchart illustrating a voice processing method according to one embodiment of the present invention. Specific details for implementing the invention
[0041] The advantages and features of the present invention, and the methods for achieving them, will become clear by referring to the embodiments described in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments presented below, but can be implemented in various different forms and should be understood to include all modifications, equivalents, and substitutions that fall within the spirit and scope of the present invention. The embodiments presented below are provided to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention. In describing the present invention, detailed descriptions of related known technologies are omitted if it is determined that such detailed descriptions may obscure the essence of the present invention.
[0042] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this application, terms such as “comprising” or “having” are intended to indicate the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof. Terms such as “first,” “second,” etc., may be used to describe various components, but the components should not be limited by these terms. These terms are used solely for the purpose of distinguishing one component from another.
[0043] Hereinafter, embodiments according to the present invention will be described in detail with reference to the attached drawings. In describing with reference to the attached drawings, identical or corresponding components are given the same reference numerals, and redundant descriptions thereof will be omitted.
[0045] FIG. 1 is an exemplary diagram of a voice processing environment including an electronic device including a voice processing device according to an embodiment of the present invention, a server, and a network connecting the same. Referring to FIG. 1, the voice processing environment (1) may include an electronic device (200) including a voice processing device (100), a server (300), and a network (400). The electronic device (200) including the voice processing device (100) and the server (300) may be connected to each other in a 5G communication environment.
[0046] A voice processing device (100) can receive speech information from a user and provide a voice recognition service through recognition and analysis. Here, the voice recognition service may include receiving speech information from a user, distinguishing between a trigger word and a spoken voice, and outputting a voice recognition processing result for the spoken voice so that the user can recognize it.
[0047] In this embodiment, the speech information may include a starter word and a spoken voice. The starter word is a specific command that activates the voice recognition function of the voice processing device (100) and may be named a wake-up word. The voice recognition function can be activated only if the starter word is included in the speech information, and if the starter word is not included in the speech information, the voice recognition function remains in a deactivated state (e.g., sleep mode). Such a starter word may be pre-configured and stored in the memory (160 in FIG. 2) described later.
[0048] Additionally, the spoken voice is processed after the voice recognition function of the voice processing device (100) is activated by the trigger word, and may include a voice command that the voice processing device (100) can substantially process to generate an output. For example, if the user's speech information is "Hi LG, turn on the air conditioner," the trigger word may be "Hi LG," and the spoken voice may be "turn on the air conditioner." The voice processing device (100) can determine the presence of the trigger word from the user's speech information and analyze the spoken voice to control the air conditioner (205) as an electronic device (200).
[0049] In the present embodiment, the voice processing device (100) can determine one of the electronic devices (200) and operate the determined electronic device (200) when the voice command included in the user's spoken voice can be processed by a plurality of electronic devices (200) while the voice recognition function is activated after receiving a start word.
[0050] To this end, the voice processing device (100) can convert a user's spoken voice containing voice commands into user spoken text. The voice processing device (100) can perform grammatical analysis or semantic analysis on the user spoken text to search for a domain to which the user spoken text belongs and an intention of the user spoken text, and search for one or more named entities as a result of named entity recognition included in the user spoken text. The voice processing device (100) can determine whether the user's spoken voice is a complete spoken voice or an incomplete spoken voice in response to the search results for the domain, the intention, and the named entities. If the user's spoken voice is an incomplete spoken voice, the voice processing device (100) can generate a query spoken voice requesting information to complete the incomplete spoken voice into a complete spoken voice and provide feedback to the user. The voice processing device (100) can complete the complete spoken voice by receiving the user's response spoken voice corresponding to the query spoken voice. When the voice processing device (100) completes the user's spoken voice into a complete spoken voice, it can determine and operate one of the multiple electronic devices in response to the command included in the complete spoken voice.
[0051] In this embodiment, the voice processing device (100) may be included in the electronic device (200). The electronic device (200) may include various devices corresponding to the Internet of Things (IoT), such as a user terminal (201), an artificial intelligence speaker (202) that acts as a hub connecting other electronic devices to the network (400), a washing machine (203), a robot vacuum cleaner (204), an air conditioner (205), and a refrigerator (206). However, examples of the electronic device (200) are not limited to those depicted in FIG. 1.
[0052] Among these electronic devices (200), a user terminal (201) can receive a service for operating or controlling a voice processing device (100) through an authentication process after accessing a voice processing device operating application or a voice processing device operating site. In this embodiment, the user terminal (201) that has completed the authentication process can operate the voice processing device (100) and control the operation of the voice processing device (100).
[0053] In this embodiment, the user terminal (201) may be a desktop computer, smartphone, laptop, tablet PC, smart TV, mobile phone, PDA (personal digital assistant), laptop, media player, micro server, GPS (global positioning system) device, e-book terminal, digital broadcasting terminal, navigation, kiosk, MP3 player, digital camera, home appliance, and other mobile or non-mobile computing devices operated by the user, but is not limited thereto. Additionally, the user terminal (201) may be a wearable terminal such as a watch, glasses, hair band, and ring equipped with communication functions and data processing functions. The user terminal (201) is not limited to the above-described contents, and any terminal capable of web browsing may be used without restriction.
[0054] The server (300) may be a database server that provides big data necessary for applying various artificial intelligence algorithms and data for operating the voice processing device (100). In addition, the server (300) may include a web server or an application server that enables remote control of the operation of the voice processing device (100) using a voice processing device operating application or a voice processing device operating web browser installed on a user terminal (201).
[0055] Here, artificial intelligence (AI) is a field of computer engineering and information technology that studies methods to enable computers to perform thinking, learning, and self-development capable of human intelligence, and it can mean enabling computers to mimic intelligent human behavior.
[0056] Furthermore, artificial intelligence does not exist in isolation but is closely related, directly and indirectly, to many other fields of computer science. Particularly in the modern era, there are very active attempts to introduce AI elements into various sectors of information technology and utilize them to solve problems within those fields.
[0057] Machine learning is a field of artificial intelligence that encompasses research areas that empower computers to learn without explicit programming. Specifically, machine learning can be defined as a technology that studies and builds systems and algorithms capable of learning, making predictions, and improving their own performance based on empirical data. Rather than executing strictly defined static program instructions, machine learning algorithms may adopt an approach of constructing specific models to derive predictions or decisions based on input data.
[0058] The server (300) can receive the user's spoken voice from the voice processing device (100) and convert it into spoken text. The server (300) can perform grammatical analysis or semantic analysis on the user's spoken text to search for the domain to which the user's spoken text belongs and the intent of the user's spoken text, and search for one or more named entities as a result of named entity recognition included in the user's spoken text. Here, the server (300) can execute a machine learning algorithm to perform grammatical analysis or semantic analysis on the user's spoken text. The server (300) can determine whether the user's spoken voice is a complete spoken voice or an incomplete spoken voice in response to the search results for the domain, intent, and named entities. If the user's spoken voice is an incomplete spoken voice, the server (300) can generate a query spoken voice that requests information to complete the incomplete spoken voice into a complete spoken voice. The server (300) can complete a complete speech voice by receiving a user's response speech voice corresponding to a query speech voice from the voice processing device (100). In this embodiment, the server (300) can transmit processing information as described above to the voice processing device (100).
[0059] Depending on the processing capability of the voice processing device (100), at least some of the following may be performed by the voice processing device (100): conversion into the user speech text described above, search for domains, intentions, and entity names, determination of whether the user's speech voice is a complete speech voice or an incomplete speech voice, generating a query speech voice and providing feedback to the user if the user's speech voice is an incomplete speech voice, and receiving the user's response speech voice to complete the complete speech voice.
[0060] The network (400) can perform the role of connecting an electronic device (200) including a voice processing device (100) and a server (300). Such a network (400) may include wired networks such as LANs (local area networks), WANs (wide area networks), MANs (metropolitan area networks), and ISDNs (integrated service digital networks), or wireless networks such as wireless LANs, CDMA, Bluetooth, and satellite communication, but the scope of the present invention is not limited thereto. In addition, the network (400) can transmit and receive information using short-range communication and / or long-range communication. Here, short-range communication may include Bluetooth, RFID (radio frequency identification), infrared communication (IrDA, infrared data association), UWB (ultra-wideband), ZigBee, and Wi-Fi (wireless fidelity) technologies, and long-range communication may include CDMA (code division multiple access), FDMA (frequency division multiple access), TDMA (time division multiple access), OFDMA (orthogonal frequency division multiple access), and SC-FDMA (single carrier frequency division multiple access) technologies.
[0061] The network (400) may include connections of network elements such as hubs, bridges, routers, switches, and gateways. The network (400) may include one or more connected networks, such as a multi-network environment, including a public network such as the Internet and a private network such as a secure corporate private network. Access to the network (400) may be provided through one or more wired or wireless access networks. Furthermore, the network (400) may support an IoT (Internet of Things) network and / or 5G communication that exchanges and processes information between distributed components such as objects.
[0063] FIG. 2 is a schematic block diagram of a voice processing device according to an embodiment of the present invention. In the following description, parts that overlap with the description of FIG. 1 will be omitted. Referring to FIG. 2, the voice processing device (100) may include a communication unit (110), a user interface unit (120) including a display unit (121) and an operation unit (122), a sensing unit (130), an audio processing unit (140) including an audio input unit (141) and an audio output unit (142), an information processing unit (150), a memory (160), and a control unit (170).
[0064] The communication unit (110) may provide a communication interface necessary to provide transmission and reception signals between a voice processing device (100) and / or an electronic device (200) and / or a server (300) in the form of packet data in conjunction with a network (400). Furthermore, the communication unit (110) may perform the role of receiving a predetermined information request signal from an electronic device (200) and may perform the role of transmitting information processed by the voice processing device (100) to the electronic device (200). Additionally, the communication unit (110) may transmit a predetermined information request signal from the electronic device (200) to the server (300), receive a response signal processed by the server (300), and transmit it to the electronic device (200). Furthermore, the communication unit (110) may be a device including hardware and software necessary to transmit and receive signals, such as control signals or data signals, through a wired or wireless connection with another network device.
[0065] In addition, the communication unit (110) can support various types of intelligent communication (IoT (internet of things), IoE (internet of everything), IoST (internet of small things), etc.) and can support M2M (machine to machine) communication, V2X (vehicle to everything communication) communication, D2D (device to device) communication, etc.
[0066] The display unit (121) of the user interface unit (120) can display the operating status of the voice processing device (100) under the control of the control unit (170). According to an embodiment, the display unit (121) may be configured as a touchscreen by forming a layered structure with a touchpad. In this case, the display unit (121) may also be used as an operation unit (122) capable of inputting information by the user's touch. To this end, the display unit (121) may be configured as a touch recognition display controller or various other input / output controllers. For example, the touch recognition display controller may provide an output interface and an input interface between the device and the user. The touch recognition display controller may transmit and receive electrical signals to and from the control unit (170). In addition, the touch recognition display controller displays visual output to the user, and the visual output may include text, graphics, images, videos, and combinations thereof. Such a display part (121) may be a specific display component, such as an OLED (organic light emitting display), an LCD (liquid crystal display), or an LED (light emitting display) capable of touch recognition.
[0067] The operation unit (122) of the user interface unit (120) is equipped with a plurality of operation buttons (not shown) and can transmit a signal corresponding to an input button to the control unit (170). This operation unit (122) may be composed of a sensor, button, or switch structure capable of recognizing a user's touch or press operation. In this embodiment, the operation unit (122) can transmit an operation signal operated by the user to the control unit (170) to check or change various information related to the operation of the voice processing device (100) displayed on the display unit (121).
[0068] The sensing unit (130) may include various sensors that sense the surrounding conditions of the voice processing device (100), and may include a proximity sensor (not shown) and an image sensor (not shown). The proximity sensor may acquire location data of an object (e.g., a user) located around the voice processing device (100) by utilizing infrared light, etc. Meanwhile, the location data of the user acquired by the proximity sensor may be stored in the memory (160).
[0069] The image sensor may include a camera (not shown) capable of capturing images around the voice processing device (100), and multiple cameras may be installed for shooting efficiency. For example, the camera may include an image sensor (e.g., CMOS image sensor) configured to include at least one optical lens and multiple photodiodes (e.g., pixels) that form an image by light passing through the optical lens, and a digital signal processor (DSP) that forms an image based on signals output from the photodiodes. The digital signal processor can generate not only still images but also video composed of frames composed of still images. Meanwhile, the image captured and acquired by the camera acting as the image sensor may be stored in memory (160).
[0070] In this embodiment, the sensing unit (130) is limited to a proximity sensor and an image sensor, but is not limited thereto and may include at least one of a sensor capable of detecting the surrounding conditions of the voice processing device (100), such as a Lidar sensor, a weight sensor, an illumination sensor, a touch sensor, an acceleration sensor, a magnetic sensor, a gravity sensor (G-sensor), a gyroscope sensor, a motion sensor, an RGB sensor, an infrared sensor (IR sensor: infrared sensor), a fingerprint sensor (finger scan sensor), an ultrasonic sensor, an optical sensor, a microphone, a battery gauge, an environmental sensor (e.g., a barometer, a hygrometer, a thermometer, a radiation sensor, a heat sensor, a gas sensor, etc.), and a chemical sensor (e.g., an electronic nose, a healthcare sensor, a biometric sensor, etc.). Meanwhile, in the present embodiment, the voice processing device (100) can combine and utilize information sensed from at least two of these sensors.
[0071] Among the audio processing unit (140), the audio input unit (141) can receive user speech information (e.g., a start word and spoken voice) and transmit it to the control unit (170), and the control unit (170) can transmit the user speech information to the information processing unit (150). To this end, the audio input unit (141) may be equipped with one or more microphones (not shown). Additionally, to receive the user's spoken voice more accurately, multiple microphones (not shown) may be provided. Here, each of the multiple microphones may be spaced apart at different locations and the received user's spoken voice may be processed into an electrical signal.
[0072] In an optional embodiment, the audio input unit (141) may use various noise removal algorithms to remove noise generated during the process of receiving the user's speech information. In an optional embodiment, the audio input unit (141) may include various components for voice signal processing, such as a filter (not shown) that removes noise when receiving the user's speech information, and an amplifier (not shown) that amplifies and outputs a signal output from the filter.
[0073] The audio output unit (142) of the audio processing unit (140) can output audio such as warning sounds, operation modes, operation status, error status, etc., notification messages, response information corresponding to the user's speech information, and processing results corresponding to the user's spoken voice (voice command) under the control of the control unit (170). The audio output unit (142) can convert an electrical signal from the control unit (170) into an audio signal and output it. To this end, a speaker, etc., may be provided.
[0074] The information processing unit (150) can convert the user's spoken voice containing a voice command into user spoken text while the voice recognition function is activated after receiving a trigger word. The information processing unit (150) can perform grammatical analysis or semantic analysis on the user spoken text to search for the domain to which the user spoken text belongs and the intention of the user spoken text, and search for one or more entities as a result of named entity recognition included in the user spoken text. The information processing unit (150) can determine whether the user's spoken voice is a complete spoken voice or an incomplete spoken voice in response to the search results for the domain, the intention, and the named entities. If the user's spoken voice is an incomplete spoken voice, the information processing unit (150) can generate a query spoken voice requesting information to complete the incomplete spoken voice into a complete spoken voice and provide feedback to the user. The information processing unit (150) can complete the complete spoken voice by receiving the user's response spoken voice corresponding to the query spoken voice. When the user's spoken voice is completed into a complete spoken voice, the information processing unit (150) transmits the complete spoken voice to the control unit (170), and the control unit (170) can determine and operate one of the multiple electronic devices in response to the command included in the complete spoken voice.
[0075] In this embodiment, the information processing unit (150) may perform learning in conjunction with the control unit (170) or receive learning results from the control unit (170). In this embodiment, the information processing unit (150) may be provided outside the control unit (170) as shown in FIG. 2, may be provided inside the control unit (170) and operate like the control unit (170), or may be provided inside the server (300) of FIG. 1. The details of the information processing unit (150) will be explained below with reference to FIGS. 3 to 5.
[0076] The memory (160) stores various information necessary for the operation of the voice processing device (100) and can store control software capable of operating the voice processing device (100), and may include a volatile or non-volatile recording medium. For example, the memory (160) may store a preset trigger word for determining the presence of a trigger word from the user's spoken voice. Meanwhile, the trigger word may be set by the manufacturer. For example, "Hi LG" may be set as the trigger word, and the setting may be changed by the user. Such a trigger word is input to activate the voice processing device (100), and the voice processing device (100), upon recognizing the trigger word spoken by the user, can switch to a voice recognition activation state.
[0077] Additionally, the memory (160) can store user speech information (starter words and speech voice) received through the audio input unit (141), can store information detected by the sensing unit (130), and can store information processed by the information processing unit (150).
[0078] Additionally, the memory (160) may store commands to be executed by the information processing unit (150), such as a command to convert a user’s spoken voice containing voice commands into user spoken text; a command to perform grammatical or semantic analysis on the user spoken text to search for a domain to which the user spoken text belongs and an intention of the user spoken text, and to search for one or more entities as a result of named entity recognition included in the user spoken text; a command to determine whether the user’s spoken voice is a complete spoken voice or an incomplete spoken voice in response to the search results for the domain, the intention, and the entities; a command to generate a query spoken voice requesting information to complete the incomplete spoken voice into a complete spoken voice and provide feedback to the user if the user’s spoken voice is an incomplete spoken voice; and a command to receive the user’s response spoken voice corresponding to the query spoken voice and complete the complete spoken voice. Additionally, the memory (160) may store various information processed by the information processing unit (150).
[0079] Here, the memory (160) may include magnetic storage media or flash storage media, but the scope of the present invention is not limited thereto. Such memory (160) may include internal memory and / or external memory, and may include volatile memory such as DRAM, SRAM, or SDRAM, non-volatile memory such as OTPROM (one time programmable ROM), PROM, EPROM, EEPROM, mask ROM, flash ROM, NAND flash memory, or NOR flash memory, flash drives such as SSD, CF (compact flash) card, SD card, Micro-SD card, Mini-SD card, Xd card, or memory stick, or storage devices such as HDD.
[0080] Here, simple voice recognition is performed by the voice processing device (100), and high-level voice recognition, such as natural language processing, can be performed by the server (300). For example, if the word spoken by the user is a preset trigger word, the voice processing device (100) can switch to a state for receiving the spoken voice as a voice command. In this case, the voice processing device (100) performs only the voice recognition process up to the point of inputting the trigger word voice, and voice recognition for subsequent utterances can be performed through the server (300). Since there are limitations to the system resources of the voice processing device (100), complex natural language recognition and processing can be performed through the server (300).
[0081] The control unit (170) transmits speech information received through the audio input unit (141) to the information processing unit (150), and can provide the voice recognition processing result from the information processing unit (150) as visual information through the display unit (121) or as auditory information through the audio output unit (142).
[0082] The control unit (170) is a type of central processing unit and can control the operation of the entire voice processing unit (100) by running control software loaded in the memory (160). The control unit (170) may include all types of devices capable of processing data, such as a processor. Here, 'processor' may refer to a data processing device embedded in hardware that has a physically structured circuit to perform functions expressed by code or instructions included in a program, for example. Examples of such data processing devices embedded in hardware may include microprocessors, central processing units (CPUs), processor cores, multiprocessors, application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc., but the scope of the present invention is not limited thereto.
[0083] In this embodiment, the control unit (170) can perform machine learning, such as deep learning, on user speech so that the voice processing device (100) outputs an optimal voice recognition processing result, and the memory (160) can store data used for machine learning, result data, etc.
[0084] Deep learning, a type of machine learning, can learn by descending to deep levels in multiple stages based on data. Deep learning can represent a set of machine learning algorithms that extract key data from multiple data as the stages are increased.
[0085] The deep learning structure may include an artificial neural network (ANN), and for example, the deep learning structure may be composed of a deep neural network (DNN), such as a convolutional neural network (CNN), a recurrent neural network (RNN), or a deep belief network (DBN). The deep learning structure according to the present embodiment may utilize various known structures. For example, the deep learning structure according to the present invention may include a CNN, an RNN, a DBN, etc. RNNs are widely used in natural language processing and are effective structures for processing time-series data that changes over time; they can be constructed by stacking layers at every moment. DBNs may include a deep learning structure composed of multiple layers of a restricted Boltzmann machine (RBM), which is a deep learning technique. By repeating RBM training, a DBN having that number of layers can be constructed once a certain number of layers is reached. CNNs can include models that mimic human brain functions, based on the assumption that when a person recognizes an object, basic features of the object are extracted, then undergo complex calculations within the brain to recognize the object based on the results.
[0086] Meanwhile, the training of artificial neural networks can be achieved by adjusting the weights of connections between nodes (and bias values if necessary) to produce a desired output for a given input. Furthermore, artificial neural networks can continuously update weight values through learning. Additionally, methods such as backpropagation can be used for the training of artificial neural networks.
[0087] Meanwhile, the control unit (170) may be equipped with an artificial neural network and can perform machine learning-based user recognition and user voice tone recognition using the received voice input signal as input data.
[0088] The control unit (170) may include an artificial neural network, such as a deep neural network (DNN), CNN, RNN, DBN, etc., and may learn the deep neural network. Both unsupervised learning and supervised learning may be used as machine learning methods for such artificial neural networks. The control unit (170) may control the structure of the tone recognition artificial neural network to be updated after learning according to the settings.
[0090] FIG. 3 is a schematic block diagram of an information processing unit according to one embodiment of the speech recognition device of FIG. 2. In the following description, parts that overlap with the description of FIG. 1 and FIG. 2 will be omitted. Referring to FIG. 3, the information processing unit (150) may include an automatic speech recognition processing unit (151), a natural language understanding processing unit (152), a dialogue manager processing unit (153), a natural language generation processing unit (154), a text-to-speech conversion processing unit (155), and a first database (156) to a third database (158). As an optional embodiment, the information processing unit (150) may include one or more processors. As an optional embodiment, the automatic speech recognition processing unit (151) to the third database (158) may correspond to one or more processors. As an optional embodiment, the automatic speech recognition processing unit (151) to the third database (158) may correspond to software components configured to be executed by one or more processors.
[0091] The automatic voice recognition processing unit (151) can generate user speech text by converting the user’s spoken voice, which includes a voice command, into text. Here, the user’s spoken voice may include a voice command that instructs the operation of one of the multiple electronic devices (200) from the user’s perspective, but may also include a voice command that can operate multiple electronic devices (200) from the perspective of the voice processing device (100). For example, if the user’s spoken voice is "Turn it on," from the user’s perspective, it may be a voice command that instructs the operation of the air conditioner (205) among the multiple electronic devices (200), but from the perspective of the voice processing device (100), it may be a voice command that can operate multiple electronic devices (200), including the user terminal (201), washing machine (203), robot vacuum cleaner (204), and air conditioner (205).
[0092] In this embodiment, the automatic speech recognition processing unit (151) can perform speech-to-text conversion (STT). The automatic speech recognition processing unit (151) can convert user speech voice input through the audio input unit (141) into user speech text. In this embodiment, the automatic speech recognition processing unit (151) may include a speech recognition unit (not shown). The speech recognition unit may include an acoustic model and a language model. For example, the acoustic model may include information related to vocalization, and the language model may include information regarding unit phoneme information and combinations of unit phoneme information. The speech recognition unit can convert user speech voice into user speech text using information related to vocalization and information regarding unit phoneme information. Information regarding the acoustic model and the language model may be stored in the first database (156), that is, the automatic speech recognition database.
[0093] The natural language understanding processing unit (152) can perform syntactic analysis or semantic analysis on the user's utterance text to explore the domain and intent of the user's utterance speech. Here, syntactic analysis can divide the user's utterance text into grammatical units (e.g., words, phrases, morphemes, etc.) and identify what grammatical elements the divided units have. Additionally, semantic analysis can be performed using semantic matching, rule matching, formula matching, etc. In this embodiment, the domain may include information specifying the type (product) of an electronic device (200) that the user intends to operate. Additionally, in this embodiment, the intent may include information indicating how to operate the electronic device (200) included in the domain. For example, if the user utterance text is <lower the temperature of the air conditioner by 5 degrees>, the natural language understanding processing unit (152) can search for air conditioner as the domain and lower the temperature as the intention.
[0094] Additionally, the natural language understanding processing unit (152) can search for one or more named entities as a result of named entity recognition for user utterance text. Rule-based named entities can be recognized using a named entity dictionary and a combined word dictionary stored in the second database (157) for named entity recognition, and one or more named entities can be obtained as a result of recognition. Here, a named entity may refer to a noun or number that has a unique meaning appearing in text. Named entities can be broadly classified into name expressions such as personal names, place names, and organization names; time expressions such as dates and times; and numeric expressions such as amounts or percentages. For example, if the text is "<Lower the air conditioner temperature by 5 degrees>", the named entities may include air conditioner, temperature, and 5 degrees.
[0095] Additionally, the natural language understanding processing unit (152) can determine essential slots from the user utterance text. Here, the term "essential slot" refers to the information that the voice processing device (100) needs to know to provide a response regarding the user utterance text. In this embodiment, the essential slots may include first to third slots. The first slot may include a slot related to a domain, the second slot may include a slot related to an intention, and the third slot may include a slot related to a quantity entity name. For example, if the user utterance text is "<Lower the temperature of the air conditioner by 5 degrees>", the first slot may include "air conditioner", the second slot may include "lowering the temperature", and the third slot may include "5 degrees". Also, for example, if the user utterance text is "<lower the temperature by 5 degrees>", it can be seen that the first slot does not exist, the second slot includes "lowering the temperature", and the third slot does not exist. Here, the determination of the required slot may also be performed in the dialogue manager processing unit (153) in addition to the natural language understanding processing unit (152).
[0096] In one embodiment, the natural language understanding processing unit (152) may use matching rules stored in the second database (157), that is, the natural language understanding database, to search for domains and intentions. The natural language understanding processing unit (152) may identify the meaning of words (named entities) extracted from user utterance text using linguistic features (e.g., grammatical elements) such as morphemes and phrases, and may search for user intentions by matching the identified meaning of the words to domains and intentions. For example, the natural language understanding processing unit (152) may search for domains and intentions by calculating how many words (named entities) extracted from user utterance text are included in each domain and intention.
[0097] In one embodiment, the natural language understanding processing unit (152) may utilize a statistical model to explore domains, intentions, and named entities. The statistical model may refer to various types of machine learning models. In this embodiment, the natural language understanding processing unit (152) may refer to a domain classifier model to explore domains, an intent classifier model to explore intentions, and a named entity recognizer model to explore named entities.
[0098] The conversation manager processing unit (153) generally controls the conversation between the user and the voice processing device (100) and can determine the query text to be generated using the user speech text understanding result received from the automatic speech recognition processing unit (151), or can have the natural language generation processing unit (154) generate language text in the user's language when generating the query text to be fed back to the user.
[0099] In this embodiment, the conversation manager processing unit (153) receives the domain, intent, and entity name searched by the natural language understanding processing unit (152) and can determine whether the user's speech is a complete speech or an incomplete speech.
[0100] The conversation manager processing unit (153) can determine a required slot including a first slot related to a domain, a second slot related to an intention, and a third slot related to a quantity entity name from the user utterance text.
[0101] The conversation manager processing unit (153) can determine that the user's utterance is a complete utterance when all required slots in the user's utterance text are filled with entity names. For example, if the user's utterance text is "<Please lower the temperature of the air conditioner by 5 degrees>", the first slot related to the domain (air conditioner), the second slot related to the intention (lowering the temperature), and the third slot related to the quantity entity name (5 degrees) are all filled with entity names, so it can be said to be a complete utterance.
[0102] The conversation manager processing unit (153) can determine that the user's spoken voice is an incomplete spoken voice if there is one or more missing required slots in the user's spoken text. For example, if the user's spoken text is "<lower the temperature further>", the first slot (air conditioner) and the third slot (5 degrees) are missing when compared to the complete spoken voice described above. The user's spoken text "<lower the temperature further>" described above can be considered an incomplete spoken voice with two required slots missing.
[0103] When the conversation manager processing unit (153) determines that the user's speech voice is an incomplete speech voice, it can generate a query speech voice requesting information to complete the incomplete speech voice into a complete speech voice and provide feedback to the user.
[0104] Here, the conversation manager processing unit (153) may request the natural language generation processing unit (154) to generate query text in order to generate query speech voice and provide feedback to the user. The natural language generation processing unit (154) may generate query text in response to the request to generate query text. The text-to-speech conversion processing unit (155) may convert the query text into query speech voice and then provide feedback to the user through the audio output unit (142).
[0105] In this embodiment, the conversation manager processing unit (153) can perform a request to the natural language generation processing unit (154) to generate query text by using a state table already established in the third database (158). The state table already established in the third database (158) may include a slot required in the user command text according to the domain and intent, and an action state indicating the query text to be requested corresponding to the required slot.
[0106] The conversation manager processing unit (153) can generate query text corresponding to slots where entity names are missing in user utterance text based on a pre-established state table. Here, if there are many missing slots, the slots where entity names are missing can be determined in the order of the first slot, second slot, and third slot.
[0107] First, the conversation manager processing unit (153) can generate query text related to the domain based on the status table if the entity name is missing in the first slot. Next, the conversation manager processing unit (153) can generate query text related to the intent based on the status table if the entity name is missing in the second slot after the entity name is filled in the first slot. Finally, the conversation manager processing unit (153) can generate query text related to the third slot based on the status table if the entity name is missing in the third slot after the entity name is filled in the second slot.
[0108] The conversation manager processing unit (153) can complete a complete speech voice by receiving a user's response speech voice corresponding to a query speech voice. The conversation manager processing unit (153) can fill in the required slots where entity names are missing with entity names through feedback of the query speech voice and repeated reception of the user's response speech voice. Subsequently, the conversation manager processing unit (153) can complete the user's speech voice into a complete speech voice when the required slots where entity names are missing are all filled with entity names. When the user's speech voice is completed into a complete speech voice, the control unit (170) can determine and operate one of the multiple electronic devices (200) in response to the command included in the complete speech voice.
[0109] The natural language generation processing unit (154) can generate query text using a knowledge base at the request of the dialogue manager processing unit (153). Additionally, the natural language generation processing unit (154) can generate text indicating the completion of the operation of the electronic device (200) when the operation of the determined electronic device (200) is completed.
[0110] The text-to-speech conversion processing unit (155) can convert the query text generated by the natural language generation processing unit (154) into query speech voice and provide feedback to the user through the audio output unit (142). Additionally, the text-to-speech conversion processing unit (155) can convert the text of the operation completion of the electronic device (200) generated by the natural language generation processing unit (154) into the speech of the operation completion of the electronic device (200) and provide feedback to the user through the audio output unit (142).
[0112] FIG. 4 is an illustrative diagram explaining an embodiment in which the user's incomplete speech voice in FIG. 3 is completed into a complete speech voice. In the following description, parts that overlap with the description of FIG. 1 to FIG. 3 will be omitted.
[0113] Referring to FIG. 4a, the utterance information (401) "Hi LG, lower the temperature" uttered by the user can be transmitted to the information processing unit (150) through the audio input unit (141). Once the conversion of the user utterance text for the voice command "lower the temperature" is completed, the user utterance text can be input to the natural language understanding processing unit (152). The natural language understanding processing unit (152) can perform grammatical analysis or semantic analysis on the user utterance text to search for domains, intentions, and named entities (402).
[0114] The conversation manager processing unit (153) determines whether the user's spoken voice is a complete spoken voice or an incomplete spoken voice, and if it is an incomplete spoken voice, it may request the generation of a query text. Because the first slot and the third slot are missing from the aforementioned user spoken text, the conversation manager processing unit (153) may determine the aforementioned user's spoken voice as an incomplete spoken voice. The conversation manager processing unit (153) may request the generation of a first query text (403) to fill the first slot related to the domain in order of priority by referring to a pre-established state table. The natural language generation processing unit (154) may generate <Which device should I lower it on?> as the first query text. The text-to-speech conversion processing unit (155) converts the first query text into a first query spoken voice (404) and may provide feedback to the user through the audio output unit (142).
[0115] Referring to FIG. 4b, a user who receives a first query speech voice may utter "<air conditioner>" as a response speech voice (405), which may be converted into user speech text and input to a natural language understanding processing unit (152). The natural language understanding processing unit (152) may add (406) an entity name obtained from the current user speech text (air conditioner) to the domain, intent, and entity name obtained from the previous user speech text (lower the temperature).
[0116] The conversation manager processing unit (153) can determine that the aforementioned user's spoken voice is an incomplete spoken voice because the domain was obtained as air conditioner and the intention was obtained as lowering the temperature from the previous and current user speech text, but the third slot was omitted.
[0117] The conversation manager processing unit (153) may request the natural language generation processing unit (154) to generate a second query text (407) to fill a third slot representing an entity name representing a quantity by referring to a pre-established state table. The natural language generation processing unit (154) may generate <How many degrees should I lower the air conditioner temperature?> as the second query text. The text-to-speech conversion processing unit (155) may convert the second query text into a second query speech voice (408) and provide feedback to the user through the audio output unit (142).
[0118] Referring to FIG. 4c, a user who receives a second query speech voice may utter <5 degrees> as a response speech voice (409), which can be converted into user speech text and input to a natural language understanding processing unit (152). The natural language understanding processing unit (152) may add (410) the entity name obtained from the current user speech text (5 degrees) to the domain, intent, and entity name obtained from the previous user speech text (lower the air conditioner temperature).
[0119] The conversation manager processing unit (153) can determine that the conversation manager processing unit (153) is a complete speech voice because the domain is obtained as air conditioner from the previous and current user speech text, the intention is obtained as lowering the temperature, and 5 degrees is obtained as an entity name indicating the quantity.
[0120] The conversation manager processing unit (153) transmits the completion of a complete spoken voice to the control unit (170), and the control unit (170) determines the air conditioner (205) among the plurality of electronic devices (200) and can request (411) that the air conditioner (205) lower the temperature by 5 degrees. The control unit (170) can receive a response signal (412) that the air conditioner (205) has lowered the temperature by 5 degrees and transmit it to the conversation manager processing unit (153). In response to this response signal, the conversation manager processing unit (153) can request (413) the natural language generation processing unit (154) to generate a command execution completion text. The natural language generation processing unit (154) can generate "<The temperature of the air conditioner has been lowered by 5 degrees>" as the command execution completion text. The text-to-speech conversion processing unit (155) converts the command execution completion text into the command execution completion speech voice (414) and can provide feedback to the user through the audio output unit (142).
[0122] FIG. 5 is an example diagram illustrating the query text to be fed back through state table matching for the user's incomplete speech voice in FIG. 3. In the following description, parts that overlap with the descriptions of FIG. 1 to 4 will be omitted.
[0123] Referring to FIG. 5, 510 represents the type of slot for user speech voice, 520 represents a state table already established in the third database (158), and 530 represents the feedback query text generated by the natural language generation processing unit (154) in response to a request from the conversation manager processing unit (153). In the type of slot (510) for user speech voice (e.g., lower the air conditioner temperature by 5 degrees), a question mark indicates a required slot where the entity name is missing, a circle indicates a slot that is already filled and meaningless, Prd(product) indicates a state where the entity name is filled in the first slot related to the domain, Temp indicates a state where the entity name is filled in the second slot related to the intent, and Num indicates a state where the entity name is filled in the third slot related to the entity name representing the quantity.
[0124] In FIG. 5, assuming the user's complete speech voice is "Lower the air conditioner temperature by 5 degrees," the required slots may include air conditioner (first slot), lower temperature (second slot), and 5 degrees (third slot). For example, the user speech voice may include a first user speech voice (511) of "Lower it further," a second user speech voice (512) of "Lower the air conditioner further," a third user speech voice (513) of "Lower the temperature further," a fourth user speech voice (514) of "Lower the temperature by 5 degrees further," and a fifth speech voice (515) of "Lower the air conditioner temperature further."
[0125] The voice processing device (100) can determine the existence of an essential slot in which an entity name is missing from each of the first user speech voice (511) to the fifth user speech voice (515), and determine that each of the first user speech voice (511) to the fifth user speech voice (515) is an incomplete speech voice.
[0126] The voice processing device (100) can determine that the first user speech voice (511) is an incomplete speech voice because entity names are missing in the first slot, the second slot, and the third slot. The voice processing device (100) can determine that the second user speech voice (512) is an incomplete speech voice because entity names are missing in the second slot and the third slot. The voice processing device (100) can determine that the third user speech voice (513) is an incomplete speech voice because entity names are missing in the first slot and the third slot. The voice processing device (100) can determine that the fourth user speech voice (511) is an incomplete speech voice because entity names are missing in the first slot. The voice processing device (100) can determine that the fifth user speech voice (515) is an incomplete speech voice because entity names are missing in the third slot.
[0127] In the first column (521) of the status table (520), a domain corresponding to priority (e.g., Prd) is designated as the required slot corresponding to a plurality of domains (unclear domains: air conditioner / washing machine / TV / smartphone) and a plurality of intentions (unclear intentions: temperature down (Temp_down) / volume down (Volume_down)) that can be obtained from the first user speech voice (511), and an action status (Request_domain) representing the query text to be requested corresponding to the domain that is the required slot is designated.
[0128] In the second column (522) of the status table (520), corresponding to a plurality of intents (unclear intent: lower temperature (Temp_down) / lower volume (Volume_down)) obtainable from the second user speech voice (512), the next-ranked intent (e.g., Temp) is designated as the required slot, and an action state (Request_intent) representing the query text to be requested is designated corresponding to the intent (Temp, temperature) that is the required slot.
[0129] In the third column (523) and the fourth column (524) of the status table (520), a domain (e.g., Prd) is designated as the required slot corresponding to a plurality of domains (unclear domains: air conditioner / washing machine / TV / smartphone) that can be obtained from the third user speech voice (513) and the fourth user speech voice (514), and an action status (Request_domain) representing the query text to be requested is designated corresponding to the domain that is the required slot.
[0130] In the fifth column (525) of the status table (520), corresponding to the clear domain and clear intent obtainable from the fifth user speech voice (515), a number (e.g., Num) corresponding to the quantity is designated as the required slot, and an action status (Request_entity) representing the query text to be requested is designated corresponding to the quantity (Num) which is the required slot.
[0131] The voice processing device (100) can generate <Which device should I lower?> as a first query text (531) to be fed back to the user using the action state (Request_domain) of the first column (521). The voice processing device (100) can generate <Which part of the air conditioner should I lower?> as a second query text (532) to be fed back to the user using the action state (Request_intent) of the second column (522).
[0132] The voice processing device (100) can generate <Which device should I lower it?> as a third query text (533) and a fourth query text (534) to be fed back to the user using the action status (Request_domain) of the third column (523) and the fourth column (524).
[0133] The voice processing device (100) can generate <How many degrees should I lower the air conditioner temperature?> as the fifth query text (535) to be fed back to the user using the action state (Request_entity) of the fifth column (525).
[0134] The voice processing device (100) can convert the first query text (531) to the fifth query text (535) into the first query speech voice to the fifth query speech voice and provide feedback to the user. Subsequently, the voice processing device (100) can receive a response speech voice corresponding to the query speech voice from the user to complete the complete speech voice. If the complete speech voice is not completed, the feedback of the query speech voice and the reception of the user's response speech voice can be repeated to complete the complete speech voice.
[0136] FIG. 6 is a flowchart illustrating a voice processing method according to an embodiment of the present invention. In the following description, parts that overlap with the description of FIG. 1 to FIG. 5 will be omitted.
[0137] Referring to FIG. 6, in step S610, the voice processing device (100) converts the user’s spoken voice, which includes a voice command, into user spoken text. Here, the user’s spoken voice may include a voice command that directs the operation of one of the multiple electronic devices (200) from the user’s perspective, but may also include a voice command that can operate the multiple electronic devices (200) from the perspective of the voice processing device (100).
[0138] In step S620, the voice processing device (100) performs grammatical or semantic analysis on the user utterance text to search for the domain to which the user utterance text belongs and the intent of the user utterance text, and searches for one or more named entities as a result of the named entity recognition contained in the user utterance text. The voice processing device (100) may search for a domain specifying the type of electronic device that the user intends to operate, an intent indicating how to operate the electronic device, and named entities including nouns or numbers having unique meanings appearing in the user utterance text.
[0139] In step S630, the voice processing device (100) determines whether the user's spoken voice is a complete spoken voice or an incomplete spoken voice in response to the search results for the domain, the intention, and the named entity. The voice processing device (100) may determine from the user's spoken text a required slot including a first slot related to the domain, a second slot related to the intention, and a third slot related to the quantity named entity. The voice processing device (100) may determine the user's spoken voice as a complete spoken voice if all required slots in the user's spoken text are filled with named entities, and may determine the user's spoken voice as an incomplete spoken voice if there are required slots in the user's spoken text where the named entity is missing.
[0140] In step S640, if the user's spoken voice is an incomplete spoken voice, the voice processing device (100) generates a query spoken voice that requests information to complete the incomplete spoken voice into a complete spoken voice and feeds it back to the user.
[0141] The voice processing device (100) can determine which slots among the necessary slots in the user utterance text are missing named entities based on a pre-established state table containing the necessary slots in the user command text according to the domain and intent, and the query text to be requested corresponding to the necessary slots. The voice processing device (100) can generate query text corresponding to the missing slots based on the state table. The voice processing device (100) can convert the generated query text into query utterance voice and provide feedback to the user.
[0142] When determining slots where entity names are missing, the voice processing device (100) can determine slots where entity names are missing in the order of the first slot, second slot, and third slot among the slots required in the user speech text.
[0143] When generating query text, the voice processing device (100) can generate query text related to a domain if an entity name is missing in the first slot, generate query text related to an intent if an entity name is missing in the second slot after the entity name is filled in the first slot, and generate query text related to the slot where the entity name is missing if an entity name is missing in any of the third slots after the entity name is filled in the second slot.
[0144] In step S650, the voice processing device (100) receives the user's response voice corresponding to the query voice and completes the complete voice. Through repeated feedback of the query voice and reception of the user's response voice, the voice processing device (100) fills the required slots where the entity name is missing with the entity name, and completes the user's voice when the required slots where the entity name is missing are all filled with the entity name, into a complete voice.
[0145] Afterwards, when the user's spoken voice is completed as a complete spoken voice, the voice processing device (100) can determine and operate one of the multiple electronic devices in response to the command included in the complete spoken voice.
[0147] The embodiments according to the present invention described above may be implemented in the form of a computer program that can be executed through various components on a computer, and such a computer program may be recorded on a computer-readable medium. In this case, the medium may include a magnetic medium such as a hard disk, a floppy disk, and a magnetic tape, an optical recording medium such as a CD-ROM and a DVD, a magneto-optical medium such as a floptical disk, and a hardware device specifically configured to store and execute program instructions, such as a ROM, RAM, or flash memory.
[0148] Meanwhile, the above-mentioned computer program may be one specifically designed and configured for the present invention, or one known and available to those skilled in the art of computer software. Examples of computer programs may include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.
[0149] In the specification of the present invention (particularly in the claims), the use of the term "above" and similar descriptive terms may be in both singular and plural. Furthermore, where a range is described in the present invention, it is to include an invention to which individual values belonging to said range are applied (unless otherwise stated), and this is equivalent to describing each individual value constituting said range in the detailed description of the invention.
[0150] Unless explicitly stated or contrary to the order of the steps constituting the method according to the present invention, said steps may be performed in a suitable order. The present invention is not necessarily limited by the order in which said steps are described. The use of all examples or exemplary terms (e.g., etc.) in the present invention is merely for the purpose of describing the present invention in detail, and the scope of the present invention is not limited by said examples or exemplary terms unless limited by the claims. Furthermore, those skilled in the art will understand that various modifications, combinations, and changes may be made according to design conditions and factors within the scope of the claims or equivalents to which they are added.
[0151] Accordingly, the scope of the present invention should not be limited to the embodiments described above, and all scopes equivalent to or equivalently modified from the claims set forth below, as well as the claims set forth below, shall be considered to fall within the scope of the concept of the present invention. Explanation of the symbols
[0152] 100: Voice processing unit 200: Electronic Devices 300: Server 400: Network
Claims
Claim 1 A voice processing method performed by a voice processing device, comprising: a step of converting a user’s spoken voice containing voice commands into user spoken text; a step of performing grammatical analysis or semantic analysis on the user spoken text to search for a domain to which the user spoken text belongs and an intent of the user spoken text, and searching for one or more named entities as a result of named entity recognition included in the user spoken text; a step of determining whether the user’s spoken voice is a complete spoken voice or an incomplete spoken voice in correspondence with the domain, the intent, and the search result for the named entities; and a step of, if the user’s spoken voice is an incomplete spoken voice, generating a query spoken voice requesting information to complete the incomplete spoken voice into the complete spoken voice and providing feedback to the user. The method includes the step of receiving the user's response speech voice corresponding to the query speech voice to complete the complete speech voice, wherein the feedback step comprises: determining a slot in which an entity name is missing among the slots required in the user speech text based on a pre-established state table including a slot required in the user command text according to a domain and intent and a query text to be requested corresponding to the required slot; and generating a query text corresponding to the missing slot based on the state table.A voice processing method comprising the step of converting the generated query text into the query utterance voice and providing feedback to the user, wherein among the slots required in the user utterance text, the first slot is related to a domain, among the slots required in the user utterance text, the second slot is related to an intention, and among the slots required in the user utterance text, the third slot is related to a quantity entity name, and the step of determining the slot where the entity name is missing includes the step of determining the slot where the entity name is missing in the order of the first slot, the second slot, and the third slot among the slots required in the user utterance text. Claim 2 A voice processing method according to claim 1, wherein the converting step comprises converting the user’s spoken voice, which includes a voice command instructing the operation of at least one of a plurality of electronic devices, into the user’s spoken text. Claim 3 A voice processing method according to claim 2, wherein the searching step comprises a domain specifying the type of electronic device that the user intends to operate, an intention indicating how to operate the electronic device, and a step of searching for an entity name including a noun or a number having a unique meaning appearing in the user's utterance text. Claim 4 A voice processing method according to claim 1, wherein the determining step comprises: determining a required slot from the user utterance text including a first slot related to the domain, a second slot related to the intention, and a third slot related to a quantity entity name; determining the user's utterance voice as a complete utterance voice if the required slot in the user utterance text is filled with entity names; and determining the user's utterance voice as an incomplete utterance voice if there exists a required slot in the user utterance text in which entity names are missing. Claim 5 delete Claim 6 delete Claim 7 A voice processing method according to claim 1, wherein the step of generating the query text comprises: generating a query text related to the domain when the entity name is missing in the first slot; generating a query text related to the intent when the entity name is missing in the second slot after the entity name is filled in the first slot; and generating a query text related to the slot where the quantity entity name is missing when the entity name is missing in the third slot after the entity name is filled in the second slot. Claim 8 A voice processing method according to claim 1, wherein the completing step comprises: a step of filling the required slots where the entity names are missing with entity names through feedback of the query utterance voice and the repeated reception of the user's response utterance voice; and a step of completing the user's utterance voice into the complete utterance voice when the required slots where the entity names are missing are all filled with entity names. Claim 9 A voice processing method according to claim 8, further comprising the step of determining and operating one of a plurality of electronic devices in response to a command included in the complete voice of the user when the voice of the user is completed as the complete voice. Claim 10 A computer-readable recording medium storing a computer program for executing any one of the methods of claims 1 to 4 and claims 7 to 9 using a computer. Claim 11 A voice processing device comprising one or more processors, wherein the one or more processors convert a user’s spoken voice containing voice commands into user spoken text, perform grammatical analysis or semantic analysis on the user spoken text to search for a domain to which the user spoken text belongs and an intent possessed by the user spoken text, search for one or more named entities as a result of named entity recognition included in the user spoken text, determine whether the user’s spoken voice is a complete spoken voice or an incomplete spoken voice in response to the search results for the domain, the intent, and the named entities, and if the user’s spoken voice is an incomplete spoken voice, generate a query spoken voice requesting information to complete the incomplete spoken voice into the complete spoken voice and feed it back to the user, and receive the user’s response spoken voice corresponding to the query spoken voice to complete the complete spoken voice, wherein the one or more processors, when feeding back to the user, provide a necessary slot in the user command text according to the domain and intent, and corresponding to the necessary slot Based on a pre-established state table containing a query text to be requested, a slot among the necessary slots in the user utterance text in which an entity name is missing is determined; based on the state table, a query text corresponding to the missing slot is generated; and the generated query text is converted into the query utterance voice and provided as feedback to the user, wherein among the necessary slots in the user utterance text, a first slot is related to a domain, among the necessary slots in the user utterance text, a second slot is related to an intent, and among the necessary slots in the user utterance text, a third slot is related to a quantity entity name, and when the one or more processors determine the slot in which the entity name is missing, among the necessary slots in the user utterance text, the first slot,A voice processing device configured to determine the slot in which the entity name is omitted in the order of the second slot and the third slot. Claim 12 A voice processing device according to claim 11, wherein one or more processors are configured to convert a user’s spoken voice into user spoken text, the voice command including a voice command that directs the operation of at least one of a plurality of electronic devices when converting into user spoken text. Claim 13 In claim 12, the one or more processors are configured to search for an entity name including a noun or number having a unique meaning appearing in the user speech text, the domain, the intent, the domain specifying the type of electronic device that the user intends to operate when searching for the entity name, the intent indicating how to operate the electronic device. Claim 14 A voice processing device according to claim 11, wherein one or more processors, when determining whether the voice is complete or incomplete, determine a required slot including a first slot related to the domain, a second slot related to the intention, and a third slot related to a quantity entity name from the user voice text, and determine the user's voice as complete voice if the required slot in the user voice text is filled with entity names, and determine the user's voice as incomplete voice if there is a required slot in the user voice text in which entity names are missing. Claim 15 delete Claim 16 delete Claim 17 A voice processing device according to claim 11, wherein the one or more processors are configured to generate a query text related to the domain when the entity name is missing in the first slot when the query text is generated, generate a query text related to the intent when the entity name is missing in the second slot after the entity name is filled in the first slot, and generate a query text related to the slot where the quantity entity name is missing when the entity name is missing in the third slot after the entity name is filled in the second slot. Claim 18 A voice processing device according to claim 11, wherein the one or more processors are configured to fill the required slots where the entity names are missing with entity names through feedback of the query voice and repeated reception of the user's response voice when the complete voice speech is completed, and to complete the user's voice speech when the required slots where the entity names are missing are all filled with entity names. Claim 19 A voice processing device according to claim 18, wherein the one or more processors are further configured to determine and operate one of a plurality of electronic devices in response to a command included in the complete voice when the user's spoken voice is completed as the complete voice.