Multi-assistant device control
By introducing multiple assistant components into the multi-assistant voice processing system, the problem of limited information exchange between different voice processing systems is solved, and better cross-system command processing capabilities and user experience are achieved.
Patent Information
- Application Number
- CN202380071253.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-14
- Filing Date
- 2023-08-17
- Publication Date
- 2025-05-13
AI Technical Summary
In multi-assistant voice processing systems, the user's cross-assistant command processing capability is limited by information exchange between different voice processing systems, resulting in poor user experience, especially when processing device process control involving multiple systems.
Multiple assistant components are introduced, responsible for passing information and commands between different voice processing systems and arbitrating the use of device resources. Through the multi-assistant component, a voice processing system can process commands related to device control, even if the command is initially activated by another system.
Improves the ability to process commands across different voice processing systems, enhances the user experience, and ensures that users can seamlessly control the process of device involved in multiple systems, even when these processes are started by different voice processing systems.
Smart Images

Figure CN119998870A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Patent Application No. 17 / 944,600, filed on September 14, 2022, and entitled “Multiple Assistant Device Control.” Said application and this application also claim priority to U.S. Provisional Patent Application No. 63 / 400,636, filed on August 24, 2022, and entitled “Multiple Assistant Device Control.” The contents of the above applications are expressly incorporated herein by reference in their entirety. Background Art
[0003] Speech recognition systems have been developed to the point where humans can use their voices to interact with computing devices. Such systems employ techniques to recognize words spoken by a human user based on received audio input of various qualities. Speech recognition, in conjunction with natural language understanding processing techniques, enables voice-based user control of computing devices to perform tasks based on the user's spoken commands. Speech recognition and natural language understanding processing techniques may be referred to herein together or separately as speech processing. Speech processing may also involve converting a user's speech into text data, which may then be provided to various text-based software applications.
[0004] Speech processing can be used by computers, handheld devices, telephone-computer systems, kiosks, and a wide variety of other devices to improve human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] For a more complete understanding of the present disclosure, reference is now made to the following description taken in conjunction with the accompanying drawings.
[0006] Figure 1 is a conceptual diagram illustrating components of a virtual assistant system having a protected cross-assistant command processing feature according to an embodiment of the present disclosure;
[0007] FIG. 2A to FIG. 2B is a signal flow diagram illustrating example operations for protected cross-assistant command processing according to an embodiment of the present disclosure;
[0008] FIG. 3A to FIG. 3B is a flow chart illustrating example operations for performing cross-assistant command processing according to an embodiment of the present disclosure;
[0009] Figure 4 is a conceptual diagram of components of a speech processing system according to an embodiment of the present disclosure;
[0010] Figure 5 is a conceptual diagram showing components that may be included in a device according to an embodiment of the present disclosure;
[0011] Figure 6is a conceptual diagram of an automatic speech processing component according to an embodiment of the present disclosure;
[0012] Figure 7 is a conceptual diagram of how natural language processing is performed according to an embodiment of the present disclosure;
[0013] Figure 8 is a conceptual diagram of how natural language processing is performed according to an embodiment of the present disclosure;
[0014] Fig. 9 is a conceptual diagram of a text-to-speech component according to an embodiment of the present disclosure;
[0015] Fig.10 is a block diagram conceptually illustrating example components of an apparatus according to an embodiment of the present disclosure;
[0016] Fig.11 is a block diagram conceptually illustrating example components of a system according to an embodiment of the present disclosure; and
[0017] Fig.12 An example of a computer network for use with the overall system according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0018] The speech processing system and the speech generation system can be combined with other services to create a virtual "assistant" that users can interact with using natural language input (such as voice, text input, etc.). The assistant can utilize different computerized speech-enabled technologies. Automatic speech recognition (ASR) is a field of computer science, artificial intelligence, and linguistics that involves converting audio data associated with speech into text or other types of words representing the speech. Similarly, natural language understanding (NLU) is a field of computer science, artificial intelligence, and linguistics that involves enabling computers to derive meaning from text or other natural language meaning-expressing data. ASR and NLU can be used together as part of a speech processing system, sometimes also referred to as a spoken language understanding (SLU) system. Text-to-speech (TTS) is a field of computer science that involves converting text and / or other meaning-expressing data into audio data that is synthesized to resemble human speech. ASR, NLU, and TTS can be used together to act as a virtual assistant that responds to verbal commands and responds with synthesized speech. For example, an audio-controlled user device and / or one or more speech processing systems may be configured to receive human speech and detect wake-up words and / or other natural language inputs for activating the device. The device and / or system may determine the command represented by the user input and use TTS and / or other system commands to provide a response (e.g., in the form of synthesized speech, a command to send audio to a different device / system component, etc.).
[0019] Some audio control devices may provide access to more than one speech processing system, each of which may provide services associated with a different virtual assistant. In such a multi-assistant system, one or more speech processing systems may be associated with its own set of wake-up words (which are used to invoke the speech processing system) and other observable characteristics (such as voice characteristics and other auditory or visual indicators) that allow the user to identify which speech processing system the user is interacting with.
[0020] In some cases, the entire system can be configured in a way that does not allow certain communications / operations between speech processing systems. There may be many reasons for this. First, a user may have a set of permissions for a first speech processing system that does not allow certain data to be shared with a second speech processing system. In addition, centralized user settings may not allow user speech information to be shared with speech processing systems that are not directly called by the user. Therefore, in order to enhance the perceived protection of user privacy, communication between different speech processing system components can be blocked. Secondly, although the first speech processing system and the second speech processing system can each be called from the first device, the system itself may not want to share information directly between them. This may be true in the following case: the first speech processing system and the second speech processing system have different speech processing architectures / pipelines (rather than sharing many speech processing components), so that the command to call the first speech processing system goes to one set of devices / components for processing, while the command to call the second speech processing system goes to another set of devices / components for processing. This may be true in the case where the first speech processing system and the second speech processing system are managed by competing entities. Therefore, direct sharing of information between different speech processing system components may not be allowed. Therefore, in some overall system configurations, components of the first speech processing system may not be configured to communicate with components of the second speech processing system.
[0021] However, in some cases, a user may speak a command using the wake word of one system, when the command is actually related to a different system. For example, a user may use the wake word of a first assistant / system to start a timer (e.g., "Alexa, set a 10-minute timer"), but when the timer ends and the device beeps, the user may attempt to stop the timer using the wake word of a second assistant / system (e.g., "Kitchen, cancel timer"). If a first speech processing system (such as the system associated with the wake word "Alexa") is separate from a second speech processing system (such as the system associated with the wake word "Kitchen"), the user's command to stop the timer may not be understood by the second speech processing system because the second speech processing system may not be able to correctly process the command "cancel timer" because it does not have any information about the ongoing timer because the initial command to set the timer was processed by the first speech processing system. In this case, the second speech processing system may return an error to the user (e.g., outputting the audio "Sorry, no timer set"), which may cause frustration to the user, especially if the user may not remember which wake word was said when starting the timer, and may therefore be unsure how to stop the beeping.
[0022] In another example, a first user may use a first assistant / system to start a process (e.g., a timer, playing media, opening a camera feed, etc.), but a second user may use a second assistant / system to attempt to control the process (e.g., by stopping the process, skipping a song, etc.). This may occur on a shared device where the first user is interacting with the shared device using the first assistant and the second user is interacting with the shared device using the second assistant. If the first user starts a process using the device and leaves the room, the second user may experience difficulty using the second assistant to control the ongoing process. This occurs because, similar to the timer example above, the second assistant may not have information about the process, which can result in output errors and a poor user experience.
[0023] Techniques and components are provided for improving the ability to process commands across speech processing systems, particularly when the speech processing systems may have limitations on the information that can be directly exchanged between them. A device may include multi-assistant components that can pass information and commands between components dedicated to specific speech processing systems and / or arbitrate the use of device resources. A single speech processing system may also have information indicating which commands are related to device control (e.g., which commands should be processed by a device control skill) so that commands related to device control can be processed by one system even though the commands may be related to device control of another system. Although the disclosure below describes operation with respect to two speech processing systems, the teachings herein may be applicable to configurations with more than two speech processing systems.
[0024] The speech processing system may be configured to determine when user input requests control of a device process. In this case, the system sends the requested data to a special device skill, which can communicate with a dedicated component on the device that can manage device process commands, communicate with multiple speech processing systems, and arbitrate between the systems. The device skill sends the command to the dedicated component on the device, enabling one speech processing system to control a device process, even though the device process may have been initiated by a different speech processing system.
[0025] A device process may involve controlling a process involving an action to be performed by a device. Such device process control may include, for example, starting / stopping a timer, setting / stopping an alarm, playing / stopping media content (such as a song, video, podcast, etc.), controlling output content (such as skipping a song, rewinding a song, extending / pausing a timer / alarm, stopping synthesized speech output, etc.), setting a temperature (e.g., if the device can operate as a thermostat), activating / deactivating components of the device (such as a camera, lights, etc.), controlling device settings (such as volume, brightness, sensitivity, etc.), setting / controlling reminders, initiating / controlling / terminating a call or call request, etc. Thus, a device process control may control a device to transition from a first state (e.g., outputting audio, displaying something on a display) to a second state (e.g., stopping audio output, outputting audio at a different volume, displaying something else on a display, removing something from a display, etc.).
[0026] Although various systems are referred to herein as "speech processing systems," they may also be considered natural language processing systems because they may be configured to process natural language input, which is not necessarily verbal and may be input using some other method, such as text input to an application (or the like), where the application may correspond to a particular assistant / system. Thus, the input and output of the device need not be (or represent) spoken language. In some implementations, depending on the system configuration, a user may be able to enter natural language input via text, Braille, American Sign Language (ASL), etc. Other inputs that trigger the processing system are also possible, such as sound events (e.g., a baby crying, footsteps), button presses, etc.
[0027] The system may be configured to incorporate user permissions and only perform the activities disclosed herein if approved by the user. Thus, the systems, devices, components, and techniques described herein will generally be configured to limit processing where appropriate and only process user information in a manner that ensures compliance with all appropriate laws, regulations, standards, etc. The systems and techniques may be implemented on a geographic basis to ensure compliance with laws in various jurisdictions and entities where components of the system and / or users are located.
[0028] Figure 1 is a conceptual diagram showing components of a virtual assistant system 100 having a protected cross-assistant command processing feature according to an embodiment of the present disclosure. The virtual assistant system 100 may include an audio-enabled device 110, a first natural language / speech processing system 120a (which may be abbreviated as "first system 120a"), and a second natural language / speech processing system 120b (which may be abbreviated as "second system 120b"). The first system 120a and the second system 120b may be collectively referred to as "system 120". Although Figure 1 The first system 120a and the second system 120b are shown to have similar components arranged in a similar manner, but the components, functions, and / or architectures of the first system 120a and the second system 120b may be different. In addition, some or all of the components and / or functions of one or both of the first system 120a and / or the second system 120b may reside on or be performed by the device 110. Figure 4 and Figure 5 Other possible arrangements of components and functionality of device 110 and system 120 are described in greater detail.As noted above, although two systems 120 are shown for purposes of illustration, any number and variety of systems 120 may be supported and coordinated consistent with the principles and processes described herein.
[0029] The device 110 may receive audio corresponding to a spoken natural language input originating from a user (not shown). The device 110 may process the audio after detecting a wake word. The wake word may be a word or a phrase that, when detected, causes the device 110 to invoke the speech processing system 120 to process audio data accompanying or containing the wake word. The wake word may be specific to a particular speech processing system 120. Thus, if the device 110 detects a first wake word, it may route data corresponding to speech to the first speech processing system, and if the device 110 detects a second wake word, it may route data corresponding to speech to the second speech processing system. (The device 110 may also be configured to detect any number of wake words having any correlation with a set of available speech processing systems 120, so that no wake word is associated with more than one speech processing system 120.) The device 110 may generate audio data corresponding to the audio / speech, and may send the audio data to the first system 120a and / or the second system 120b. The device 110 may send the audio data to the system 120 via one or more applications installed on the device 110. An example of such an application is the Amazon Alexa application that can be installed on a smartphone, tablet, etc. In some implementations, the device 110 can receive text data corresponding to natural language input from the user 5 and send the text data to one of the systems 120. The device 110 can receive output data from the system 120 and generate synthesized speech output and / or perform an action. The device 110 may include a camera for capturing image and / or video data for processing by the system 120. Fig.12 Examples of various devices 110 are further shown in FIG.
[0030] The system 120 may include supporting components of local systems and / or remote systems, such as a set of computing components located geographically away from the device 110 but accessible via the network 199 (e.g., a server accessible via the Internet). The system 120 may also include remote systems that are physically separated from the device 110 but located geographically close to the device 110 and accessible via the network 199 (e.g., a home server located in the same residence as the device 110). The system 120 may also include some combination thereof, such as certain components / operations being performed by device components or a home server, while other components / operations are performed by a geographically remote server. Although the figures and discussions of the present disclosure describe certain steps in a particular order, the steps described may be performed in a different order (and certain steps may be removed or added) without departing from the present disclosure.
[0031] The device 110 may include a microphone 114 for receiving audio and a speaker 112 for transmitting audio. The device 110 may include one or more wake-up word detectors 121 capable of detecting one or more wake-up words. In some implementations, the wake-up word detector 121 may be embedded in a processor chip; for example, a digital signal processor (DSP). In some implementations, the wake-up word detector 121 may be an application-driven software component. In some cases, a single wake-up word detector 121 may be capable of detecting multiple wake-up words for more than one system. In other cases, the device 110 may include multiple wake-up word detectors, such as a first wake-up word detector 121a and a second wake-up word detector 121b, each of which is capable of detecting its own wake-up word. For example, the first wake-up word detector 121a may detect one or more wake-up words associated with the first system 120a, and the second wake-up word detector 121b may detect one or more wake-up words associated with the second system 120b.
[0032] The device may include one or more assistant components 140, including a first assistant component 140a and a second assistant component 140b. The assistant component 140 may be coupled to one or more systems in the system 120. Figure 1 In the example system 100 shown, a first assistant component 140a communicates with a first system 120a, and a second assistant component 140b communicates with a second system 120b. In some implementations, a single assistant component 140 may handle communications with more than one system 120. The device 110 may have a dedicated assistant component 140 for one system 120, or a single assistant component 140 that communicates with all systems 120. The device may include a multi-assistant component 115 for managing multi-assistants and cross-assistant operations of the device 110 as described herein. The device may also include a set of components for storing / tracking state data 194. (As described below, the state data 194 may be tracked and maintained separately by each assistant component 140 and by the multi-assistant component 115.) Such state data 194 may indicate the state of the device 110 (and / or a user profile corresponding to the device 110) and may correspond to one or more processes of the device. Examples of state data may include volume level, data indicating what is being displayed on a display, time data, network access data, timer status, etc. State data 194 may be stored on device 110 or on another device (such as a remote device, a home server, etc.). Fig.10 Additional components of device 110 are described in more detail.
[0033] In some configurations, in order to maintain a sense of privacy and / or other separation between speech processing systems, the first assistant component 140a may not be configured to communicate with the second assistant component 140b without routing the communication through the multi-assistant component 115. In this way, the multi-assistant component 115 can mediate interactions between speech processing system components. Similarly, the multi-assistant component 115 (or other remote / cloud component) can mediate communications between the first system 120a and the second system 120b. Therefore, the speech processing systems may not be configured to communicate directly, particularly when such communications may involve specific utterances being processed. Although illustrated as physically running on the device 110, the multi-assistant component 115 can run on different physical devices (e.g., a home server, etc.). In this (or other) cases, the multi-assistant component 115 can coordinate multi-assistant operations of multiple devices 110, where such devices 110 can be associated with one or more user accounts. For example, a single multi-assistant component 115 can coordinate multi-assistant operations of multiple devices associated with a particular user / user profile, a family / family profile / multiple user profiles, etc.
[0034] As part of this separation, in certain configurations, each speech processing system and / or components associated therewith may store / manage their own state data 194 about the device. For example, a first assistant component 140a associated with a first system 120a may store / manage state data 194a including data about the interaction / operation of the device 110 (and / or a user profile associated with the device 110) and the first system 120a. For example, if a user interacts with the device 110 to invoke the first system 120a (e.g., by speaking a first wake word associated with the first system 120a), the first assistant component 140a may save certain information about the interaction between the device 110 and the first system 120a as state data 194a. Thus, if a device process is initiated as a result of a command to the first system 120a, the first assistant component 140a may store information about the device process as state data 194a. For example, if the user starts a timer by invoking the first assistant associated with the first system 120a, the state data 194a may reflect the start time of the timer, the remaining time, a tag associated with the timer, etc. In another example, if the user starts playing music by invoking the first assistant associated with the first system 120a, the state data 194a may reflect the start of the music, the source of the music content (e.g., a music service), information about the currently playing music, information about the previously played music, etc.
[0035] The second assistant component 140b may also store / manage its own state data 194b regarding the interaction / operation of the device 110 (and / or a user profile associated with the device 110) and the second system 120b. Such management of the state data 194b regarding the second system 120b may operate similarly to the management described above regarding 194a and the first system 120a. However, as part of the system separation, the first assistant component 140a may not be able to access the state data 194b, and the second assistant component 140b may not be able to access the state data 194a. Therefore, each system / assistant component may only track state data 194 regarding its own operations.
[0036] Certain status data 194 may also be stored / managed by the multi-assistant component 115. Such status data may be stored / managed by the multi-assistant component 115. Figure 1 194m. Such status data 194m may include information related to the device processes being performed on the device, and may include some (certain) portion of the information stored in 194a / 194b and / or other information related to the management of device processes. For example, if a timer is in progress, the status data 194m may include an indicator that the timer is in progress and a system 120 for invoking the timer, but may not include as many timer details as the status data 194 that invokes the system. Similarly, if music is being output, the status data 194m may indicate that the music is playing but may not include all the details of the music playing. The status data 194m may indicate which device processes are active at any particular point in time (e.g., a timer is in progress, a timer ends and emits a beep, music is playing, etc.). The status data 194m may indicate which device controls are executable for a particular device process (whether or not in progress). For example, the status data 194m may indicate whether the device 110 is able to stop, extend, pause a timer; stop, pause, adjust the volume of music playing, etc. The state data 194m may also indicate which channels (e.g., hardware components) are currently being used by what processes, etc. Information may be exchanged between the multi-assistant component 115 and the single assistant component 140 to update the corresponding state data and / or to perform control of the device 110 / device process. For example, an application programming interface (API) or other interface, a registration process, etc. may be used to coordinate between the multi-assistant component 115 and the single assistant component 140 to exchange information about the state / process of the device 110.
[0037] The system 120 may include various components for processing natural language commands. The system 120 may include a language processing component 192 for performing operations related to understanding natural language (such as ASR, NLU, entity resolution, etc.). The system 120 may include a language output component 193 for performing operations related to generating natural language output (such as TTS). The system 120 may also include a component for tracking system state data 195. Such system state data 195 may indicate the state of the operation of each system 120 (e.g., with respect to a specific device 110, user profile, etc.). For example, the state data 195 may include conversation data, an indication of previous utterances, whether the system 120 has any ongoing processes for the device 110 / user profile, etc. The system 120 may include one or more skill components 190. The skill component 190 may perform various operations related to executing commands, such as online shopping, streaming media, controlling smart appliances, etc.
[0038] One of the skills available to a system 120 may include a device skill 191. Such a device skill may be configured to process and manage specific utterances related to controlling a device process or device state. Each system 120 may have its own device skill 191 and / or a central device skill 191 may be accessible to multiple systems 120. Each device skill 191 may be associated with its own skill processing component 125 (discussed below).
[0039] The device skill 191 may be configured to communicate with the device 110 through the assistant component 140. Thus, the device skill 191 may send commands to control the device 110 (in coordination with the multiple assistant component 115) through the assistant component 140. Thus, the system 120 may send commands / messages / directions to the multiple assistant components 115 by routing such communications with the respective assistant components 140. For example, a device control command from the first system 120a may be sent from the device skill 191a to the first assistant component 140a, which then routes the device control command to the multiple assistant components 115.
[0040] FIG. 2A to FIG. 2B is a signal flow diagram illustrating example operations for protected cross-assistant command processing according to an embodiment of the present disclosure. Figure 2A It is shown that the user can instruct the first system 120a to start the operation of the device process. Figure 2B A number of operations are shown in which a user can direct the second system 120b to use components (such as 115, 191, etc.) to control Figure 2A In the embodiment of the present invention, the first system 120a and the second system 120b are connected to each other by using the same device process started by the first system 120a, rather than by allowing direct communication between the first system 120a and the second system 120b. Specifically, Figure 2AOperations between the microphone 114, the speaker 112, the wake-up word detector 121, the first assistant component 140a, the multi-assistant component 115, and the second assistant component 140b of the device 110 and the first system 120a, the device skill 191a (which may be of the first system 120a), and the second system 120b are shown.
[0041] like Figure 2A As shown, microphone 114 may receive an audio signal and send (202) the audio data to wake-up word detector 121. The audio data may represent, for example, a natural language command, such as: "Alexa, set a timer for 10 minutes." Wake-up word detector 121 may detect the wake-up word "Alexa" corresponding to first speech processing system 120a and first assistant component 140a. Wake-up word detector 121 may notify (204) multi-assistant component 115 that the first wake-up word was detected in the input.
[0042] In some implementations, the device 110 may receive input data in other formats, such as typed or scanned text, Braille, or American Sign Language (ASL) (e.g., detected by processing image data and / or sensor data representing a user communicating in ASL). The device 110 may determine that the input data is to be processed by the first system 120a based on other indications (such as a button press) or because the first system 120a represents the default system 120 for executing commands from the device 110.
[0043] The multi-assistant component 115 may signal (206) the first assistant component 140a that the first assistant component 140a may send data (e.g., audio data) representing a command to the first system 120a. After the multi-assistant component 115 confirms the call of the first system 120a (e.g., by detecting the first wake word through the wake word detector 121), the audio data of the utterance may be sent to the first assistant component 140a. The audio data may be sent (207) to the first assistant component 140a by the multi-assistant component 115 or another component. The multi-assistant component 115 may also send (208) state data 194 corresponding to the state of the device 110 and / or a user profile corresponding to the device 110 to the first assistant component 140a. The state data being sent (208) may be data available to the multi-assistant component 115 (e.g., obtained from the state data 194m). Such state data may indicate one or more process controls that can be executed by the device 110. In some cases, such state data being sent may only apply to the active process of the device 110. Thus, the state data sent may include metadata corresponding to processes and / or process controls of device 110. First assistant component 140a may send (210) audio data representing a command to first system 120a. First assistant component 140a may also send (212) state data to first system 120a. The state data sent (212) from first assistant component 140a to first system 120a may include all or some of the state data sent from multi-assistant component 115 to first assistant component 140a in step 208 above (e.g., a portion of state data 194m). The state data sent (212) from first assistant component 140a to first system 120a may also include certain state data (e.g., state data 194a) available to first assistant component 140a, which may be different from state data 194m. The state data may indicate one or more active hardware components of device 110, which may indicate activity on certain processing channels of the device. The system may use specific identifiers (such as speech identifiers) to track speech-related data across different components. For example, when receiving audio data of an utterance, the device 110 may assign a specific identifier to the audio data. The identifier may be sent together with the WW detection indication (204), the WW signal (206), the status data (208 / 212), the audio data (210), etc., so that the device 110, the system 120a, etc. can track information related to the specific utterance.
[0044] The device / profile state data 194 (corresponding to one of the assistant components 140 and / or multiple assistant components 115) that is ultimately sent to the speech processing system 120 may be obtained from storage on the device 110, or from storage on a different device corresponding to the same user profile. The state data 194 (particularly the state data 194m) may indicate what processes are ongoing and / or controllable using the device 110 (and / or other devices associated with the user profile associated with the state data 194). For example, if the device 110 has an ongoing conversation and a timer, the state data 194m may indicate that <dialog><dialogue>; <timer>. The status data 194m may also indicate what commands are executable (which may be related to the device process). For example, the status data 194m may indicate <stop dialogue>; <stop timer>; <adjust timer>, etc. The status data 194m may indicate many different controllable device processes, or may only indicate active device processes. The status data 194m may also indicate the type of command. For example, certain commands may be related to dialogue control, alarm control, content control, etc. These types may be indicated in the status data 194m. The status data 194m may also indicate priority information associated with a specific device process. For example, if the device is playing music and also outputs a beep corresponding to an expired timer, the status data 194m may indicate that the timer has a higher priority than the music playback. Depending on the system configuration, user permissions / privacy settings, etc., the status data sent (208 / 212) may be limited in some way, such as only indicating ongoing device processes, to reduce the amount of shared status data.
[0045] The first system 120a may perform speech processing (e.g., using the language processing component 192a and corresponding operations described herein) to determine whether an incoming user request corresponds to a request to control a device process and, therefore, should be routed (214) to a device skill. Such a determination may involve processing both audio data and other data indicating which potential commands / requests may be associated with control of a device process. The other data may include state data (e.g., state data 194) corresponding to a device / user profile. State data 194 may indicate what commands correspond to a device process, thereby allowing the first system 120a to correctly determine when an incoming request corresponds to a device process, thereby indicating whether processing for the request should be handled by a device skill. Such a device skill may be configured to send commands to the multi-assistant component 115 (using the routing discussed herein) in order to indicate to the device one or more instructions for controlling the device process. Such instructions allow the multi-assistant component 115 to coordinate control of the device with the assistant component 140 and other components. Although Figure 2A (and Figure 2B ) shows that the device / profile state data 194 is sent from the multi-assistant component 115, but the first system 120a can also obtain the device / profile state data 194 from another source, such as an off-device storage component in communication with the first system 120a that can provide access to state data, which can be similar to the state data 194m (or other state data 194). Such a determination may also involve processing other data indicating the processing capabilities of the device 110 and / or other devices corresponding to the user profile.
[0046] If the incoming request does correspond to a device skill (e.g., the request corresponds to control of a device process) as determined through language processing (e.g., using other data / state data 194), the first system 120a can route (216) data related to the request to the first system's device skill 191a. Such data can include NLU result data (such as NLU output data 885 / sorted output data 825, as discussed below) and / or can include processed data that is based on such NLU result data but has been converted into specific commands specific to the device process.
[0047] In the example of the utterance "Alexa, set a timer for 10 minutes," the first system 120a may use state data 194 to determine that the device 110 is capable of operating a timer, and may determine that the command corresponds to a request to control a device process and may determine that a device skill should be invoked. Accordingly, the first system 120a may send data representing the request to the device skill 191a, such as an instruction to start a timer, etc. The device skill 191a may then generate (218) output data to be included in a message to the device. The device skill 191a may send (220) the message data back to the device 110. Such a message may be routed through the first assistant component 140a. The message data may include instructions for controlling a device process, Figure 2A An example of may be instructions to start a timer. The first assistant component 140a may receive the message data and may determine that the message data includes instructions corresponding to a device process control. The message data may also (or alternatively) include information indicating that the message corresponds to a process to be managed by the multi-assistant component 115. The first assistant component 140a may process the message data to determine (based on the instructions and / or routing instructions) that the request should be processed by the multi-assistant component 115. The first assistant component 140a may then route the message / instruction data (or a portion thereof) (223) to the multi-assistant component 115. Therefore, since the message data may come from a component associated with the first system 120a (device skill 191a), the message data may first be routed through the first assistant component 140a and then sent to the multi-assistant component 115.
[0048] The message / output data may include an utterance identifier linking the particular output data to the input data (e.g., audio data 202). In this manner, the first system 120a may indicate to the device 110 that the particular message / output data matches the particular input received by the device 110. The message / output data may include instructions / commands to the device 110 to start a 10 minute timer. The instructions may also include instructions as to the assistant / system that the user invoked during the request (e.g., "Alexa") so that the appropriate system can track the timer for management purposes. Thus, in Figure 2A In an example, the message data may include an indication of the first system 120a.
[0049] The message data may also include data to be output as a confirmation of the command, such as synthesized speech confirming the command, data confirming the command to be displayed on a display of device 110, etc. If such a confirmation is included, the confirmation may be output by the device, for example by sending (222) output audio of the confirmation message from first assistant component 140a (or other component) to speaker 112 for output as synthesized speech (e.g., "Start your timer now").
[0050] First assistant component 140a and / or multi-assistant component 115 may then execute (224) the instructions received from device skill 191a. In the above example, this may include executing a request to start a timer. Multi-assistant component 115 may receive the instruction / message data and process it to determine that the requested control is a control that can be executed by device 110. Multi-assistant component 115 may also determine that the process to be controlled is related to first assistant component 140a. For example, the request may also indicate the assistant that the user originally invoked for the timer. Alternatively (or in addition), the request may indicate an identifier of the original utterance. The identifier may be used by multi-assistant component 115 to identify the wake word that accompanied the original utterance, thereby allowing multi-assistant component 115 to determine the assistant that was originally invoked. Multi-assistant component 115 may then instruct first assistant component 140a (and / or other components) to take the actions required to start the timer (or otherwise execute the instructions to control the device process). First assistant component 140a may then take actions to execute the instructions (e.g., start the timer) and may indicate to multi-assistant component 115 that the timer has been started. For example, the first assistant component 140a may register information about timer controls with the multi-assistant component 115, such as what controls are available for the timer, and / or other timer data.
[0051] The multi-assistant component 115 may then update its device / user profile state data 194m to record data related to the timer. For example, the multi-assistant component 115 may update (225) the state data 194m to indicate that the timer is active and that the device 110 may be configured to respond to commands to control / stop a particular active timer. The first assistant component 140a may also update (226) its own state data 194a to indicate information about the timer, and may send (228) a message to the first system 120a regarding the requested start of the timer. The first system 120a may then take any desired actions related to the start of the timer, and may update (230) its own state data (e.g., system state data 195a) locally and remotely to indicate the requested timer (e.g., the length of the timer, the timer start time, the timer end time, the associated device and / or user profile), etc.
[0052] In this way, if Figure 2A As reflected, the user may invoke the first assistant to initiate a command to control a device process (e.g., set a 10 minute timer). The invoked first system 120a may process the request, determine that the request is related to a process to control the device 110, and route information about the request through the multi-assistant component 115, which in turn may utilize the first system 120a to manage the process.
[0053] like Figure 2B As shown, the device flow may also be controlled using a second, invoked assistant different from the initially invoked assistant, even if the respective systems of those assistants do not directly communicate or are not aware that another assistant is accessible via the device 110. The multiple assistant component 115 may allow operation as described herein.
[0054] continue Figure 2A In the example shown in , after 10 minutes, the timer may expire and the device 110 may be outputting audio corresponding to the end of the timer, such as a beep, etc. Figure 2B In an example of , a user (who may or may not be the same user who started the timer) may speak a command to end the timer, only accompanied by a different wake word (for a different assistant / system) than the original request to start the timer. For example, the command to terminate the timer may be "kitchen, turn off timer," where "kitchen" is a second wake word associated with a second assistant / second system 120b that is different from the first assistant / first system 120a.
[0055] like Figure 2B As shown, microphone 114 detects the audio of the speech and sends (232) the audio data to wake-up word detector 121. Wake-up word detector 121 may detect the wake-up word "kitchen" corresponding to second speech processing system 120b and second assistant component 140b. Wake-up word detector 121 may notify (234) multi-assistant component 115 that the second wake-up word is detected in the input.
[0056] The multi-assistant component 115 may signal (236) the second assistant component 140b that the second assistant component 140b may send data representing the command to the second system 120b. After the multi-assistant component 115 confirms the invocation of the second system 120b (e.g., by detecting the second wake word through the wake word detector 121), the audio data of the speech may be sent to the second assistant component 140b. The audio data may be sent (237) to the second assistant component 140b by the multi-assistant component 115 or another component. The multi-assistant component 115 may also send (238) to the second assistant component 140b state data 194 corresponding to the state of the device 110, including processes and state data shared by other assistants, and / or a user profile corresponding to the device 110. The state data being sent (238) may be data available to the multi-assistant component 115 (e.g., obtained from the state data 194m). Such state data may indicate one or more process controls that can be executed by the device 110. In some cases, such state data being sent may only apply to the active process of the device 110. Thus, the transmitted state data may include metadata corresponding to the process and / or process control of device 110. Figure 2B In an example of , the status data being sent may indicate an active timer and one or more controls that the device 110 may perform regarding the timer (e.g., stopping the timer, pausing the timer, etc.). The second assistant component 140b may send (240) audio data representing a command to the second system 120b. The second assistant component 140b may also send (242) status data to the second system 120b. The status data sent (242) from the second assistant component 140b to the second system 120b may include all or some of the status data sent from the multi-assistant component 115 to the second assistant component 140b in step 238 above (e.g., a portion of the status data 194m). The status data sent (242) from the second assistant component 140b to the second system 120b may also include certain status data available to the second assistant component 140b (e.g., status data 194b), which may be different from the status data 194m. The status data may indicate one or more active hardware components of the device 110, which may indicate activity on certain processing channels of the device. and Figure 2A As with the first utterance of the present invention, the system may use a specific identifier (such as an utterance identifier) to track data related to the second utterance across different components. The second identifier may be sent together with the WW detection indication (234), the WW signal (236), the status data (238 / 242), the audio data (240), etc., so that the device 110, the system 120b, etc. can track information related to the second utterance.
[0057] The status data 194 sent (238) to the second assistant component 140b and / or sent (242) to the second system 120b may indicate that a timer control is available, that an initially requested timer has expired and is causing the output of a corresponding audio (e.g., a beep), or similar status information. In some implementations, the status data shared by the multi-assistant component 115 may be selected to be as minimal as possible so that each assistant is aware of the shared status data from the other assistants, such as a list of shared controls, processes, focus channels, etc. As described above with respect to Figure 2A As described in the example, Figure 2B In an example, second system 120b may obtain status data from device 110 (e.g., via second assistant component 140b) and / or from another source of status data 194.
[0058] The second system 120b may perform speech processing (e.g., using language processing component 192b and corresponding operations described herein) to determine whether an input user request corresponds to a request to control a device process and, therefore, should be routed (244) to a device skill. Such a determination may involve processing both audio data and other data indicating which potential commands / requests may be related to control of a device process. The other data may include state data corresponding to the device / user profile (e.g., state data 194), including state data shared by other assistants via a multi-assistant component. Such a determination may also involve processing other data indicating processing capabilities of the device 110 and / or other devices corresponding to the user profile.
[0059] exist Figure 2B In an example of , the second system 120b / language processing determines that the incoming request does correspond to a device skill (e.g., the request corresponds to control of a device process, such as ending a timer), and therefore the second system 120b can route (246) data related to the request to the device skill 191b of the second system 120b. Such data can include NLU result data (such as NLU output data 885 / sorted output data 825 as discussed below) and / or can include processed data that is based on such NLU result data but has been converted into specific commands specific to the device process.
[0060] In the example of the utterance "Kitchen, turn off the timer," the second system 120b may use the state data 194 to determine that the device 110 can process the <stop timer> command, and therefore the second system 120b may determine that the command corresponds to a request to control a device process. Note that in this particular example, the second system 120b may not have information indicating whether the timer is active relative to the first system 120a or which assistant controls the timer. Instead, the second system 120b may have interpreted the command to end the timer, determined that such a command is related to a device process, and therefore routed processing of the command to the device skill 191b (and ultimately to the multi-assistant component 115), even though the second system 120b may not have information about the original timer. In some implementations, the selected device skill 191b is specific to a process managed or arbitrated by the multi-assistant component 115.
[0061] The second system 120b may thus invoke the device skill. The second system 120b may send (246) data representing the request to the device skill 191b, such as instructions to send a timer, etc. The device skill 191b may then generate (248) output data to be included in a message to the device. The device skill 191b may send (250) the message data back to the device 110. Such a message may be routed through the second assistant component 140b. The message data may include instructions to control a device process, in this example, instructions to stop a timer. The second assistant component 140b may receive the message data and may determine that the message data includes instructions corresponding to a device process control (e.g., to stop a timer). The message data may also (or alternatively) include information indicating that the message corresponds to a process to be managed by the multi-assistant component 115. The second assistant component 140b may process the message data to determine (based on the instructions and / or routing instructions) that the request should be processed by the multi-assistant component 115. The second assistant component 140b may then route the message / instruction data (or a portion thereof) (253) to the multi-assistant component 115. Therefore, since the message data may come from a component associated with the second system 120b (device skill 191b), the message data may first be routed through the second assistant component 140b and then sent to the multi-assistant component 115.
[0062] The message / output data may include an utterance identifier that links the particular output data to the input data (e.g., audio data 232). In this manner, the second system 120b may indicate to the device 110 that the particular message / output data matches the particular input received by the device 110. In this example, the message / output data may include an instruction / command to the device 110 to stop a timer.
[0063] The message data may also include data to be output as a confirmation of the command, such as synthesized speech confirming the command, data confirming the command to be displayed on a display of device 110, etc. If such a confirmation is included, the confirmation may be output by the device, such as by sending (252) output audio of the confirmation message from second assistant component 140b to speaker 112 for output as synthesized speech (e.g., "Stop timer"). The confirmation message output to the user may also be visual, such as a display element indicating that the timer is being canceled.
[0064] The multi-assistant component 115 may then take action to execute the instructions received from the device skill 191b. In the above example, this may include executing a request to terminate a timer. Since the request to stop the timer comes from the second system 120b, the message data 250 may not indicate which timer to stop. The multi-assistant component 115 may take further action to determine which timer to stop.
[0065] Specifically, the multi-assistant component 115 may receive the message data / directive and determine that the directive is related to controlling a specific device process (i.e., controlling a timer). The multi-assistant component 115 may evaluate its state data 194m to determine the timer information. Since the first assistant component 140a registers the timer information (e.g., about Figure 2A operation), such information may be available in the status data 194m. By evaluating the status data 194m, the multi-assistant component 115 may determine whether a timer associated with the device 110 (and / or a user profile for the device) is active, a timer control for the device, an assistant corresponding to the timer (e.g., the first assistant component 140a), etc. Therefore, by referring to the status data 194m, the multi-assistant component 115 may determine that the command to end the timer corresponds to a timer initially started using the first system 120a. The multi-assistant component 115 may then coordinate with the first assistant component 140a to execute (254) the instructions to end the timer. Specifically, the multi-assistant component 115 may send an instruction to the first assistant component 140a to end the timer. The first assistant component 140a may then take steps to end the timer in the normal manner (e.g., if the user presses a button to end the timer, the command to end the timer comes from the first system 120a, etc.). First assistant component 140a may also update (256) its own status data 194a to indicate information about the stopped timer, and may send (258) a message to first system 120a to inform that the requested timer has been stopped. First system 120a may then update (260) its own status data (e.g., system status data 195a) to indicate that the requested timer has been stopped. Multi-assistant component 115 may also update (255) device / user profile status data 194m to record that the timer has been stopped.
[0066] Thus, the multi-assistant component 115 may coordinate with other components of the device 110 to stop the timer, which may include stopping the output of an audio (eg, a beep) associated with the timer.
[0067] The above article about Figure 2A and Figure 2B The techniques shown can also be implemented for many other device processes, where such cross-assistant device control is coordinated by multi-assistant component 115 and device skill 191. These configurations and techniques allow the entire system to successfully process commands that would otherwise require the use of one assistant, even if the command invokes the wrong assistant, without directly sharing information between assistant / speech processing systems.
[0068] Using an example similar to that discussed above, if a user uses the first speech processing system 120a (such as the one described above using Figure 2A If, for example, a user utters a "stop" command when invoking the second speech processing system 120b, the device / profile state data 194 for processing the "stop" command may indicate what process is active / what "stop" type commands are available for the particular device / user profile. Thus, the second speech processing system 120b may use the device / profile state data 194 to interpret the command as being associated with one of the available stop commands, route processing through the device skill 191b and ultimately return to the multi-assistant component 115, so that the multi-assistant component 115 may coordinate with the first assistant 140a to ultimately stop the active process that was initiated using the first speech processing system 120a.
[0069] In some configurations, the device that captures the speech audio may be different from the actual device to be controlled. For example, a user may control a smartwatch 110c (eg, Fig.12 110c) says an utterance similar to "Alexa, stop the music." The utterance may refer to the fact that music is not playing on the smart watch 110c, but is playing on another device (e.g., on a home audio system). As described herein, a speech processing system (e.g., a first system 120a associated with the wake word "Alexa") that receives the audio data of the utterance may determine that the utterance corresponds to a device control (e.g., using a language processing component 192a) and may send data corresponding to the utterance to the device skill 191a. The first system 120a may also determine that the utterance originates from a device / user associated with a particular user profile. The first system 120a may send an indication of the user profile to the device skill 191a. Alternatively (or in addition), the device skill 191a may determine that the utterance originates from a particular device (e.g., smart watch 110c) or a user associated with a particular user profile. The first system 120a and / or the device skill 191a may determine that the user profile is also associated with an audio output device 110a that is capable of playing music, is playing music, and / or is generally capable of performing music control operations (e.g., as indicated by the state data 194 of the user profile), regardless of whether such music playback was initiated due to a command to the first system 120a or to a different system (such as the second system 120b). The device skill 191a may then determine that output data includes a command to stop music playback, and may send the output data to the audio output device 110a via the multi-assistant component 115 of the device, for example, using the method described above. Figure 2B The device skill 191a may also send other output data to the smart watch 110c to confirm receipt and / or processing of the request to stop music playback.
[0070] Figure 3A A flow chart illustrating example operations of a system for performing cross-assistant command processing is shown. As shown, system 120 may receive (302) audio data representing an utterance from a device 110 capable of interacting with multiple different assistant / speech processing systems. System 120 may receive (304) state data corresponding to a device process control that can be performed by device 110 or another device mentioned in the utterance (e.g., another device associated with the same user profile as device 110). Such state data (e.g., device / profile state data 194) may be received from device 110 or from some other data source (e.g., profile storage 470 discussed below, device skill 191a, and / or other source). System 120 may then perform (306) speech processing on the audio data to determine NLU result data. Such speech processing may use the state data 194 received above. The system may also use the state data 194 to determine (308) that the NLU result data corresponds to a request to control a device process. Thus, the system may send (310) data representing the request to device control skill 191. The data may include NLU result data or some other version of data indicating a device process to be controlled, a control to be performed, a device to be controlled, a user profile or other identifier corresponding to the request, and / or other information related to the request. The device skill 191 and / or other components may then send (312) the output data to the device 110 (e.g., the multi-assistant component 115). The output data may include commands to control the process and / or other data, such as the confirmation data discussed above. Execution of the relevant commands / directives may be coordinated between the multi-assistant component 115 and the assistant component 140. Additional details of these operations are discussed herein.
[0071] Figure 3B A flowchart illustrating example operations of a device for performing cross-assistant command processing is shown. As shown, the device 110 may capture (322) an utterance requesting control of a device process. The device 110 may determine (324) that the utterance corresponds to a call to a first speech processing system. Such a call may be the result of detecting a specific wake-up word of the first system 120a, pressing a button corresponding to the first system 120a, detecting a gesture associated with the first system 120a, and interacting with a device corresponding to the first system 120a. The device 110 may then send (326) audio data representing the utterance from the first assistant component 140a to the first system 120a. The device 110 may also send (328) state data 194 to the first system 120a. Such state data 194 may be sent by the multi-assistant component 115 (which may use the first assistant component 140a as an intermediary). The state data 194 may indicate the capabilities of the device 110, but may not indicate the system that was initially called to initiate the process to be controlled. The device 110 may then receive (330) a command to control a device process from the first system 120a (e.g., from the device skill 191a) to the multi-assistant component 115. The multi-assistant component 115 may receive the command and cause the device or associated assistant component 140 to coordinate (332) control of the requested device process (e.g., adjust volume, stop music playback, control a timer, etc.). As described above, the multi-assistant component 115 may coordinate with the assistant component 140 to perform device control / specific instructions. The multi-assistant component 115 may determine (334) that the device process is associated with the second system 120b. For example, the multi-assistant component 115 may evaluate the state data 194m, or other data indicating the source of the original command. The multi-assistant component 115 may then send (336) a process control indication to the second assistant component 140b, which may then send (338) the indication to the second system 120b, thereby allowing the second system 120b to update its own state data 195b to track control (e.g., termination, adjustment, etc.) of the process initially initiated as a result of a command involving the second system 120b.
[0072] System 100 may be used as Figure 4 The various components described above operate. The various components may be located on the same or different physical devices. Communication between the various components may be performed directly or across a network 199. The device 110 may include an audio capture component (such as a microphone or microphone array of the device 110) to capture the audio 11 and create corresponding audio data. Once a voice is detected in the audio data representing the audio 11, the device 110 may determine whether the voice is directed to the device 110 / system 120. In at least some embodiments, this determination may be made using a wake-up word detection component 420. The wake-up word detection component 420 may be configured to detect various wake-up words. In at least some instances, the wake-up word may correspond to the name of a different digital assistant. An example wake-up word / digital assistant name is "Alexa". In another example, the input to the system may be in the form of text data 413, for example as a result of a user typing input in a user interface of the device 110. Other forms of input may include indications that the user has pressed a physical or virtual button on the device 110, that the user has made a gesture, and the like. Device 110 may also capture images using camera 1018 of device 110 and may send image data 421 representing those images to system 120. Image data 421 may include raw image data or image data processed by device 110 before being sent to system 120.
[0073] The wake-up word detector 420 of the device 110 may process the audio data representing the audio 11 to determine whether speech is detected therein. The device 110 may use various techniques to determine whether the audio data includes speech. In some instances, the device 110 may apply a voice activity detection (VAD) technique. Such a technique may determine whether there is speech in the audio data based on various quantitative aspects of the audio data, such as the spectral slope between one or more frames of the audio data; the energy level of the audio data in one or more spectral bands; the signal-to-noise ratio of the audio data in one or more spectral bands; or other quantitative aspects. In other instances, the device 110 may implement a classifier configured to distinguish speech from background noise. The classifier may be implemented by techniques such as linear classifiers, support vector machines, and decision trees. In other further instances, the device 110 may apply a hidden Markov model (HMM) or a Gaussian mixture model (GMM) technique to compare the audio data with one or more acoustic models in storage, which may include models corresponding to speech, noise (e.g., ambient noise or background noise), or silence. Other additional techniques may be used to determine whether there is speech in the audio data.
[0074] Wake-up word detection can be performed without performing linguistic analysis, text analysis, or semantic analysis. Instead, the audio data representing the audio 11 is analyzed to determine whether specific characteristics of the audio data match a pre-configured acoustic waveform, audio signature, or other data corresponding to the wake-up word.
[0075] Therefore, the wake-up word detection component 420 can compare the audio data with the stored data to detect the wake-up word. A method for wake-up word detection applies a general large vocabulary continuous speech recognition (LVCSR) system to decode the audio signal, wherein the wake-up word search is performed in the resulting grid or confusion network. Another method for wake-up word detection constructs an HMM for each wake-up word and non-wake-up word speech signal respectively. Non-wake-up word speech includes other spoken words, background noise, etc. One or more HMMs can be constructed to model the non-wake-up word speech characteristics known as the filling model. Viterbi decoding is used to search for the best path in the decoding graph, and the decoded output is further processed to make a decision on the presence of keywords. This method can be extended to include discriminant information by combining a hybrid DNN-HMM decoding framework. In another example, the wake-up word detection component 420 can be built directly on a deep neural network (DNN) / recursive neural network (RNN) structure without involving HMM. This architecture can estimate the wake-up word posterior with context data by stacking frames in a context window of a DNN or using an RNN. Subsequent posterior threshold adjustment or smoothing is applied to make a decision. Other techniques for wake word detection, such as those known in the art, may also be used.
[0076] Once the wake word detector 420 detects the wake word and / or the input detector detects the input, the device 110 may "wake up" and begin transmitting audio data 411 representing the audio 11 to the system 120. The audio data 411 may include data corresponding to the wake word; in other embodiments, the portion of the audio corresponding to the wake word is removed by the device 110 before the audio data 411 is sent to the system 120. In the case of touch input detection or gesture-based input detection, the audio data may not include the wake word.
[0077] In some implementations, system 100 may include more than one system 120. System 120 may respond to different wake-up words and / or perform different categories of tasks. System 120 may be associated with its own wake-up word, so that saying a certain wake-up word causes audio data to be sent to and processed by a specific system. For example, detection of the wake-up word "Alexa" by the wake-up word detector 420 may cause audio data to be sent to system 120a for processing, while detection of the wake-up word "Mandy" by the wake-up word detector may cause audio data to be sent to system 120b for processing. Systems may have separate wake-up words and systems for different skills / systems (e.g., "Dungeon Master" for gameplay skill / system 120c), and / or these skills / systems may be coordinated by one or more skills 490 of one or more systems 120.
[0078] After being received by system 120, audio data 411 may be sent to orchestrator component 430. Orchestrator component 430 may include memory and logic that enables orchestrator component 430 to transmit various data segments and various forms of data to various components of the system and to perform other operations as described herein.
[0079] The orchestrator component 430 may send the audio data 411 to the language processing component 192. The language processing component 192 (sometimes also referred to as a spoken language understanding (SLU) component) includes an automatic speech recognition (ASR) component 450 and a natural language understanding (NLU) component 460. The ASR component 450 transcribes the audio data 411 into text data. The text data output by the ASR component 450 represents one or more (e.g., in the form of an N-best list) ASR hypotheses that represent the speech represented in the audio data 411. The ASR component 450 interprets the speech in the audio data 411 based on similarities between the audio data 411 and a pre-established language model. For example, the ASR component 450 may compare the audio data 411 with models of sounds (e.g., acoustic units such as phonemes, phonemes, phonemes, etc.) and sound sequences to identify words that match the sound sequences of the speech represented in the audio data 411. The ASR component 450, in some embodiments, sends the text data generated thereby to the NLU component 460 via the orchestrator component 430. The text data sent from the ASR component 450 to the NLU component 460 may include a single highest scoring ASR hypothesis, or may include an N-best list that includes multiple ASR hypotheses. The N-best list may additionally include individual scores associated with each ASR hypothesis represented therein. Figure 6 The ASR component 450 is described in more detail.
[0080] The speech processing system 192 may also include an NLU component 460. The NLU component 460 may receive text data from the ASR component. The NLU component 460 may attempt to semantically interpret the phrases or sentences represented in the text data input into the component by determining one or more meanings associated with the phrases or sentences represented in the text data. The NLU component 460 may determine an intent representing an action that the user wishes to perform, and may determine information that allows the device (e.g., the device 110, the system 120, the skill component 490, the skill processing component 125, etc.) to perform the intent. For example, if the text data corresponds to "play Beethoven's 5th symphony", the NLU component 460 may determine the intent of the system to output music, and may identify "Beethoven" as the artist / composer and "5th symphony" as a piece of music to be played. For another example, if the text data corresponds to "how is the weather", the NLU component 460 may determine the intent of the system to output weather information associated with the geographic location of the device 110. In another example, if the text data corresponds to “turn off the lights,” the NLU component 460 may determine that the system intends to turn off the lights associated with the device 110 or the user 5.
[0081] As described above, in some cases, a user may issue a request to control a device process. In this case, the user may speak a request to control a device process to one assistant system 120 (e.g., second system 120b), which is actually started with a command to another assistant system (e.g., first system 120a). In order to properly handle such a request when the second system 120b does not have information about the process associated with the first system 120a, the NLU 460 may access the state data 194, which allows the system 120 to determine that the request corresponds to a device control that can be performed by the device 110 through the interface with the device skill 191.
[0082] The NLU component 460 may process the NLU result data 885 / 825 (hereinafter referred to as Figure 8 Further discussed and may include tagged text data, intent indications, etc.) is returned to orchestrator 430. Orchestrator 430 may forward the NLU result data to skill component 490. If the NLU result data includes a single NLU hypothesis, NLU component 460 and orchestrator component 430 may direct the NLU result data to the skill component 490 associated with the NLU hypothesis. If the NLU result data 885 / 825 includes an N-best list of NLU hypotheses, NLU component 460 and orchestrator component 430 may direct the highest scoring NLU hypothesis to the skill component 490 associated with the highest scoring NLU hypothesis. The system may also include an NLU post-ranker 465, which may rank the potential interpretations determined by NLU component 460 in conjunction with other information. Local device 110 may also include its own NLU post-ranker, which may operate in a manner similar to NLU post-ranker 465. Figure 7 and Figure 8 The NLU component 460, NLU post-ranser 465, and other components are described in more detail.
[0083] Skill components can be software that runs on system 120, which is similar to a software application. That is, skill components 490 can enable system 120 to perform specific functions in order to provide data or generate output of some other request. As used herein, "skill components" can refer to software that can be placed on a machine or virtual machine (e.g., software that can be started in a virtual instance when called). Skill components can be customized software for performing one or more actions indicated by a business entity, device manufacturer, user, etc. The skill components described herein can be referred to using many different terms (such as actions, robots, applications, etc.). System 120 can be configured with more than one skill component 490. For example, a weather service skill component can enable system 120 to provide weather information, a car service skill component can enable system 120 to arrange a trip with respect to a taxi or ride-sharing service, and a restaurant skill component can enable system 120 to order pizza with respect to an online ordering system for a restaurant, etc. Skill components 490 can operate in coordination between system 120 and other devices (such as device 110) to complete certain functions. The input to the skill component 490 may come from a speech processing interaction or through other interaction or input sources. The skill component 490 may include hardware, software, firmware, etc. that may be dedicated to a particular skill component 490 or shared between different skill components 490.
[0084] The skill processing component 125 may communicate with the skill component 490 within the system 120 and / or directly with the orchestrator component 430 or other components. The skill processing component 125 may be configured to perform one or more actions. The ability to perform such actions may sometimes be referred to as a "skill". That is, a skill may enable the skill processing component 125 to perform a specific function in order to provide data or perform some other action requested by a user. For example, a weather service skill may enable the skill processing component 125 to provide weather information to the system 120, a car service skill may enable the skill processing component 125 to schedule a trip with a taxi or ride-sharing service, an ordering pizza skill may enable the skill processing component 125 to order a pizza with a restaurant online ordering system, and the like. Additional types of skills include home automation skills (e.g., skills that enable a user to control home devices such as lights, door locks, cameras, thermostats, etc.), entertainment device skills (e.g., skills that enable a user to control entertainment devices such as smart TVs), video skills, newsletter skills, and custom skills that are not associated with any preconfigured type of skill.
[0085] The system 120 may be configured with a skill component 490 that is specifically designed to interact with the skill processing component 125. Unless otherwise explicitly stated, references to a skill, skill device, or skill component may include the skill component 490 operated by the system 120 and / or the skill operated by the skill processing component 125. In addition, the functions or skills described herein as skills may be referred to using many different terms, such as actions, robots, applications, etc. The skill 490 and or skill processing component 125 may return output data to the orchestrator 430.
[0086] Dialog processing is a field of computer science that deals with exchanges between computing systems and humans via text, audio, and / or other forms of communication. While some dialog processing involves simply generating a response based on the user's most recent input (i.e., a single-turn dialog), more complex dialog processing involves determining and taking action as appropriate on one or more goals expressed by the user in a multi-turn dialog, such as making a restaurant reservation and / or booking an airline ticket. These multi-turn "goal-directed" dialog systems can recognize, retain, and use information gathered during more than one input during a back-and-forth or "multi-turn" interaction with a user; for example, information about the language in which the dialog was conducted.
[0087] The system 100 may include a dialog manager component 572 that manages and / or tracks dialogs between a user and a device and, in some cases, between a user and one or more systems 120. As used herein, a "dialog" may refer to data transmissions (such as involving multiple user inputs and system 100 outputs) between the system 100 and a user (e.g., via the device 110) that are all associated with a single "dialog" between the system and the user, which may have been initiated with a single user input that started the dialog. Thus, data transmissions for a dialog may be associated with the same dialog identifier, which components throughout the system 100 may use to track information for the entire dialog. Subsequent user inputs to the same dialog may or may not begin with saying a wake word. Each natural language input to a dialog may be associated with a different natural language input identifier, such that multiple natural language input identifiers may be associated with a single dialog identifier. In addition, other non-natural language inputs (e.g., image data, gestures, button presses, etc.) may be associated with a particular dialog, depending on the context of the input. For example, a user may open a dialog box with system 100 to verbally request a meal delivery, and the system may respond by displaying images of food available for order, and the user may speak a response (e.g., "item 1" or "that") or may respond by gesture (e.g., pointing to an item on the screen or giving a thumbs up) or may touch the desired item on the screen to select. Non-voice input (e.g., gestures, screen touches, etc.) may be part of the dialog, and data associated therewith may be associated with the dialog identifier for the dialog.
[0088] After identifying that the user is having a conversation with the user, the dialogue manager component 572 may associate the dialogue session identifier with the dialogue. The dialogue manager component 572 may track the user input and the corresponding system-generated response to the user input as a turn. The dialogue session identifier may correspond to multiple rounds of user input and the corresponding system-generated response. The dialogue manager component 572 may directly transmit the data identified by the dialogue session identifier to the arranger component 430 or other components. Depending on the system configuration, the dialogue manager 572 may determine the appropriate system-generated response to give a round of specific speech or user input. Alternatively, the creation of the system-generated response may be managed by another component of the system (e.g., language output component 193, NLG479, arranger 430, etc.), while the dialogue manager 572 selects the appropriate response. Alternatively, another component of the system 120 may select the response using the technology discussed herein. The text of the system-generated response may be sent to the TTS component 480 for creating audio data corresponding to the response. Then, the audio data may be sent to the user device (e.g., device 110) to be ultimately output to the user. Alternatively (or additionally), the dialog responses may be returned in text or some other form.
[0089] The dialogue manager 572 may receive an ASR hypothesis (i.e., text data) and semantically interpret the phrases or sentences represented therein. That is, the dialogue manager 572 determines one or more meanings associated with the phrases or sentences represented in the text data based on the words represented in the text data. The dialogue manager 572 determines a target corresponding to an action that the user wishes to perform and a text data segment that allows the device (e.g., device 110, system 120, skill 490, skill processing component 125, etc.) to perform the intent. For example, if the text data corresponds to "what is the weather like", the dialogue manager 572 may determine that the system 120 is to output weather information associated with the geographic location of the device 110. In another example, if the text data corresponds to "turn off the lights", the dialogue manager 572 may determine that the system 120 is to turn off the lights associated with the device 110 or the user 5.
[0090] Dialog manager 572 may send the result data to one or more skills 490. If the result data includes a single hypothesis, orchestrator component 430 may send the result data to the skill 490 associated with the hypothesis. If the result data includes an N-best list of hypotheses, orchestrator component 430 may send the highest scoring hypothesis to the skill 490 associated with the highest scoring hypothesis.
[0091] The system 120 includes a language output component 193. The language output component 193 includes a natural language generation (NLG) component 479 and a text-to-speech (TTS) component 480. The NLG component 479 can generate text to achieve TTS output to the user. For example, the NLG component 479 can generate text corresponding to instructions corresponding to a specific action to be performed by the user. The NLG component 479 can generate appropriate text for various outputs as described herein. The NLG component 479 may include one or more trained models that are configured to output text suitable for specific inputs. The text output by the NLG component 479 can become an input to the TTS component 480 (e.g., the output text data 1010 discussed below). Alternatively or in addition, the TTS component 480 can receive text data from the skills 490 or other system components for output.
[0092] The NLG component 479 may include a trained model. The NLG component 479 generates text data 1010 based on the dialogue data received by the dialogue manager 572, so that the output text data 1010 has a natural feel and includes words and / or phrases specifically formatted for the requesting individual in some embodiments. NLG can use templates to formulate responses. And / or the NLG system may include models trained according to various templates for forming the output text data 1010. For example, the NLG system may analyze transcripts of local news programs, television programs, sports events, or any other media programs to obtain common components of the relevant language and / or region. As an illustrative example, the NLG system may analyze transcriptions of regional sports programs to determine common words or phrases used to describe scores or other sports news in a particular region. NLG may also receive a dialogue history, a formality indication, and / or a command history or other user history (such as a dialogue history) as input.
[0093] The NLG system may generate dialog data based on one or more response templates. Further continuing the example above, the NLG system may select a template to answer the question "What is the weather like now?" in the form of: "The current weather is $weather_information$." The NLG system may analyze the logical form of the template to generate one or more text responses (including tags and annotations) to familiarize the generated responses. In some embodiments, the NLG system may determine which response is the most appropriate response to select. Thus, the selection may be based on past reactions, past questions, formality, and / or any other feature, or any other combination thereof. Then, the text-to-speech component 480 may be used to generate response audio data representing the response generated by the NLG system.
[0094] TTS components 480 can use one or more different methods to generate audio data (e.g., synthesized speech) according to text data. The text data input to TTS components 480 may come from skill components 490, arranger components 430 or another component of the system. In a synthesis method known as unit selection, TTS components 480 matches text data with a database of recorded speech. TTS components 480 selects the matching unit of the recorded speech and links the units together to form audio data. In another synthesis method known as parameter synthesis, TTS components 480 changes parameters (such as frequency, volume and noise) to create audio data including artificial speech waveforms. Parameter synthesis uses a computerized speech generator, sometimes referred to as a vocoder. TTS components 480 may be able to generate output audio representing natural language speech of one or more natural languages (e.g., English, Mandarin, French, etc.).
[0095] System 100 (on device 110, system 120, or a combination thereof) may include a profile store for storing various information related to individual users, groups of users, devices, etc., that interact with the system. As used herein, a "profile" refers to a set of data associated with a user, group of users, device, etc. The data for a profile may include preferences specific to a user, device, etc.; input and output capabilities of the device; Internet connection information; user bibliographic information; subscription information, and other information.
[0096] The profile storage 470 may include one or more user profiles, each of which is associated with a different user identifier / user profile identifier. Each user profile may include various user identification data. Each user profile may also include data corresponding to user preferences. Each user profile may also include user preferences and / or one or more device identifiers representing one or more devices of the user. For example, a user account may include one or more IP addresses, MAC addresses, and / or device identifiers (such as serial numbers) for each additional electronic device associated with the identified user account. When a user logs into an application installed on the device 110, the user profile (associated with the presented login information) may be updated to include information about the device 110, such as an indication that the device is currently in use. Each user profile may include identifiers of skills that the user has enabled. When a user enables a skill, the user may grant the system 120 permission to allow the skill to perform user input according to the user's natural language. If the user does not enable the skill, the system 120 may not call the skill to perform user input according to the user's natural language.
[0097] Profile storage 470 may include one or more group profiles. Each group profile may be associated with a different group identifier. A group profile may be specific to a group of users. That is, a group profile may be associated with two or more individual user profiles. For example, a group profile may be a family profile associated with a user profile associated with multiple users of a single family. A group profile may include preferences shared by all user profiles associated therewith. Each user profile associated with a group profile may additionally include preferences specific to the user associated therewith. That is, each user profile may include preferences unique to one or more other user profiles associated with the same group profile. A user profile may be an independent profile, or may be associated with a group profile.
[0098] The profile storage 470 may include one or more device profiles. Each device profile may be associated with a different device identifier. Each device profile may include various device identification information. Each device profile may also include one or more user identifiers representing one or more users associated with the device. For example, a profile for a family device may include user identifiers for family users.
[0099] The profile store 470 may include data corresponding to the state data 194. For example, the profile store 470 may indicate the device process control capabilities of one or more devices 110 associated with a particular user profile. Such state data 194 may be updated by one or more devices 110 as the user interacts with the device to maintain an updated record of the device state. Alternatively (or in addition), for a particular user profile, the profile store 470 may include state data 194 reflecting capability data indicating device process control operations that may be performed by the device 110.
[0100] although Figure 4 The components may be shown as part of system 120, device 110, or others, but the components may also be arranged in other devices (e.g., if shown in system 120, then arranged in device 110, vice versa, or arranged in other devices together) without departing from the present disclosure. Figure 5 An apparatus 110 so configured is shown.
[0101] exist Figure 5 In the example system 100 shown, the device 110 includes a first assistant component 140a and a second assistant component 140b. The first assistant component 140a can communicate with the backend components of the first system 120a (for example, via the network 199). The first assistant component 140a can also communicate with the language processing component 592, the language output component 593, the first wake-up word detector 121a and / or the hybrid selector 524. The first system 120a can be associated with one or more local skill components 190a1, 190a2 and 190a3 (collectively referred to as "skill components 190"). The local skill component 190 can communicate with one or more skill processing components 125. The second assistant component 140b can be associated with the second system 120b, which can be a separate computing system that is separate and remote from the device 110. The first system 120a and the second system 120b can be configured as described herein; for example, as described with respect to Figure 1 and Figure 4 described.
[0102] The second assistant component 140b may be logically or otherwise isolated from certain components of the device 110. For example, the second assistant component 140b may not be able to communicate directly with the first assistant component 140a; such communications may need to be mediated by the multi-assistant component 115. The second assistant component 140b may include or be associated with its own proprietary components. For example, the second assistant component 140b may be associated with the second wake-up word detector 121b. In addition, the second assistant component 140b may utilize a separate language processing component and a language output component, which may reside in the device 110 or the second system 120b. However, the second assistant component 140b may be coupled with the multi-assistant component 115 and / or the dialogue manager 472, which may be between the first assistant component 140a and the second assistant component 140b.
[0103] In some implementations, speech processing of input audio data directed to the first system 120a may be performed on the device 110. The device 110 may send a message represented in the input audio data to the second system 120b without first sending the input audio data to the first system 120a. For example, the device 110 may receive the input audio data and detect a wake-up word corresponding to the first system 120a using the first wake-up word detection component 121a. The language processing component 592 of the device 110 may process the input audio data and determine that the input audio data represents a request to generate a message and send the message to the second system 120b. The first assistant component 140a may receive the output of the language processing component 592 and forward the output to the multi-assistant component 115. The first assistant component 140a may include output metadata indicating that the multi-assistant component 115 is to forward the output to the second system 120b (e.g., via the second assistant component 140b). In some cases, the first assistant component 140a may send the output to the language output component 593 to generate an output in the form of output audio data representing the output (e.g., TTS output). The multi-assistant component 593 may receive the output (or output audio data) and metadata and determine that the output is to be processed by the second system 120b. The multi-assistant component 115 may send the output to the second assistant component 140b. The second assistant component 140b may send the output to the second system 120b. The second system 120b may process the output by, for example, executing the command represented in the output. The system 120b may return response data to the device 110; for example, by sending the response output audio data to the multi-assistant component 115 for output by the device's speaker.
[0104] In some cases, the multi-assistant component 115 may determine (e.g., based on state data about an active conversation including input audio data) that response data from the second system 120b is to be translated back into the language of the input audio data. The multi-assistant component 115 may send the response data and an indication that the response data needs to be translated to the first system 120a via the first assistant component 140a. For example, the response data may be audio data and / or text data. The first system 120a may return the translated response data. The translated response data may be audio data and / or text data. If the translated response data is text data, the multi-assistant component 115 may send it to the language output component 593 for conversion into synthesized speech for output by the device 110.
[0105] In at least some embodiments, system 120 can receive audio data 411 from device 110 to recognize speech corresponding to verbal input in the received audio data 411 and perform functions in response to the recognized speech. In at least some embodiments, these functions involve sending instructions (e.g., commands) from system 120 to device 110 (and / or other devices 110) to cause device 110 to perform an action, such as outputting an audible response to the verbal input through a speaker, and / or controlling secondary devices in the environment by sending control commands to the secondary devices.
[0106] Thus, when device 110 is capable of communicating with system 120 via network 199, some or all of the functions capable of being performed by system 120 may be performed by sending one or more instructions to device 110 via network 199, which in turn may process the instructions and perform one or more corresponding actions. For example, system 120 may use remote instructions (e.g., remote responses) included in the response data to direct device 110 to output an audible response to a user question via a speaker of device 110 (or otherwise associated therewith) (e.g., using TTS processing performed by on-device TTS component 580), output content (e.g., music) via a speaker of device 110 (or otherwise associated therewith), display content on a display of device 110 (or otherwise associated therewith), and / or send instructions to a secondary device (e.g., instructions to turn on a smart light). It should be appreciated that system 120 may be configured to provide other functionality in addition to the functionality discussed herein, such as, but not limited to, providing step-by-step instructions for navigating from a departure location to a destination location, conducting e-commerce transactions on behalf of user 5 as part of a shopping function, establishing a communication session (e.g., a video call) between user 5 and another user, and the like.
[0107] The device 110 may include one or more wake-up word detection components 121 (and / or 121a and / or 121b) configured to compare the audio data 411 with a stored model for detecting a wake-up word (e.g., "Alexa"), which indicates to the device 110 that the audio data 411 is to be processed to determine NLU output data (e.g., slot data corresponding to a named entity, tag data, and / or intent data, etc.). In at least some embodiments, the hybrid selector 524 of the device 110 may send the audio data 411 to the wake-up word detection component 121a. If the wake-up word detection component 121a detects a wake-up word in the audio data 411, the wake-up word detection component 121a may send an indication of such detection to the hybrid selector 524. In response to receiving the indication, the hybrid selector 524 may send the audio data 411 to the system 120 and / or the ASR component 550. The wake-up word detection component 121a may also send an indication to the hybrid selector 524 that the wake-up word was not detected. In response to receiving such an indication, hybrid selector 524 may refrain from sending audio data 411 to system 120 and may prevent ASR component 550 from further processing audio data 411. In this case, audio data 411 may be discarded.
[0108] Device 110 may use an on-device language processing component such as SLU / language processing component 592 (which may include ASR component 550 and NLU 560) to perform its own speech processing in a manner similar to that discussed herein with respect to SLU component 192 (or ASR component 450 and NLU component 460) of system 120. Language processing component 592 may operate similarly to language processing component 192, ASR component 550 may operate similarly to ASR component 450, and NLU component 560 may operate similarly to NLU component 460. Device 110 may also internally include or otherwise have access to other components such as one or more skill components 190 (which may operate similarly to skill component 490) capable of executing commands based on NLU output data or other results determined by device 110 / system 120, profile store 570 (configured to store profile data similar to the profile data discussed herein with respect to profile store 470 of system 120), or other components. In at least some embodiments, profile storage 570 may store profile data only for users or groups of users specifically associated with device 110. Skills component 190 may communicate with skills processing component 125, similar to that described above with respect to skills component 490. Device 110 may also have its own language output component 593, which may include NLG component 579 and TTS component 580. Language output component 593 may operate similarly to language processing component 192, NLG component 579 may operate similarly to NLG component 479, and TTS component 580 may operate similarly to TTS component 480.
[0109] In at least some embodiments, the on-device language processing component may not have the same capabilities as the language processing component of system 120. For example, the on-device language processing component may be configured to process only a subset of the natural language user input that can be processed by system 120. For example, this subset of natural language user input may correspond to local-type natural language user input, such as user input that controls a device or component associated with the user's home. In such a case, the on-device language processing component may be able to interpret and respond to the local-type natural language user input more quickly, for example, than processing involving system 120. If device 110 attempts to process natural language user input and the on-device language processing component is not necessarily the best fit for the input, the language processing result determined by device 110 may indicate a low confidence or other metric indicating that the processing performed by device 110 may not be as accurate as the processing performed by system 120.
[0110] The hybrid selector 524 of the device 110 may include a hybrid proxy (HP) 526 configured to proxy traffic to / from the system 120. For example, the HP 526 may be configured to send messages to / from a hybrid execution controller (HEC) 527 of the hybrid selector 524. For example, command / direction data received from the system 120 may be sent to the HEC 527 using the HP 526. The HP 526 may also be configured to allow the audio data 411 to pass to the system 120 while also receiving (e.g., intercepting) the entire audio data 411 and sending the audio data 411 to the HEC 527.
[0111] In at least some embodiments, the hybrid selector 524 may also include a local request orchestrator (LRO) 528, which is configured to notify the ASR component 550 about the availability of new audio data 411 representing the user's voice and otherwise initiate operation of the local language processing when the new audio data 411 becomes available. In general, the hybrid selector 524 may control the execution of the local language processing, such as by sending "execute" and "terminate" events / instructions. An "execute" event may direct the component to continue any suspended execution (e.g., by directing the component to execute according to a previously determined intent in order to determine the guidance). Meanwhile, a "terminate" event may direct the component to terminate further execution, such as when the device 110 receives guidance data from the system 120 and chooses to use the remotely determined guidance data.
[0112] Thus, when receiving the audio data 411, the HP 526 may allow the audio data 411 to be passed to the system 120, and the HP 526 may also input the audio data 411 to the on-device ASR component 550 by routing the audio data 411 to the HEC 527 of the hybrid selector 524, whereby the LRO 528 notifies the ASR component 550 of the audio data 411. At this point, the hybrid selector 524 may wait for response data from either or both of the system 120 or the local language processing component. However, the present disclosure is not limited thereto, and in some instances, the hybrid selector 524 may send the audio data 411 only to the local ASR component 550 without departing from the present disclosure. For example, the device 110 may process the audio data 411 locally without sending the audio data 411 to the system 120.
[0113] The local ASR component 550 is configured to receive the audio data 411 from the hybrid selector 524 and recognize the speech in the audio data 411, and the local NLU component 560 is configured to determine the user intent based on the recognized speech and determine how to act based on the user intent by generating NLU output data that may include guidance data (e.g., instructing the component to perform an action). Such NLU output data may take a form similar to that determined by the NLU component 460 of the system 120. In some cases, the guidance may include a description of the intent (e.g., the intent to turn off {device A}). In some cases, the instructions may include (e.g., encode) an identifier of the second device (such as a kitchen light) and an operation to be performed on the second device. The guidance data may be formatted using Java (such as JavaScript syntax or a JavaScript-based syntax). This may include formatting the guidance using JSON. In at least some embodiments, the device-determined guidance may be serialized, just as the remotely determined guidance may be serialized for transmission in a data packet over the network 199. In at least some embodiments, the device-determined guidance may be formatted as a program application programming interface (API) call having the same logical operation as the remotely determined guidance. In other words, the device-determined directions may mimic the remotely-determined directions by using the same or similar format as the remotely-determined directions.
[0114] An NLU hypothesis (output by NLU component 560) may be selected for use in response to the natural language user input, and local response data (e.g., local NLU output data, local knowledge base information, Internet search results, and / or local guidance data) may be sent to hybrid selector 524, such as a "ReadyToExecute" response. Hybrid selector 524 may then determine whether to respond to the natural language user input using guidance data from an on-device component, whether to use guidance data received from system 120, assuming even a remote response is received (e.g., when device 110 is able to access system 120 via network 199), or determine output audio requesting additional information from user 5.
[0115] Device 110 and / or system 120 may associate a unique identifier with each natural language user input. Device 110 may include the unique identifier when sending audio data 411 to system 120, and response data from system 120 may include the unique identifier to identify which natural language user input the response data corresponds to.
[0116] In at least some embodiments, device 110 may include or be configured to use one or more skill components 190, which may operate in a manner similar to skill components 490 implemented by system 120. Skill components 190 may correspond to one or more domains that are used to determine how to act on verbal input in a particular manner, such as by outputting instructions that correspond to a determined intent and that can be processed to achieve the desired action. Skill components 190 installed on device 110 may include, but are not limited to, smart home skill components (or smart home domains) and / or device control skill components (or device control domains) that are executed in response to verbal input corresponding to an intent to control a second device in an environment, music skill components (or music domains) that are executed in response to verbal input corresponding to an intent to play music, navigation skill components (or navigation domains) that are executed in response to verbal input corresponding to an intent to get directions, shopping skill components (or shopping domains) that are executed in response to verbal input corresponding to an intent to purchase items from an electronic market, and the like.
[0117] Additionally or alternatively, the device 110 may communicate with one or more skill processing components 125. For example, the skill processing components 125 may be located in a remote environment (e.g., a separate location) such that the device 110 can communicate only with the skill processing components 125 via the network 199. However, the present disclosure is not limited in this regard. For example, in at least some embodiments, the skill processing components 125 may be configured in a local environment (e.g., a home server and / or the like) such that the device 110 can communicate with the skill processing components 125 via a dedicated network such as a local area network (LAN).
[0118] As used herein, a "skill" may refer to a skill component 190 / 490, a skill processing component 125, or a combination of a skill component 190 / 490 and a corresponding skill processing component 125. Similar to the approach discussed herein, the local device 110 may be configured to recognize multiple different wake-up words and / or perform different categories of tasks based on the wake-up words. These different wake-up words may invoke different processing components ( Figure 5 ). For example, detection of the wake word "Alexa" by the wake word detector 121a may cause the audio data to be sent to certain language processing components 592 / skills 190 for processing, while detection of the wake word "computer" by the wake word detector may cause the audio data to be sent to different language processing components 592 / skills 190 for processing.
[0119] Figure 6 6 is a conceptual diagram of an ASR component 450 according to an embodiment of the present disclosure. The ASR component 450 may receive audio data 631 and process it to recognize and transcribe the speech contained therein. The ASR component 450 may output a transcript as ASR output data 615. In some cases, the ASR component 450 may generate more than one ASR hypothesis (e.g., representing possible transcripts) for a single spoken natural language input. The ASR hypothesis may be assigned a score (e.g., a probability score, a confidence score, etc.) that represents the likelihood that the corresponding ASR hypothesis matches the spoken natural language input (e.g., representing the likelihood that a particular set of words matches the words spoken in the natural language input). The score may be based on many factors, including, for example, the similarity of the sounds in the spoken natural language input to a language sound model (e.g., an acoustic model 653 stored in an ASR model storage 652), and the likelihood that a particular word matching the sounds will be included in a sentence at a particular position (e.g., using a language or grammar model 654). Based on the considered factors and assigned confidence scores, ASR component 450 may output the ASR hypothesis that most likely matches the spoken natural language input, or may output multiple ASR hypotheses in the form of a grid or N-best list, where each ASR hypothesis corresponds to a respective score.
[0120] The ASR component 450 may interpret the spoken natural language input using one or more models in the ASR model storage 652. These models may consist of NN-based end-to-end models such as the ASR model 650. Some models may process the audio data 631 based on similarities between the spoken natural language input and acoustic units (e.g., representing subword units or phonemes) in the acoustic model 653, and use a language model 654 to predict the word / phrase / sentence that a sequence of acoustic units may represent. In some implementations, a finite state transducer (FST) 655 may perform language model functions.
[0121] The ASR component 450 may include a speech recognition engine 658. The ASR component 450 may receive audio data 631 from, for example, a microphone 114 of a user device 110. In some cases, the audio data 631 may be processed audio that has been detected by an acoustic front end (AFE) or other component. The speech recognition engine 658 may process the audio data 631 using one or more of an ASR model 650, an acoustic model 653, a language model 654, an FST 655, and / or other data models and information to recognize the speech conveyed in the audio data. The audio data 631 may be a frame that has been digitized (e.g., by an AFE) to represent a time interval, and the AFE determines a plurality of values (referred to as features) representing the quality of the audio data for the frame, and a set of values (referred to as feature vectors) representing the features / quality of the audio data within the frame. In at least some embodiments, the audio frame may be 10 ms per frame. In some embodiments, the audio frame may represent a larger audio window; for example, about 2 ms. Many different features may be determined, as is known in the art, and each feature may represent a certain quality of the audio that may be useful for ASR processing. The AFE may use a variety of methods to process the audio data, such as log filter bank energy (LFBE), Mel-frequency cepstral coefficients (MFCC), perceptual linear prediction (PLP) techniques, neural network feature vector techniques, linear discriminant analysis, semi-knot covariance matrices, or other methods known to those skilled in the art. In some cases, the feature vector of the audio data 631 may be encoded at the processing system 120, in which case the feature vector may be decoded by the speech recognition engine 658 and / or decoded prior to processing by the speech recognition engine 658.
[0122] In some embodiments, the ASR component 450 may process the audio data 631 using an ASR model 650. The ASR model 650 may be, for example, a recurrent neural network, such as an RNN-T. The ASR model 650 may predict the probability (y|x) of a label y=(1,...,yu) given an acoustic feature x=(1,...,t). During reasoning, the ASR model 650 may generate an N-best list using, for example, a beam search decoding algorithm. The ASR model 650 may include an encoder 610, a prediction network 620, a joint network 630, and a softmax 640. The encoder 610 may be similar or analogous to an acoustic model (e.g., similar to an acoustic model 653 described below), and may process a series of acoustic input features to generate an encoded hidden representation. The prediction network 620 may be similar or analogous to a language model (e.g., similar to a language model 654 described below), and may process previous output label predictions and map them to corresponding hidden representations. The joint network 630 may be, for example, a feed-forward NN that processes the hidden representations from the encoder 610 and the prediction network 620 and predicts output label probabilities. The softmax 640 may be a function implemented to normalize the predicted output probabilities (e.g., as a layer of the joint network 630).
[0123] In some implementations, the speech recognition engine 658 may attempt to match the feature vectors received in the audio data 631 with known acoustic units (e.g., phonemes) and words of the language in the stored acoustic model 653, language model 654, and / or FST 655. For example, the audio data 631 may be processed by one or more acoustic models 653 to determine acoustic unit data. The acoustic unit data may include an indicator of a sound unit detected in the audio data 631 by the ASR component 450. For example, an acoustic unit may be composed of one or more of a phoneme, a diphoneme, a tone, a phoneme, a diphoneme, a triphoneme, etc. The acoustic unit data may be represented using one or a series of symbols of a phonetic symbol (such as X-SAMPA, the International Phonetic Alphabet, or the Primary Instructional Alphabet (ITA) phonetic symbol). In some embodiments, the phonemic representation of the audio data may be analyzed using an n-gram-based tokenization procedure. An entity or a slot representing one or more entities may be represented by a series of n-grams.
[0124] The acoustic unit data may be processed using the language model 654 (and / or using the FST 655) to determine the ASR output data 615. The ASR output data 615 may include one or more hypotheses. One or more of the hypotheses represented in the ASR output data 615 may then be sent to other components (such as the NLU component 460 / 560) for further processing, as described herein. The ASR output data 615 may include representations of the utterance text, such as words, subword units, etc.
[0125] The speech recognition engine 658 calculates the score of the feature vector based on the acoustic information and the language information. Acoustic information (such as the identifier of the acoustic unit and / or the corresponding score) is used to calculate the acoustic score, which represents the probability that the expected sound represented by a set of feature vectors matches the language phoneme. The language information is used to adjust the acoustic score by considering the sounds and / or words used in context with each other, thereby increasing the probability that the ASR component 450 will output a grammatically meaningful ASR hypothesis. The specific model used may be a general model, or it may be a model corresponding to a specific domain (such as music, banking, etc.).
[0126] The speech recognition engine 658 can use a variety of techniques to match feature vectors with phonemes, such as using a hidden Markov model (HMM) to determine the probability that a feature vector may match a phoneme. The received sound can be represented as a path between HMM states, and multiple paths can represent multiple possible text matches for the same sound. Other techniques can also be used, such as using FST.
[0127] The speech recognition engine 658 can use the acoustic model 653 to attempt to match the received audio feature vector with a word or sub-word acoustic unit. The acoustic unit can be a phoneme, a phoneme, a phoneme in context, a syllable, a part of a syllable, a syllable in context, or any other such part of a word. The speech recognition engine 658 calculates the recognition score of the feature vector based on acoustic information and language information. The acoustic information is used to calculate the acoustic score, which represents the possibility of matching the expected sound represented by a set of feature vectors with the sub-word unit. The language information is used to adjust the acoustic score by considering the sounds and / or words used mutually in the context, thereby improving the possibility of the ASR component 450 outputting a grammatically meaningful ASR hypothesis.
[0128] The speech recognition engine 658 can use a variety of techniques to match feature vectors with phonemes or other acoustic units (such as diphones, triphones, etc.). A common technique is to use a hidden Markov model (HMM). HMM is used to determine the probability that a feature vector may match a phoneme. HMM is used to present many states, wherein the states collectively represent potential phonemes (or other acoustic units, such as triphones), and each state is associated with a model (such as a Gaussian mixture model or a deep belief network). The transition between states may also have an associated probability, which represents the possibility of reaching the current state from the previous state. The received sound can be represented as a path between HMM states, and multiple paths can represent multiple possible text matches of the same sound. Each phoneme can be represented by multiple potential states, which correspond to different known pronunciations of the phoneme and its parts (such as the beginning, middle and end of the spoken language sound). The preliminary determination of the probability of a potential phoneme can be associated with a state. When a new feature vector is processed by the speech recognition engine 658, the state can change or remain unchanged based on the processing of the new feature vector. The Viterbi algorithm can be used to find the most likely sequence of states based on the processed feature vectors.
[0129] Possible phonemes and associated states / state transitions (e.g., HMM states) can form a path that traverses a grid of possible phonemes. Each path represents a series of phonemes that may be matched to audio data represented by a feature vector. Depending on the recognition score calculated for each phoneme, a path may overlap with one or more other paths. Certain probabilities are associated with each transition between states. A cumulative path score for each path can also be calculated. This process of determining scores based on feature vectors can be referred to as acoustic modeling. When scores are combined as part of ASR processing, the scores can be multiplied (or otherwise combined) to achieve the desired combined score, or the probabilities can be converted to the log domain and added to assist in processing.
[0130] The speech recognition engine 658 can also calculate the score of the path branch based on the language model or grammar. Language modeling involves determining which words may be used together to form the score of coherent words and sentences. The application language model can improve the possibility of the ASR component 450 correctly interpreting the voice contained in the audio data. For example, for the input audio that sounds like "hello", the acoustic model processing of the potential phoneme path of "HELO", "HALO" and "YELO" can be adjusted by the language model to adjust the recognition score of "HELO" (interpreted as the word "hello"), "HALO" (interpreted as the word "halo") and "YELO" (interpreted as the word "yellow") based on the language context of each word in the spoken utterance.
[0131] Figure 7 and Figure 8 It shows how the NLU component 460 can perform NLU processing. Figure 7 is a conceptual diagram of how natural language processing is performed according to an embodiment of the present disclosure. And Figure 8 is a conceptual diagram of how natural language processing is performed according to an embodiment of the present disclosure.
[0132] Figure 7 How to perform NLU processing on text data is described. NLU component 460 can process text data including several ASR hypotheses for a single user input. For example, if ASR component 450 outputs text data including an n-best list of ASR hypotheses, NLU component 460 can process the text data with respect to all (or a portion) of the ASR hypotheses represented in the text data.
[0133] NLU component 460 can annotate text data by parsing and / or marking text data. For example, for the text data "Tell me the weather in Seattle", NLU component 460 can mark "Tell me the weather in Seattle" as the <output weather> intent, and mark "Seattle" separately as the location of weather information.
[0134] NLU component 460 can include candidate list component 750. Candidate list component 750 selects skills executable with respect to ASR output data 615 input to NLU component 460 (e.g., applications executable with respect to user input). ASR output data 615 (also referred to as ASR data 615) can include representations of utterance text, such as words, sub-word units, etc. Thus, candidate list component 750 limits more resource-intensive downstream NLU processes to be executed with respect to skills executable with respect to user input.
[0135] Without the candidate list component 750, the NLU component 460 can process the ASR output data 615 input to the component in parallel, serially, or using some combination thereof for each skill of the system. By implementing the candidate list component 750, the NLU component 460 can process the ASR output data 615 only for skills that can be executed with respect to the user input. This reduces the overall computing power and latency due to NLU processing.
[0136] The candidate list component 750 may include one or more trained models. The models may be trained to recognize various forms of user input that may be received by the system 120. For example, during a training period, the skill processing component 125 associated with a skill may provide the system 120 with training text data representing sample user input that may be provided by a user to invoke the skill. For example, for a ridesharing skill, the skill processing component 125 associated with the ridesharing skill may provide the system 120 with training text data including text corresponding to "Call me a taxi to [location]," "Call me a car to [location]," "Book me a taxi to [location]," "Book me a car to [location]," and the like. The one or more training models to be used by the candidate list component 750 may be trained using the training text data representing sample user input to determine other potentially relevant user input structures that a user may attempt to use to invoke a particular skill. During training, the system 120 may consult the skill processing component 125 associated with the skill as to whether other user input structures determined from the perspective of the skill processing component 125 are allowed to be used to invoke the skill. Alternative user input structures may be derived by one or more trained models during model training and / or may be based on user input structures provided by different skills. The skill processing component 125 associated with a particular skill may also provide the system 120 with training text data indicating grammar and annotations. The system 120 may use training text data representing sample user input, determined relevant user input, grammar and annotations to train a model that indicates when the user input can be directed to a skill / processed by a skill based at least in part on the structure of the user input. Each trained model of the candidate list component 750 may be trained on a different skill. Alternatively, the candidate list component 750 may use one trained model per domain, such as one trained model for skills associated with the weather domain, one trained model for skills associated with the ridesharing domain, and the like.
[0137] The system 120 can use the sample user inputs provided by the skill processing component 125 and the related sample user inputs that may be determined during training as binary examples to train a model associated with a skill associated with the skill processing component 125. The model associated with a particular skill can then be operated by the candidate list component 750 at runtime. For example, some sample user inputs may be positive examples (e.g., user inputs that can be used to invoke a skill). Other sample user inputs may be negative examples (e.g., user inputs that may not be used to invoke a skill).
[0138] As described above, the candidate list component 750 may include a different trained model for each skill of the system, a different trained model for each domain, or some other combination of trained models. For example, the candidate list component 750 may alternatively include a single model. The single model may include a portion trained on features (e.g., semantic features) shared by all skills of the system. The single model may also include skill-specific portions, each of which is trained on a specific skill of the system. Implementing a single model with skill-specific portions may result in less latency than implementing a different trained model for each skill because the single model with skill-specific portions limits the number of features processed at each skill level.
[0139] The sections trained on characteristics shared by more than one skill may be clustered based on domains. For example, a first section of the section trained on multiple skills may be trained on weather domain skills, a second section of the section trained on multiple skills may be trained on music domain skills, a third section of the section trained on multiple skills may be trained on travel domain skills, etc.
[0140] Clustering may not be beneficial in every case because it may cause the candidate list component 750 to output an indication of only a portion of the skills that the ASR output data 615 may relate to. For example, the user input may correspond to "Tell me about Tom Collins." If the model clusters based on domain, the candidate list component 750 may determine that the user input corresponds to a recipe skill (e.g., drink recipes) even though the user input may also correspond to an information skill (e.g., including information about a person named Tom Collins).
[0141] The NLU component 460 can include one or more recognizers 763. In at least some embodiments, the recognizer 763 can be associated with the skill processing component 125 (e.g., the recognizer can be configured to interpret text data to correspond to the skill processing component 125). In at least some other instances, the recognizer 763 can be associated with a domain such as smart home, video, music, weather, customization, etc. (e.g., the recognizer can be configured to interpret text data to correspond to the domain).
[0142] If the candidate list component 750 determines that the ASR output data 615 may be associated with multiple domains, the recognizers 763 associated with the domains may process the ASR output data 615, while recognizers 763 not indicated in the output of the candidate list component 750 may not process the ASR output data 615. The "finalist" recognizers 763 may process the ASR output data 615 in parallel, serially, partially in parallel, etc. For example, if the ASR output data 615 may be associated with both a communications domain and a music domain, the recognizers associated with the communications domain may process the ASR output data 615 in parallel or partially in parallel, while the recognizers associated with the music domain process the ASR output data 615.
[0143] Each recognizer 763 may include a named entity recognition (NER) component 762. The NER component 762 attempts to identify grammatical and lexical information that can be used to interpret meaning about the text data input therein. The NER component 762 identifies portions of the text data that correspond to named entities associated with the domain associated with the recognizer 763 that implements the NER component 762. The NER component 762 (or other components of the NLU component 460) may also determine whether a word refers to an entity whose identity is not explicitly mentioned in the text data, such as "he," "she," "it," or other anaphora, exophoria, etc.
[0144] Each recognizer 763, and more specifically each NER component 762, may be associated with a specific grammar database 776 and a specific set of intents / actions 774, which may be stored in the NLU storage 773, and a specific personalized dictionary 786, which may be stored in the entity library 782. Each gazetteer 784 may include domain / skill index vocabulary information associated with a specific user and / or device 110. For example, gazetteer A (784a) includes skill index vocabulary information 786aa to 786an. For example, a user's music domain vocabulary information may include album name, artist name, and song name, while the user's communication domain vocabulary information may include contact names. Because each user's music collection and contact list are presumably different. This personalized information may improve entity resolution performed later.
[0145] The NER component 762 applies grammatical information 776 and vocabulary information 786 associated with a domain (associated with a recognizer 763 implementing the NER component 762) to determine mentions of one or more entities in the text data. In this way, the NER component 762 identifies "slots" (each slot corresponds to one or more specific words in the text data) that may be useful for later processing. The NER component 762 can also label each slot with a type (e.g., noun, place, city, artist name, song name, etc.).
[0146] Each grammar database 776 includes entity names (i.e., nouns) that are commonly found in speech for a particular domain associated with the grammar database 776, while the vocabulary information 786 is personalized to the user and / or device 110 from which the user input originated. For example, a grammar database 776 associated with the shopping domain may include a database of words that people commonly use when discussing shopping.
[0147] A downstream process called entity resolution (discussed in detail elsewhere herein) links text data slots to specific entities known to the system. To perform entity resolution, the NLU component 460 may utilize gazetteer information (784a through 784n) stored in the entity library storage 782. The gazetteer information 784 may be used to match text data (representing a portion of the user input) to text data representing known entities (such as song titles, contact names, etc.). The gazetteer 784 may be linked to users (e.g., a particular gazetteer may be associated with a particular user's music collection), may be linked to certain domains (e.g., a shopping domain, a music domain, a video domain, etc.), or may be organized in a variety of other ways.
[0148] Each recognizer 763 may also include an intent classification (IC) component 764. The IC component 764 parses the text data to determine the intent that may represent the user input (and the domain association associated with the recognizer 763 that implements the IC component 764). The intent represents the action that the user wishes to perform. The IC component 764 can communicate with a database 774 of words linked to the intent. For example, a music intent database can link words and phrases such as "quiet", "turn off the volume", and "mute" to the <mute> intent. The IC component 764 identifies potential intents by comparing words and phrases in the text data (representing at least a portion of the user input) with words and phrases in the intent database 774 (and the domain association associated with the recognizer 763 that implements the IC component 764).
[0149] Intents recognizable by a particular IC component 764 are linked to a domain-specific (i.e., a domain associated with the recognizer 763 associated with the implementation of the IC component 764) grammatical framework 776, in which there are "slots" to be filled. Each slot of the grammatical framework 776 corresponds to a portion of text data that the system considers to correspond to an entity. For example, the grammatical framework 776 corresponding to the intent of <play music> may correspond to a text data sentence structure such as "play {artist name}", "play {album name}", "play {song name}", "play {song name} by {artist name}", etc. However, in order to make entity resolution more flexible, the grammatical framework 776 may not be structured as sentences, but rather based on associating slots with grammatical tags.
[0150] For example, before identifying named entities in text data, NER component 762 may parse text data based on grammatical rules and / or models to identify words as subjects, objects, verbs, prepositions, etc. IC component 764 (implemented by the same recognizer 763 as NER component 762) may use the recognized verbs to identify intent. Then, NER component 762 may determine a grammatical model 776 associated with the recognized intent. For example, a grammatical model 776 for an intent corresponding to <play music> may specify a slot list suitable for playing the recognized "object" and any object modifiers (e.g., prepositional phrases), such as {artist name}, {album name}, {song name}, etc. Then, NER component 762 may search the corresponding fields in dictionary 786 (and the domain association associated with recognizer 763 implementing NER component 762), thereby attempting to match words and phrases in text data previously marked by NER component 762 as grammatical objects or object modifiers with words and phrases recognized in dictionary 786.
[0151] NER components 762 can perform semantic tagging, i.e., it is marked according to the type / semantic meaning of a word or word combination. NER components 762 can use heuristic grammar rules to parse text data, or models can be constructed using technologies such as hidden Markov models, maximum entropy models, log-linear models, conditional random fields (CRFs). For example, the NER components 762 implemented by the music domain identifier can parse the text data corresponding to "playing the Rolling Stones' mom's little helper" and mark it as {verb}: "play", {object}: "mom's little helper", {object preposition}: "of", and {object modifier}: "Rolling Stones". NER components 762 recognizes "playing" as a verb based on the word database associated with the music domain, and IC components 764 (also implemented by the music domain identifier) can determine that the verb corresponds to <playing music> intention. At this stage, the meaning of "mother's little helper" or "the rolling stones" has not yet been determined, but based on grammatical rules and models, the NER component 762 has determined that the text of these phrases is related to the grammatical objects (ie, entities) of the user input represented in the text data.
[0152] NER component 762 can tag text data to give it meaning. For example, NER component 762 can tag "play Mom's Little Helper by the Rolling Stones" as: {domain} music, {intent} <play music>, {artist name} the Rolling Stones, {media type} songs, and {song name} Mom's Little Helper. For another example, NER component 762 can tag "play song by the Rolling Stones" as: {domain} music, {intent} <play music>, {musician name} the Rolling Stones, and {media type} songs.
[0153] The candidate list component 750 may receive the ASR output data 615 (eg, output from the ASR component 450 or output from the device 110b). Figure 8 4. As shown in FIG. 4 , the ASR component 450 may embed the ASR output data 615 into a form that can be processed by the trained model using sentence embedding techniques known in the art. Sentence embedding causes the ASR output data 615 to include text whose structure enables the trained model of the candidate list component 850 to operate on the ASR output data 615. For example, the embedding of the ASR output data 615 may be a vector representation of the ASR output data 615.
[0154] The candidate list component 750 can make a binary decision (e.g., yes or no) regarding which domains are relevant to the ASR output data 615. The candidate list component 750 can use one or more trained models described above in this document to make such a determination. If the candidate list component 750 implements a single trained model for each domain, the candidate list component 750 can simply run the model associated with the enabled domains, as indicated in the user profile associated with the device 110 and / or user that initiated the user input.
[0155] The candidate list component 750 may generate n-best list data 815 indicating domains that may be executed with respect to the user input in the ASR output data 615. The size of the n-best list represented in the n-best list data 815 is configurable. In one example, the n-best list data 815 may indicate each domain of the system and include an indication for each domain as to whether the domain may be able to execute the user input represented in the ASR output data 615. In another example, the n-best list data 815 may indicate only domains that may be able to execute the user input represented in the ASR output data 615, rather than indicating each domain of the system. In yet another example, the candidate list component 750 may implement thresholding so that the n-best list data 815 may indicate no more than a maximum number of domains that may execute the user input represented in the ASR output data 615. In one example, the threshold number of domains that may be represented in the n-best list data 815 is ten. In another example, the domains included in the n-best list data 815 may be limited by a threshold score, where only domains indicating a likelihood of processing user input above a certain score (determined by processing the ASR output data 615 relative to such domains by the candidate list component 750) are included in the n-best list data 815.
[0156] ASR output data 615 may correspond to more than one ASR hypothesis. When this occurs, candidate list component 750 may output a different n-best list (represented in n-best list data 815) for each ASR hypothesis. Alternatively, candidate list component 750 may output a single n-best list representing domains associated with multiple ASR hypotheses represented in ASR output data 615.
[0157] As described above, candidate list component 750 may implement thresholding so that the n-best list output therefrom may include no more than a threshold number of entries. If ASR output data 615 includes more than one ASR hypothesis, the n-best list output by candidate list component 750 may include no more than a threshold number of entries, regardless of the number of ASR hypotheses output by ASR component 450. Alternatively or in addition, the n-best list output by candidate list component 750 may include no more than a threshold number of entries for each ASR hypothesis (e.g., no more than five entries for the first ASR hypothesis, no more than five entries for the second ASR hypothesis, etc.).
[0158] In addition to making a binary decision about whether a domain is likely to be associated with the ASR output data 615, the candidate list component 750 may also generate a confidence score representing the likelihood that the domain is associated with the ASR output data 615. If the candidate list component 750 implements a different trained model for each domain, the candidate list component 750 may generate a different confidence score for each separate domain trained model that is run. If the candidate list component 750 runs a model for each domain when receiving the ASR output data 615, the candidate list component 750 may generate a different confidence score for each domain of the system. If the candidate list component 750 only runs models for domains that are associated with skills indicated as enabled in a user profile associated with the device 110 and / or user that initiated the user input, the candidate list component 750 may only generate a different confidence score for each domain associated with at least one enabled skill. If the candidate list component 750 implements a single trained model with a domain-specific training portion, the candidate list component 750 may generate a different confidence score for each domain for which the specific training portion is run. The candidate list component 750 can perform matrix-vector modifications to obtain confidence scores for all domains of the system in a single instance of processing the ASR output data 615 .
[0159] The N-best list data 815 including confidence scores that may be output by the candidate list component 750 may be represented, for example, as:
[0160] Search domain, 0.67
[0161] Recipe domain, 0.62
[0162] Information domain, 0.57
[0163] Shopping domain, 0.42
[0164] As shown, the confidence score output by candidate list component 750 can be a numeric value. The confidence score output by candidate list component 750 can alternatively be a ranking value (eg, high, medium, low).
[0165] The best list may include only entries for domains whose confidence scores meet (e.g., equal or exceed) a minimum threshold confidence score. Alternatively, the candidate list component 750 may include entries for all domains associated with a user-enabled skill, even if one or more of the domains are associated with a confidence score that does not meet the minimum threshold confidence score.
[0166] The candidate list component 750 may consider other data 820 when determining which domains may be associated with the user input represented in the ASR output data 615 and the corresponding confidence scores. The other data 820 may include usage history data associated with the device 110 and / or the user that initiated the user input. For example, if the user input initiated by the device 110 and / or the user frequently calls a domain, the confidence score of the domain may be improved. Conversely, if the user input initiated by the device 110 and / or the user rarely calls a domain, the confidence score of the domain may be reduced. Therefore, the other data 820 may include, for example, an indicator of a user associated with the ASR output data 615 determined by the user recognition component.
[0167] The other data 820 may be embedded into the characters prior to input into the candidate list component 750. Alternatively, the other data 820 may be embedded prior to input into the candidate list component 750 using other techniques known in the art.
[0168] Other data 820 may also include data indicating domains associated with the device 110 that initiated the user input and / or the user-enabled skill. The candidate list component 750 may use such data to determine which specific domains' trained models to run. That is, the candidate list component 750 may determine to run only trained models associated with domains associated with the user-enabled skill. The candidate list component 750 may alternatively use such data to change the confidence score of the domain.
[0169] By way of example, considering two domains, a first domain associated with at least one enabled skill, and a second domain not associated with any user-enabled skill of the user initiating the user input, the candidate list component 750 may run a first model specific to the first domain and a second model specific to the second domain. Alternatively, the candidate list component 750 may run a model configured to determine a score for each of the first and second domains. The candidate list component 750 may determine the same confidence score for each of the first and second domains in the first instance. The candidate list component 750 may then change those confidence scores based on which domains are associated with at least one skill enabled by the current user. For example, the candidate list component 750 may increase the confidence score associated with the domain associated with at least one enabled skill, while keeping the confidence scores associated with other domains unchanged. Alternatively, the candidate list component 750 may keep the confidence score associated with the domain associated with at least one enabled skill unchanged, while reducing the confidence score associated with another domain. Additionally, candidate list component 750 can increase a confidence score associated with a domain associated with at least one enabled skill and decrease a confidence score associated with another domain.
[0170] As shown, the user profile can indicate which skills the corresponding user has enabled (e.g., authorized to be performed using data associated with the user). Such indications can be stored in profile storage 470. When candidate list component 750 receives ASR output data 615, candidate list component 750 can determine whether profile data associated with the user and / or device 110 initiating the command includes an indication of enabled skills.
[0171] Other data 820 may also include data indicating the type of device 110. The device type may indicate the output capabilities of the device. For example, the device type may correspond to a device with a visual display, a headless (e.g., no display) device, whether the device is mobile or fixed, whether the device includes audio playback capabilities, whether the device includes a camera, other device hardware configurations, etc. The candidate list component 750 may use such data to determine which domain-specific trained models to run. For example, if the device 110 corresponds to a no-display type device, the candidate list component 750 may determine not to run a trained model specific to the domain of the output video data. The candidate list component 750 may alternatively use such data to change the confidence score of the domain.
[0172] By way of example, considering two domains, one domain outputting audio data and the other domain outputting video data, the candidate list component 750 may run a first model specific to the domain generating the audio data and a second model specific to the domain generating the video data. Alternatively, the candidate list component 750 may run a model configured to determine the score for each domain. The candidate list component 750 may determine the same confidence score for each domain in the first instance. Then, the candidate list component 750 may change the original confidence score based on the type of device 110 that initiated the user input corresponding to the ASR output data 615. For example, if the device 110 is a non-display device, the candidate list component 750 may increase the confidence score associated with the domain generating the audio data while keeping the confidence score associated with the domain generating the video data unchanged. Alternatively, if the device 110 is a non-display device, the candidate list component 750 may keep the confidence score associated with the domain generating the audio data unchanged while reducing the confidence score associated with the domain generating the video data. Additionally, if device 110 is a non-display device, candidate list component 750 can increase the confidence scores associated with domains that generate audio data and decrease the confidence scores associated with domains that generate video data.
[0173] The device type information represented in the other data 820 may represent the output capabilities of a device for outputting content to a user, which is not necessarily a user input initiating device. For example, a user may input a verbal user input corresponding to "play Game of Thrones" to a device that does not include a display. The system may determine that a smart TV or other display device (associated with the same user profile) is used to output Game of Thrones. Therefore, the other data 820 may represent a smart TV or other display device, rather than a non-display device that captures the verbal user input.
[0174] Other data 820 may also include data indicating the speed, location, or other movement information of the device from which the user input originated. For example, the device may correspond to a vehicle including a display. If the vehicle is driving, the candidate list component 750 may reduce the confidence score associated with the domain that generated the video data because it may not be desirable to output video content to the user while the user is driving. The device may output data to the system 120 to indicate when the device is moving.
[0175] Other data 820 may also include data indicating the domain currently called. For example, a user may say a first (e.g., previous) user input that causes the system to call a music domain skill to output music to the user. While the system is outputting music to the user, the system may receive a second (e.g., current) user input. Candidate list component 750 may use such data to change the confidence score of a domain. For example, candidate list component 750 may run a first model specific to a first domain and a second model specific to a second domain. Alternatively, candidate list component 750 may run a model configured to determine the score of each domain. Candidate list component 750 may also determine the same confidence score for each domain in the first instance. Then, candidate list component 750 may be called based on the first domain to change the original confidence score so that the system outputs content when receiving the current user input. Based on the first domain being invoked, the candidate list component 750 may (i) increase the confidence score associated with the first domain while keeping the confidence score associated with the second domain unchanged, (ii) keep the confidence score associated with the first domain unchanged while decreasing the confidence score associated with the second domain, or (iii) increase the confidence score associated with the first domain and decrease the confidence score associated with the second domain.
[0176] The thresholding implemented with respect to the n-best list data 815 generated by the candidate list component 750 and the different types of other data 820 considered by the candidate list component 750 are configurable. For example, as more other data 820 are considered, the candidate list component 750 can update the confidence score. As another example, if thresholding is implemented, the n-best list data 815 can exclude relevant domains. Thus, for example, the candidate list component 750 can include an indication of a domain in the n-best list 815 unless the candidate list component 750 is 100% confident that the domain is likely unable to perform the user input represented in the ASR output data 615 (e.g., the candidate list component 750 determines that the confidence score for the domain is zero).
[0177] Candidate list component 750 may send ASR output data 615 to an identifier 763 associated with the domains represented in n-best list data 815. Alternatively, candidate list component 750 may send n-best list data 815 or some other indicator of a selected subset of domains to another component, such as orchestrator component 430, which in turn may send ASR output data 615 to an identifier 763 corresponding to the domains included in n-best list data 815 or otherwise indicated in the indicator. If candidate list component 750 generates an n-best list representing domains that do not have any associated confidence scores, candidate list component 750 / orchestrator component 430 may send ASR output data 615 to an identifier 763 associated with the domain that candidate list component 750 determined to be executable for the user input. If candidate list component 750 generates an n-best list representing domains having associated confidence scores, candidate list component 750 / orchestrator component 430 can send ASR output data 615 to an identifier 763 associated with domains associated with confidence scores that satisfy (e.g., meet or exceed) a threshold minimum confidence score.
[0178] The recognizer 763 can output the labeled text data generated by the NER component 762 and the IC component 764, as described above. The NLU component 460 can compile the output labeled text data of the recognizer 763 into a single cross-domain n-best list 840, and can send the cross-domain n-best list 840 to the pruning component 850. Each entry of labeled text represented in the cross-domain n-best list data 840 (e.g., each NLU hypothesis) can be associated with a corresponding score indicating the likelihood that the NLU hypothesis corresponds to the domain associated with the recognizer 763 that output the NLU hypothesis. For example, the cross-domain n-best list data 840 can be represented as (each row corresponds to a different NLU hypothesis):
[0179] [0.95] Intent: <Play music> Artist Name: Beethoven Song Title: Waldstein Sonata
[0180] [0.70] Intent: <play video> Artist Name: Beethoven Video Name: Waldstein Sonata
[0181] [0.01] Intent: <Play music> Artist Name: Beethoven Album Name: Waldstein Sonata
[0182] [0.01] Intent: <Play music> Song title: Waldstein Sonata
[0183] The pruning component 850 may sort the NLU hypotheses represented in the cross-domain n-best list data 840 according to their respective scores. The pruning component 850 may perform score thresholding on the cross-domain NLU hypotheses. For example, the pruning component 850 may select an NLU hypothesis associated with a score that satisfies (e.g., reaches and / or exceeds) a threshold score. The pruning component 850 may also or alternatively perform a number of NLU hypothesis thresholds. For example, the pruning component 850 may select the highest scoring NLU hypothesis. The pruning component 850 may output a portion of the NLU hypotheses input therein. The purpose of the pruning component 850 is to create a streamlined list of NLU hypotheses so that more resource-intensive downstream processes can operate only on the NLU hypotheses that are most likely to represent the user's intent.
[0184] The NLU component 460 may include a lightweight slot filler component 852. The lightweight slot filler component 852 may obtain text from the slot represented in the NLU hypothesis output by the pruning component 850 and modify it to make the text easier to be processed by downstream components. The lightweight slot filler component 852 may perform low-latency operations that do not involve heavy operations, such as referencing a knowledge base (e.g., 772). The purpose of the lightweight slot filler component 852 is to replace words with other words or values that may be easier for downstream components to understand. For example, if the NLU hypothesis includes the word "tomorrow", the lightweight slot filler component 852 may replace the word "tomorrow" with an actual date for downstream processing. Similarly, the lightweight slot filler component 852 may replace the word "CD" with "album" or the word "disc". The replaced words are then included in the cross-domain n best list data 860.
[0185] The cross-domain n best list data 860 can be input to the entity resolution component 870. The entity resolution component 870 can apply rules or other instructions to standardize the tags or tokens from the previous stage into intent / slot representations. The exact conversion may depend on the domain. For example, for the travel domain, the entity resolution component 870 can convert the text corresponding to "Boston Airport" into the standard BOS three-letter code referring to the airport. The entity resolution component 870 can refer to a knowledge base (e.g., 772), which is used to specifically identify the precise entities referenced in each slot of each NLU hypothesis represented in the cross-domain n best list data 860. A specific intent / slot combination may also be bound to a specific source, which can then be used to parse the text. In the example of "playing the Rolling Stones' songs", the entity resolution component 870 can reference a personal music catalog, an Amazon music account, a user profile, etc. The entity resolution component 870 can output a modified n-best list that is based on the cross-domain n-best list 860, but includes more detailed information about the specific entities mentioned in the slots (e.g., entity IDs) and / or more detailed slot data that can ultimately be used by skills. The NLU component 460 can include multiple entity resolution components 870, and each entity resolution component 870 can be specific to one or more domains.
[0186] The NLU component 460 may include a re-ranker 890. The re-ranker 890 may assign a specific confidence score to each NLU hypothesis input therein. The confidence score of a particular NLU hypothesis may be affected by whether the NLU hypothesis has unfilled slots. For example, if an NLU hypothesis includes all filled / resolved slots, the NLU hypothesis may be assigned a higher confidence score than another NLU hypothesis that includes at least some slots that were not filled / resolved by the entity resolution component 870.
[0187] The reorderer 890 may apply rescoring, biasing, or other techniques. The reorderer 890 may not only consider the data output by the entity resolution component 870, but also other data 891. Other data 891 may include various information. For example, other data 891 may include skill ratings or popularity data. For example, if a skill has a higher rating, the reorderer 890 may increase the score of the NLU hypothesis that the skill can handle. Other data 891 may also include information about the skills that the user who initiated the user input has enabled. For example, the reorderer 890 may assign a higher score to the NLU hypothesis that can be processed by the enabled skill than the NLU hypothesis that can be processed by the unenabled skill. Other data 891 may also include data indicating the user's usage history, such as whether the user who initiated the user input regularly uses a specific skill or uses a specific skill at a specific time of the day. Other data 891 may additionally include data indicating the date, time, location, weather, type of device 110, user identifier, context, and other information. For example, the reorderer 890 may consider when any specific skill is currently active (e.g., playing music, playing a game, etc.).
[0188] As shown and described, entity resolution component 870 is implemented before re-sequencer 890. Entity resolution component 870 may alternatively be implemented after re-sequencer 890. Implementing entity resolution component 870 after re-sequencer 890 limits the NLU hypotheses processed by entity resolution component 870 to only those that successfully passed re-sequencer 890.
[0189] The re-ranker 890 may be a global re-ranker (e.g., a re-ranker that is not specific to any particular domain). Alternatively, the NLU component 460 may implement one or more domain-specific re-rankers. Each domain-specific re-ranker may re-rank the NLU hypotheses associated with the domain. Each domain-specific re-ranker may output an n-best list of re-ranked hypotheses (e.g., 5 to 10 hypotheses).
[0190] NLU component 460 may perform the tasks described above with respect to skills fully implemented as part of system 120 (e.g., Figure 4 The NLU component 460 may perform NLU processing for domains associated with skills that are at least partially implemented as part of the skill processing component 125. In one example, the candidate list component 750 may process only with respect to these latter domains. The results of these two NLU processing paths may be combined into NLU output data 885, which may be sent to the NLU post-ranker 465 that may be implemented by the system 120.
[0191] The NLU post-ranker 465 may include a statistical component that generates a ranked list of intent / skill pairs with associated confidence scores. Each confidence score may indicate the adequacy of the skill to perform the intent with respect to the NLU result data associated with the skill. The NLU post-ranker 465 may operate one or more trained models configured to process the NLU result data 885, the skill result data 830, and the other data 820 to output ranked output data 825. The ranked output data 825 may include an n-best list, wherein the NLU hypotheses in the NLU result data 885 are re-ranked so that the n-best list in the ranked output data 825 represents a prioritized list of skills determined by the NLU post-ranker 465 in response to the user input. The ranked output data 825 may also include (as part of the n-best list or otherwise) respective corresponding scores corresponding to the skills, wherein each score indicates the probability that the skill (and / or its corresponding result data) corresponds to the user input.
[0192] The system may be configured with thousands, tens of thousands, etc. of skills. The NLU post-ranker 465 enables the system to better determine the best skill to execute the user input. For example, the first NLU hypothesis and the second NLU hypothesis in the NLU result data 885 may substantially correspond to each other (e.g., their scores may be very similar), even though the first NLU hypothesis may be processed by the first skill and the second NLU hypothesis may be processed by the second skill. The first NLU hypothesis may be associated with a first confidence score indicating the confidence of the system regarding the NLU processing performed to generate the first NLU hypothesis. In addition, the second NLU hypothesis may be associated with a second confidence score indicating the confidence of the system regarding the NLU processing performed to generate the second NLU hypothesis. The first confidence score may be similar or the same as the second confidence score. The first confidence score and / or the second confidence score may be a numerical value (e.g., from 0.0 to 1.0). Alternatively, the first confidence score and / or the second confidence score may be a graded value (e.g., low, medium, high).
[0193] The NLU post-ranker 465 (or other scheduling component, such as the orchestrator component 430) may solicit the first skill and the second skill to provide potential result data 830 based on the first NLU hypothesis and the second NLU hypothesis, respectively. For example, the NLU post-ranker 465 may send the first NLU hypothesis to the first skill 490a along with a request to at least partially perform the first skill 490a with respect to the first NLU hypothesis. The NLU post-ranker 465 may also send the second NLU hypothesis to the second skill 490b along with a request to at least partially perform the second skill 490b with respect to the second NLU hypothesis. The NLU post-ranker 465 receives from the first skill 490a the first result data 830a generated by performing the first skill 490a with respect to the first NLU hypothesis. The NLU post-ranker 465 also receives from the second skill 490b the second result data 830b generated by performing the second skill 490b with respect to the second NLU hypothesis.
[0194] The result data 830 may include various parts. For example, the result data 830 may include content to be output to the user (e.g., audio data, text data, and / or video data). The result data 830 may also include a unique identifier used by the system 120 and / or skill processing component 125 to locate the data to be output to the user. The result data 830 may also include instructions. For example, if the user input corresponds to "turn on the light", the result data 830 may include instructions for the system to turn on the light associated with the device (110a / 110b) and / or the user's profile.
[0195] The NLU post-ranker 465 may consider the first result data 830a and the second result data 830b to change the first confidence score of the first NLU hypothesis and the second confidence score of the second NLU hypothesis, respectively. That is, the NLU post-ranker 465 may generate a third confidence score based on the first result data 830a and the first confidence score. The third confidence score may correspond to how likely the NLU post-ranker 465 determines that the first skill will correctly respond to the user input. The NLU post-ranker 465 may also generate a fourth confidence score based on the second result data 830b and the second confidence score. Those skilled in the art will appreciate that the first difference between the third confidence score and the fourth confidence score may be greater than the second difference between the first confidence score and the second confidence score. The NLU post-ranker 465 may also consider other data 820 to generate the third confidence score and the fourth confidence score. Although it has been described that the NLU post-ranker 465 can change the confidence scores associated with the first NLU hypothesis and the second NLU hypothesis, it will be appreciated by those skilled in the art that the NLU post-ranker 465 can change the confidence scores of more than two NLU hypotheses. The NLU post-ranker 465 can select the result data 830 associated with the skill 490 having the highest changed confidence score as data output in response to the current user input. The NLU post-ranker 465 can also consider the ASR output data 615 to change the NLU hypothesis confidence score.
[0196] The orchestrator component 430 may associate the intents in the NLU hypothesis with the skills 490 before sending the NLU result data 885 to the NLU post-ranker 465. For example, if the NLU hypothesis includes the <play music> intent, the orchestrator component 430 may associate the NLU hypothesis with one or more skills 490 that can perform the <play music> intent. Thus, the orchestrator component 430 may send the NLU result data 885 including the NLU hypothesis paired with the skills 490 to the NLU post-ranker 465. In response to the ASR output data 615 corresponding to “What should I make for dinner today”, the orchestrator component 430 may generate a skill pair 490 with associated NLU hypotheses corresponding to:
[0197] Skill 1 / NLU assumes that the <help> intent is included
[0198] Skill 2 / NLU assumption includes the intent of <order food>
[0199] Skill 3 / NLU assumption includes <disk type> intent
[0200] The NLU post-ranker 465 queries each skill 490 paired with an NLU hypothesis in the NLU output data 885 to provide result data 830 based on the NLU hypothesis associated with the skill. That is, with respect to each skill, the NLU post-ranker 465 colloquially asks each skill "if given this NLU hypothesis, what would you do with it". Based on the above example, the NLU post-ranker 465 may send the following data to the skill 490:
[0201] Skill 1: First NLU assumption including the <help> intent indicator
[0202] Skill 2: Second NLU hypothesis including the <order food> intent indicator
[0203] Skill 3: Third NLU hypothesis including <disk type> intent indicator
[0204] The NLU post-ranker 465 may query each skill 490 in parallel or substantially in parallel.
[0205] In response to the NLU post-ranker 465 soliciting the skill 490 for result data 830, the skill 490 may provide various data and indications to the NLU post-ranker 465. The skill 490 may simply provide the NLU post-ranker 465 with an indication of whether the NLU hypothesis skill it received is executable. The skill 490 may also or alternatively provide the NLU post-ranker 465 with output data generated based on the NLU hypothesis it received. In some cases, the skill 490 may require further information beyond what is represented in the received NLU hypothesis to provide output data responsive to the user input. In these cases, the skill 490 may provide the NLU post-ranker 465 with result data 830 indicating the slots of the framework that the skill 490 further needs to fill or the entities that the skill 490 further needs to resolve before the skill 490 can provide the result data 830 responsive to the user input. Skill 490 may also provide instructions and / or computer generated speech to NLU post-ranker 465 indicating how skill 490 should recommend that the system solicit further information required by skill 490. Skill 490 may also provide instructions to NLU post-ranker 465: whether skill 490 has all the required information after the user provides the additional information once, or whether skill 490 requires the user to provide various additional information before skill 490 has all the required information. Based on the above example, skill 490 may provide the following to NLU post-ranker 465:
[0206] Skill 1: Indicates that the skill can perform an NLU hypothesis that includes the <help> intent indicator Skill 2: Indicates that the skill requires the system to obtain further information
[0207] Skill 3: indicates that the skill can provide an indication of many results in response to the third NLU hypothesis including the <disc type> intent indicator
[0208] Result data 830 includes indications provided by skill 490 indicating whether skill 490 can execute the NLU hypothesis; data generated by skill 490 based on the NLU hypothesis; and indications provided by skill 490 indicating that skill 490 requires further information beyond what is represented in the received NLU hypothesis.
[0209] The NLU post-ranker 465 uses the result data 830 provided by the skills 490 to change the NLU processing confidence scores generated by the re-ranker 890. That is, the NLU post-ranker 465 uses the result data 830 provided by the queried skills 490 to create a larger difference between the NLU processing confidence scores generated by the re-ranker 890. Without the NLU post-ranker 465, the system may not have enough confidence to determine the output in response to the user input, such as when the NLU hypotheses associated with multiple skills are too close for the system to confidently determine to invoke a single skill 490 in response to the user input. For example, if the system does not implement the NLU post-ranker 465, the system may not be able to determine whether to obtain output data from the general reference information skill or the medical information skill in response to the user input corresponding to "what is acne".
[0210] The NLU post-ranker 465 may prefer skills 490 that provide result data 830 that respond to NLU hypotheses over skills 490 that provide result data 830 that corresponds to an indication that further information is needed, and skills 490 that provide result data 830 that indicates that the skill can provide multiple responses to the received NLU hypothesis. For example, based on the first skill 490a providing result data 830a that includes a response to the NLU hypothesis, the NLU post-ranker 465 may generate a first score for the first skill 490a, the first score being higher than the NLU confidence score of the first skill. For another example, based on the second skill 490b providing result data 830b that indicates that the second skill 490b needs further information to provide a response to the NLU hypothesis, the NLU post-ranker 465 may generate a second score for the second skill 490b, the second score being lower than the NLU confidence score of the second skill. For another example, based on the third skill 490c providing result data 830c indicating that the third skill 490c can provide multiple responses to the NLU hypothesis, the NLU post-ranser 465 can generate a third score for the third skill 490c, which is lower than the NLU confidence score of the third skill.
[0211] The NLU post-ranker 465 may consider other data 820 when determining the score. The other data 820 may include a ranking associated with the queried skill 490. The ranking may be a system ranking or a user-specific ranking. The ranking may indicate the authenticity of the skill from the perspective of one or more users of the system. For example, based on the first skill 490a being associated with a high ranking, the NLU post-ranker 465 may generate a first score for the first skill 490a, the first score being higher than the NLU processing confidence score of the first skill. For another example, based on the second skill 490b being associated with a low ranking, the NLU post-ranker 465 may generate a second score for the second skill 490b, the second score being lower than the NLU processing confidence score of the second skill.
[0212] Other data 820 may include information indicating whether the user who initiated the user input has enabled one or more of the queried skills 490. For example, based on the first skill 490a being enabled by the user who initiated the user input, the NLU post-ranker 465 may generate a first score for the first skill 490a, the first score being higher than the NLU processing confidence score for the first skill. For another example, based on the second skill 490b not being enabled by the user who initiated the user input, the NLU post-ranker 465 may generate a second score for the second skill 490b, the second score being lower than the NLU processing confidence score for the second skill. When the NLU post-ranker 465 receives the NLU result data 885, the NLU post-ranker 465 may determine whether the profile data associated with the user and / or device that initiated the user input includes an indication of an enabled skill.
[0213] Other data 820 may include information indicating the output capabilities of a device that will be used to output content to a user in response to user input. The system may include: a device that includes a speaker but does not include a display, a device that includes a display but does not include a speaker, and a device that includes a speaker and a display. If the device that will output the content in response to the user input includes one or more speakers but does not include a display, the NLU post-sorter 465 may increase the NLU processing confidence score associated with the first skill configured to output audio data and / or reduce the NLU processing confidence score associated with the second skill configured to output visual data (e.g., image data and / or video data). If the device that will output the content in response to the user input includes a display but does not include one or more speakers, the NLU post-sorter 465 may increase the NLU processing confidence score associated with the first skill configured to output visual data and / or reduce the NLU processing confidence score associated with the second skill configured to output audio data.
[0214] The other data 820 may include information indicating the authenticity of the result data 830 provided by the skill 490. For example, if the user says "tell me a recipe for spaghetti sauce", the first skill 490a may provide the NLU post-ranker 465 with first result data 830a corresponding to the first recipe associated with a five-star rating, and the second skill 490b may provide the NLU post-ranker 465 with second result data 830b corresponding to the second recipe associated with a one-star rating. In this case, the NLU post-ranker 465 may increase the NLU processing confidence score associated with the first skill 490a based on the first skill 490a providing the first result data 830a associated with a five-star rating and / or decrease the NLU processing confidence score associated with the second skill 490b based on the second skill 490b providing the second result data 830b associated with a one-star rating.
[0215] Other data 820 may include information indicating the type of device that initiated the user input. For example, if the device is located in a hotel room, the device may correspond to a "hotel room" type. If the user enters a command corresponding to "order me food" to a device located in a hotel room, the NLU post-ranker 465 may increase the NLU processing confidence score associated with the first skill 490a corresponding to the room service skill associated with the hotel and / or decrease the NLU processing confidence score associated with the second skill 490b corresponding to the food skill not associated with the hotel.
[0216] Other data 820 may include information indicating the location of the device and / or user that initiated the user input. The system may be configured with skills 490 that operate only with respect to certain geographic locations. For example, a user may provide user input corresponding to "When is the next train to Portland?" The first skill 490a may operate with respect to trains that arrive, depart, and pass through Portland, Oregon. The second skill 490b may operate with respect to trains that arrive, depart, and pass through Portland, Maine. If the device and / or user that initiated the user input is located in Seattle, Washington, the NLU post-ranker 465 may increase the NLU processing confidence score associated with the first skill 490a and / or reduce the NLU processing confidence score associated with the second skill 490b. Similarly, if the device and / or user that initiated the user input is located in Boston, Massachusetts, the NLU post-ranker 465 may increase the NLU processing confidence score associated with the second skill 490b and / or reduce the NLU processing confidence score associated with the first skill 490a.
[0217] Other data 820 may include information indicating the time of day. The system may be configured with skills 490 that operate with respect to certain times of day. For example, a user may provide user input corresponding to "order me food." The first skill 490a may generate first result data 830a corresponding to breakfast. The second skill 490b may generate second result data 830b corresponding to dinner. If the system 120 receives the user input in the morning, the NLU post-ranker 465 may increase the NLU processing confidence score associated with the first skill 490a and / or decrease the NLU processing score associated with the second skill 490b. If the system 120 receives the user input in the afternoon or evening, the NLU post-ranker 465 may increase the NLU processing confidence score associated with the second skill 490b and / or decrease the NLU processing confidence score associated with the first skill 490a.
[0218] Other data 820 may include information indicating user preferences. The system may include multiple skills 490 configured to be executed in substantially the same manner. For example, the first skill 490a and the second skill 490b may both be configured to order food from respective restaurants. The system may store (e.g., in the profile storage 470) user preferences associated with the user who provided the user input to the system 120, and indicate that the user prefers the first skill 490a over the second skill 490b. Therefore, when the user provides user input that both the first skill 490a and the second skill 490b can perform, the NLU post-ranker 465 may increase the NLU processing confidence score associated with the first skill 490a and / or reduce the NLU processing confidence score associated with the second skill 490b.
[0219] Other data 820 may include information indicating a system usage history associated with a user initiating user input. For example, the system usage history may indicate that the user initiates user input that invokes the first skill 490a more frequently than the user initiates user input that invokes the second skill 490b. Based on this, if the current user input can be performed by both the first skill 490a and the second skill 490b, the NLU post-ranker 465 may increase the NLU processing confidence score associated with the first skill 490a and / or decrease the NLU processing confidence score associated with the second skill 490b.
[0220] Other data 820 may include information indicating the speed of travel of the device 110 that initiated the user input. For example, the device 110 may be located in a moving vehicle, or may be a moving vehicle. When the device 110 is in motion, the system may prefer audio output over visual output to reduce the possibility of distracting the user (e.g., the driver of the vehicle). Thus, for example, if the device 110 that initiated the user input is moving at or above a threshold speed (e.g., a speed above the average user walking speed), the NLU post-sorter 465 may increase the NLU processing confidence score associated with the first skill 490a that generates audio data. The NLU post-sorter 465 may also or alternatively reduce the NLU processing confidence score associated with the second skill 490b that generates image data or video data.
[0221] The other data 820 may include information indicating how long it takes a skill 490 to provide result data 830 to the NLU post-ranker 465. When the NLU post-ranker 465 uses multiple skills 490 for result data 830, the skills 490 may respond to queries at different speeds. The NLU post-ranker 465 may implement a latency budget. For example, if the NLU post-ranker 465 determines that a skill 490 responds to the NLU post-ranker 465 within a threshold amount of time after receiving a query from the NLU post-ranker 465, the NLU post-ranker 465 may increase the NLU processing confidence score associated with the skill 490. Conversely, if the NLU post-ranker 465 determines that a skill 490 does not respond to the NLU post-ranker 465 within a threshold amount of time after receiving a query from the NLU post-ranker 465, the NLU post-ranker 465 may decrease the NLU processing confidence score associated with the skill 490.
[0222] It has been described that the NLU post-ranker 465 uses other data 820 to increase and decrease the NLU processing confidence scores associated with various skills 490 for which the NLU post-ranker 465 has requested result data. Alternatively, the NLU post-ranker 465 may use other data 820 to determine which skills 490 to request result data from. For example, the NLU post-ranker 465 may use other data 820 to increase and / or decrease the NLU processing confidence scores associated with the skills 490 associated with the NLU result data 885 output by the NLU component 460. The NLU post-ranker 465 may select the n highest scoring changed NLU processing confidence scores. The NLU post-ranker 465 may then request result data 830 only from the skills 490 associated with the selected n NLU processing confidence scores.
[0223] As described above, the NLU post-ranker 465 may request result data 830 from all skills 490 associated with the NLU result data 885 output by the NLU component 460. Alternatively, the system 120 may prefer result data 830 from skills that are fully implemented by the system 120 rather than skills that are at least partially implemented by the skill processing component 125. Thus, in the first case, the NLU post-ranker 465 may request result data 830 only from skills that are associated with the NLU result data 885 and that are fully implemented by the system 120. If none of the skills (fully implemented by the system 120) provide the NLU post-ranker 465 with result data 830 (indicating a data response to the NLU result data 885), an indication that the skill can perform a user input, or an indication that further information is needed, the NLU post-ranker 465 may request result data 830 only from skills that are associated with the NLU result data 885 and that are at least partially implemented by the skill processing component 125.
[0224] As described above, the NLU post-ranker 465 may request result data 830 from multiple skills 490. If one of the skills 490 provides result data 830 indicating a response to an NLU hypothesis and the other skills provide result data 830 indicating that they cannot perform or that they need further information, the NLU post-ranker 465 may select the result data 830 including the response to the NLU hypothesis as the data to be output to the user. If more than one skill 490 provides result data 830 indicating a response to the NLU hypothesis, the NLU post-ranker 465 may consider the other data 820 to generate a varying NLU processing confidence score, and select the result data 830 of the skill associated with the highest score as the data to be output to the user.
[0225] A system that does not implement the NLU post-ranker 465 may select the highest scoring NLU hypothesis in the NLU result data 885. The system may send the NLU hypothesis to the skill 490 associated with it along with a request for output data. In some cases, the skill 490 may not be able to provide output data to the system. This causes the system to indicate to the user that it cannot process the user input, even though another skill associated with a lower ranked NLU hypothesis may provide output data in response to the user input.
[0226] The NLU post-ranker 465 reduces instances of the above situation. As described above, the NLU post-ranker 465 queries multiple skills associated with the NLU result data 885 to provide the result data 830 to the NLU post-ranker 465 before the NLU post-ranker 465 finally determines the skill 490 to be invoked in response to the user input. Some skills 490 may provide result data 830 indicating a response to the NLU hypothesis, while other skills 490 may provide result data 830 indicating that the skill cannot provide response data. Although a system that does not implement the NLU post-ranker 465 may select one of the skills 490 that cannot provide a response, the NLU post-ranker 465 only selects the skill 490 that provides the NLU post-ranker 465 with result data corresponding to the response, the result data indicating that further information is required or indicating that multiple responses can be generated.
[0227] The NLU post-ranker 465 may select the result data 830 associated with the skill 490 associated with the highest score for output to the user. Alternatively, the NLU post-ranker 465 may output ranked output data 825 indicating the skills 490 and their respective NLU post-ranker rankings. Since the NLU post-ranker 465 receives result data 830 from the skills 490 that may correspond to responses to user input before the NLU post-ranker 465 selects one of the skills or outputs ranked output data 825, little to no delay occurs from the time the skill provides the result data 830 to the time the system outputs the response to the user.
[0228] If the NLU post-ranker 465 selects the resulting audio data to be output to the user and the system determines that the content should be output in an audible manner, the NLU post-ranker 465 (or another component of the system 120) may cause the device 110a and / or the device 110b to output audio corresponding to the resulting audio data. If the NLU post-ranker 465 selects the resulting text data to be output to the user and the system determines that the content should be output in a visual manner, the NLU post-ranker 465 (or another component of the system 120) may cause the device 110b to display text corresponding to the resulting text data. If the NLU post-ranker 465 selects the resulting audio data to be output to the user and the system determines that the content should be output in a visual manner, the NLU post-ranker 465 (or another component of the system 120) may send the resulting audio data to the ASR component 450. The ASR component 450 may generate output text data corresponding to the resulting audio data. Then, the system 120 may cause the device 110b to display text corresponding to the output text data. If NLU post-ranker 465 selects the result text data to be output to the user and the system determines that the content should be output in an audible manner, NLU post-ranker 465 (or another component of system 120) can send the result text data to TTS component 480. TTS component 480 can generate output audio data (corresponding to computer-generated speech) based on the result text data. System 120 can then cause device 110a and / or device 110b to output audio corresponding to the output audio data.
[0229] As described above, skill 490 may provide result data 830 that either indicates a response to the user input, indicates that skill 490 requires more information to provide a response to the user input, or indicates that skill 490 is unable to provide a response to the user input. If the skill 490 associated with the highest NLU post-ranker score provides NLU post-ranker 465 with result data 830 indicating a response to the user input, then NLU post-ranker 465 (or another component of system 120, such as orchestrator component 430) may simply cause content corresponding to result data 830 to be output to the user. For example, NLU post-ranker 465 may send result data 830 to orchestrator component 430. Orchestrator component 430 may cause result data 830 to be sent to a device (110a / 110b), which may output audio and / or display text corresponding to result data 830. As appropriate, arranger component 430 may send result data 830 to ASR component 450 to generate output text data and / or may send result data 830 to TTS component 480 to generate output audio data.
[0230] The skill 490 associated with the highest NLU post-ranker score may provide the NLU post-ranker 465 with result data 830 indicating that more information is needed, as well as instruction data. The instruction data may indicate how the skill 490 recommends that the system obtain the required information. For example, the instruction data may correspond to text data or audio data (i.e., computer-generated speech) corresponding to "Please indicate ________________". The instruction data may be in a format (e.g., text data or audio data) that the device (110a / 110b) is able to output. When this occurs, the NLU post-ranker 465 may simply cause the received instruction data to be output by the device (110a / 110b). Alternatively, the instruction data may be in a format that the device (110a / 110b) cannot output. When this occurs, the NLU post-ranker 465 may cause the ASR component 450 or the TTS component 480 to process the instruction data, as appropriate, to generate instruction data that can be output by the device (110a / 110b). Once the user has provided the system with all further information required by the skill 490, the skill 490 may provide the system with result data 830 indicating a response to the user input, which may be output by the system as described in detail above.
[0231] The system may include "informational" skills 490, which simply provide information to the system, which the system outputs to the user. The system may also include "transactional" skills 490, which require system instructions to execute user input. Transactional skills 490 include ride sharing skills, flight booking skills, etc. Transactional skills 490 may simply provide NLU post-sequencer 465 with result data 830 indicating that the transactional skill 490 can execute the user input. The NLU post-sequencer 465 may then cause the system to solicit an indication from the user to allow the system to cause the transactional skill 490 to execute the user input. The indication provided by the user may be an audible indication or a tactile indication (e.g., activation of a virtual button or text input through a virtual keyboard). In response to receiving the indication provided by the user, the system may provide data corresponding to the indication to the transactional skill 490. In response, the transactional skill 490 may execute the command (e.g., book a flight, book a train ticket, etc.). Thus, while the system may not further engage the informative skill 490 after the informative skill 490 provides result data 830 to the NLU post-ranker 465, the system may further engage the transactional skill 490 after the transactional skill 490 provides result data 830 to the NLU post-ranker 465 indicating that the transactional skill 490 can execute the user input.
[0232] In some cases, the NLU post-ranker 465 may generate respective scores for the first skill and the second skill that are too close (e.g., at least not a threshold difference) for the NLU post-ranker 465 to confidently determine which skill should perform the user input. When this occurs, the system may request the user to indicate which skill the user prefers to perform the user input. The system may output TTS-generated speech to the user to solicit which skill the user wants to perform the user input.
[0233] Other data 820 and / or other data 891 may include state data 195 that indicates the state of a particular operation of a particular speech processing system 120. For example, a first system 120a may have access to first system state data 195a that indicates what interactions a particular user profile / device 110 has had with the first system 120a, while a second system 120b may have access to second system state data 195b that indicates what interactions a particular user profile / device 110 has had with the second system 120b. It will be appreciated that the first system state data 195a will be different from the second system state data 195b, and that the first system 120a may not have access to the second system state data 195b, while the second system 120b may not have access to the first system state data 195a. (The same is true for the other systems and their respective state data 195.) The state data 195 of the respective systems as part of the other data 820 and / or other data 891 allows the particular system 120 to interpret and / or sort incoming requests in a manner consistent with previous interactions with the particular system 120, because such interactions may be represented in the state data 195.
[0234] Speech processing (such as the speech processing described above) may be based on device / profile state data 194. Specifically, other data 820 and / or other data 891 may also include state data 194, which indicates which device process controls can be performed by the requesting device 110 and / or a device associated with the user profile of the requesting device (e.g., as indicated by state data 194m). As described above, certain state data 194 related to the device 110 and / or user profile may be available to the assistant system 120. Such state data 194 (e.g., 194a, 194b, etc.) may have some overlap with state data 194m available to device components (e.g., multi-assistant component 115). State data available to the first assistant system 120a may not include information related to the second assistant system 120b. For example, the state data 194b available to the second system 120b may not include information identifying the operations performed by the first system 120a. This may be due to privacy perception issues, security issues, system configuration, etc. However, allowing access to certain limited state data 194 enables the system 100 to allow multiple assistant systems 120, even assistant systems that may not have initiated a particular device process, to control a device process.
[0235] As described above, a device process may involve controlling a process involving an action to be performed by the device 110. Such device process control may include, for example, starting / stopping a timer, setting / stopping an alarm, playing / stopping media content (such as a song, video, podcast, etc.), controlling output content (such as skipping a song, rewinding a song, extending / pausing a timer / alarm, stopping synthesized speech output, etc.), setting a temperature (e.g., if the device can operate as a thermostat), activating / deactivating components of the device (such as a camera, lights, etc.), controlling device settings (such as volume, brightness, sensitivity, etc.), setting / controlling reminders, initiating / controlling / terminating a call or call request, etc. Thus, a device process control may control the device to transition from a first state (e.g., outputting audio, displaying something on a display) to a second state (e.g., stopping audio output, outputting audio at a different volume, displaying something else on a display, removing something from a display, etc.).
[0236] Thus, state data 194 (e.g., 194m) included in other data 820 and / or other data 891 may indicate which such device process controls may be performed by device 110 and / or other devices associated with a particular user profile. To determine this state data 194, system 120 may obtain information from the device (e.g., as described above with reference to Figure 2A and Figure 2B As shown, the system 120 receives state data 194 from the multi-assistant component 115). In another example, the system 120 may receive an indicator of the device 110 / user profile of the incoming request and may use the indicator to obtain the appropriate state data 194 from another source. For example, the system 120 may communicate with the storage component corresponding to the device / user profile to obtain the relevant state data 194.
[0237] The resequencer 890 and / or NLU post-sequencer 465 may use the device / profile state data 194 to interpret the ASR data 615 / select a particular NLU hypothesis for interpreting the incoming utterance as an utterance requesting control of a device process. For example, a user may utter a command such as “Alexa, stop.” Without access to the device / profile state data 194, the NLU component 460b of the second system 120b may determine a potential NLU hypothesis of “stop music,” but if the state data 195b does not indicate active music playback with respect to the second system 120b and the particular requesting device 110, then the interpretation of “stop music” may be ranked lower (e.g., by the resequencer 890b and / or NLU post-sequencer 465b) because the second system 120b is not aware of any active music playback. However, if second system 120b has access to device / profile state data 194, and device / profile state data 194 indicates that device 110 (or another device requesting the profile) is capable of executing a stop music command, then re-sequencer 890 and / or NLU post-sequencer 465 may rank the interpretation of “stop music” higher. Thus, in the case where the “play music” command was previously spoken to an assistant / system other than second system 120b, second system 120b may still be able to correctly interpret a command such as “Alexa, stop” as a command to control a device process controllable by device 110 (as indicated by device / profile state data 194), even if second system 120b has no information about the specific active music playback, what music is playing, how to start it, etc.
[0238] In the above example, by using device / profile state data 194, NLU 460b and / or NLU post-sequencer 465b may correctly interpret "Alexa, stop" as a command to control a device process. Thus, the selected hypothesis may correspond to NLU result data 825 / 885 indicating a <stop music> command and also indicating that the destination of such command should be device skill 191b. NLU result data 825 / 885 may be output by NLU 460b and / or NLU post-sequencer 465b along with an indicator linked to the requesting device 110 / user profile. The indicator may be an indicator of the requesting device 110, the requesting user profile, a particular utterance, or other indicator that may be used by another component (e.g., orchestrator 230) to link NLU result data 825 / 885 to the original request / utterance.
[0239] NLU 460 may also use device / profile state data 194 to correctly determine which device an incoming request corresponds to. For example, an incoming request such as “Alexa, stop” may be received by a user such as smartwatch 110c (in Fig.12 10c). The speech processing system 120 may receive an indication of a user profile corresponding to the smart watch 110c and device / profile state data 194 for the smart watch 110c and / or other devices 110 associated with the particular user profile. The device / profile state data 194 may indicate that the music playing device (e.g., the speech detection device 110a, the vehicle 110e, the home audio system, etc.) is capable of executing the <stop music> command, while the smart watch 110c itself may not be able to execute such a command. The NLU 460 and / or the NLU post-ranker 465 may use the device / profile state data 194 to interpret the incoming request / its ranking hypothesis to determine that the "Alexa, stop" request corresponds to a different target device than the smart watch 110c that captured the input utterance. Thus, the resulting NLU result data 825 / 885 may indicate the target device so that the device skill 191 may send output data to the appropriate device (e.g., the speech detection device 110a, the vehicle 110e, the home audio system, etc.) to execute the stop music command.
[0240] NLU result data 825 / 885 may be sent (e.g., by orchestrator 230 and / or another component) to device skill 191b. Alternatively or in addition, orchestrator 230 and / or another component may process NLU result data 825 / 885 to determine other data representing the requested device process control (e.g., <stop music>), where the other data is in a different form that device skill 191b can process. Device skill 191b may then obtain input data representing the user request (where the input data may be NLU result data 825 / 885 or data in some other form), and may determine output data that a device component (e.g., multi-assistant component 115) may act on to cause the device process control to be performed, and may send the output data to a particular device 110, e.g., as described above with reference to Figure 2A and 2B Explained.
[0241] In order to reduce the amount of information shared between speech processing systems 120, the state data 194 available to the speech processing systems 120 may be configured to relate only to certain (certain) possible device control processes for a particular device. In one embodiment, the device / profile state data 194 available to the speech processing systems 120 may include only device executable controls related to active device processes. For example, if a device 110 is outputting music, the state data 194 associated with the device 110 available to the speech processing system 120 may include information related to specific controls executable for the music processing of the particular device 110 (e.g., stop music, volume control, pause playback, etc.), but may not include information related to controls executable by the particular device 110 but not related to active music playback (e.g., stop alarm, extend timer, etc.). In a different example, if a device 110 is outputting a beep associated with an expired timer, the state data 194 associated with that device 110 that is available to the speech processing system 120 may include information related to specific controls of timer processes executable by that particular device 110 (e.g., stopping a timer, extending a timer, etc.), but does not include any information related to other inactive device processes. The determination of which device processes are active can be made by a component of the respective device 110 (e.g., by the multi-assistant component 115 or other component). The multi-assistant component 115 can then select a portion of the state data 194 associated with those active processes and send the selected portion to the speech processing system 120 (e.g., as a Figure 2A and Figure 2B shown).
[0242] In some cases, state data 194 may include priority data corresponding to one or more device processes. For example, if a device has multiple controllable processes, some of which may be active at a particular time, state data 194 may indicate the relative priorities of those processes. Speech processing system 120 may process priority data to determine an interpretation of an input request (e.g., using NLU 460) and / or a ranking of different NLU hypotheses for the interpretation (e.g., using reranker 890 and / or NLU postranker 465). Example priority data may take the following form:
[0243] 1.Activity Process A
[0244] 2. Activity Process B
[0245] 3. Inactive process C
[0246] 4. Inactive process D
[0247] Thus, if device 110 is capable of controlling four processes (A through D in the above example), two of which are active (A and B), if a user speaks a command to the device "Alexa stop", and speech processing system 120 (e.g., using state data 194) determines that the "stop" command may apply to either process A or process B, then, based on the priority data, the NLU hypothesis corresponding to process A may be ranked higher than the NLU hypothesis corresponding to process B. In a specific example, process A may correspond to a beep timer and process B may correspond to music playback. Thus, if a user speaks to device 110 "Alexa stop" without indicating what should be stopped, the system may prioritize stopping the timer and causing the timer to stop instead of the music playback.
[0248] Priority may also be affected by the system state data 195. Thus, for example, if the first system state data 195a of the first system 120a indicates that there is an ongoing process associated with the first system 120a that can be controlled with a "stop" command, and the user invokes the first system 120a (e.g., by saying the wake word "Alexa," where "Alexa" invokes the first system 120a), the process known to the first system 120a may be prioritized. Taking the specific example above as an example, if the process known to the first system 120a is actually music playback, even if the device / profile state data 194 indicates that timer control has a higher priority, if the first system 120a is performing voice processing, its available information in the first system state data 195a may cause the NLU hypothesis corresponding to music playback to be ranked higher than the NLU hypothesis corresponding to time control. Such a priority order may be configured in various ways depending on the configuration of the system 100.
[0249] In some configurations, priority data corresponding to one or more device control processes may influence the score given to a particular NLU hypothesis by the re-ranker 890 and / or the NLU post-ranker 465. For example, the priority data may take the following form:
[0250] 1.Activity process A[0.85]
[0251] 2.Activity process B[0.75]
[0252] 3. Inactive processes C[0.55]
[0253] 4. Inactive processes D[0.45]
[0254] Modifiers (e.g., [0.85] for process A, [0.75] for process B, etc.) may be used to adjust the scores of hypotheses corresponding to particular processes to create result scores for use by the re-ranker 890 and / or the NLU post-ranker 465. In this manner, priority data may be used to determine the scores and / or rankings of hypotheses corresponding to one or more device processes.
[0255] The device / profile state data 194 may also indicate the types of device processes that the device 110 may control. For example, one type may be a conversation, another type may be a notification, another type may be a timer, another type may be a media output (e.g., audio or video), another type may involve a call (e.g., video or audio communication), etc. The types may also correspond to components of the device used for the process, such as a speaker, a display, etc. The device / profile state data 194 may also indicate which components of the device 110 may be active at any particular time, which the speech processing system 120 may use to interpret the utterance. For example, if a user says "Alexa, quiet!", the speech processing system 120 may use the device / profile state data 194 to prioritize control of device processes involving audio output.
[0256] Fig. 9 Components of a system that may be used to perform unit selection, parametric TTS processing, and / or model-based audio synthesis are shown in FIG. Fig. 9 9 is a conceptual diagram illustrating the operation of generating synthesized speech using a TTS system 480 according to an embodiment of the present disclosure. The TTS system 480 can receive text data 915 and process it using one or more TTS models 980 to generate synthesized speech in the form of spectrogram data 945. The vocoder 990 can convert the spectrogram data 945 into output speech audio data 995, which can represent a time domain waveform suitable for amplification and output as audio (e.g., from a speaker).
[0257] The TTS system 480 may additionally receive other input data 925. The other input data 925 may include, for example, identifiers and / or tags corresponding to speaker identities, voice characteristics, emotions, voice styles, etc., required for synthesized speech. In some implementations, the other input data 925 may include text tags or text metadata, which may indicate, for example, how a particular word should be pronounced, for example, by indicating the desired output voice quality in tags formatted according to Speech Synthesis Markup Language (SSML) or in some other form. For example, a first text tag may be included in the text marking when the text should be whispered (e.g., <begin whisper>), and a second tag may be included in the text marking when the text should be whispered (e.g., <end whisper>). Tags may be included in the text data 915 and / or other input data 925, such as metadata that accompanies a TTS request and indicates what text should be whispered (or has some other indicated audio characteristics).
[0258] The TTS system 480 may include a pre-processing component 920 that can convert text data 915 and / or other input data 925 into a form suitable for processing by the TTS model 980. The text data 915 may come from, for example, an application, a skill component (further described below), an NLG component, another device or source, or may be input by a user. The text data 915 received by the TTS system 480 is not necessarily text, but may include other data (such as symbols, codes, other data, etc.) that may reference the text to be synthesized (such as indicators of words and / or phonemes). The pre-processing component 920 can convert the text data 915 into, for example, a symbolic language representation for processing by the TTS system 480, and the symbolic language representation may include language context features, such as phoneme data, punctuation data, syllable-level features, word-level features, and / or emotions, speakers, accents, or other features. Syllable-level features may include syllable emphasis, syllable speech rate, syllable pitch change, or other such syllable-level features; word-level features may include word emphasis, word speech rate, word pitch change, or other such word-level features. Emotional features may include data corresponding to an emotion associated with the text data 915, such as surprise, anger, or fear. Speaker features may include data corresponding to a speaker type, such as gender, age, or occupation. Accent features may include data corresponding to an accent associated with a speaker, such as a Southern accent, a Boston accent, a British accent, a French accent, or other such accents. Style features may include reading style, poetry recitation style, news anchor style, sports commentator style, various singing styles, and the like.
[0259] The preprocessing component 920 may include functions and / or components for performing text normalization, linguistic analysis, language prosody generation, or other such operations. During text normalization, the preprocessing component 920 may first process the text data 915 and generate standard text, converting content such as numbers, abbreviations (such as Apt., St., etc.), symbols ($, %, etc.) into equivalents of written words.
[0260] During linguistic analysis, the pre-processing component 920 can analyze the language in the normalized text to generate a sequence of speech units corresponding to the input text. This process can be referred to as grapheme to phoneme conversion. The speech unit includes a symbolic representation of a sound unit that will eventually be combined by the system and output as speech. In order to perform speech synthesis, various sound units can be used to divide the text. In some implementations, the TTS model 980 can process speech based on phonemes (single sounds), semiphones, diphones (the second half of a phoneme combined with the first half of an adjacent phoneme), diphones (two consecutive phonemes), syllables, words, phrases, sentences, or other units. Each word can be mapped to one or more speech units. Such mapping can be performed using a language dictionary stored by the system (e.g., in a storage component). The linguistic analysis performed by the pre-processing component 920 can also identify different grammatical components, such as prefixes, suffixes, phrases, punctuation, syntactic boundaries, etc. Such grammatical components can be used by the TTS system 480 to produce an audio waveform output that sounds natural. The language dictionary may also include letter-to-sound rules and other tools that may be used to pronounce previously unrecognized words or letter combinations that may be encountered by the TTS system 480. Generally speaking, the more information included in the language dictionary, the higher the quality of the speech output.
[0261] The output of the preprocessing component 920 may be a symbolic language representation that may include a sequence of speech units. In some implementations, the sequence of speech units may be annotated with prosodic features. In some implementations, the prosody may be applied in part or in whole by the TTS model 980. This symbolic language representation may be sent to the TTS model 980 for conversion into audio data (e.g., in the form of a mel-spectrogram or other frequency content data format).
[0262] The TTS system 480 may retrieve one or more previously trained and / or configured TTS models 980 from the voice profile storage 985. The TTS model 980 may be, for example, a neural network architecture, which may be described as interconnected artificial neurons or "cells" interconnected in layers and / or blocks. In general, the neural network model architecture may be broadly described by hyperparameters that describe the number of layers and / or blocks, how many units each layer and / or block contains, the activation functions they implement, how they are interconnected, and the like. The neural network model includes trainable parameters (e.g., "weights") that indicate how much weight (e.g., in the form of arithmetic multipliers) a unit should give to a particular input when generating an output. In some implementations, the neural network model may include other features such as self-attention mechanisms that may determine certain parameters at runtime based on the inputs rather than, for example, based on loss calculations during training. Various data describing a particular TTS model 980 may be stored in the voice profile storage 985. The TTS model 980 may represent a particular speaker identity and may be adjusted based on speaking style, emotion, and the like. In some implementations, a particular speaker identity may be associated with more than one TTS model 980; for example, with different models representing different speaking styles, languages, emotions, etc. In some implementations, a particular TTS model 980 may be associated with more than one speaker identity; that is, capable of producing synthesized speech that reproduces the speech characteristics of more than one character. Thus, a first TTS model 980a may be used to create synthesized speech for a first speech processing system 120a, while a different second TTS model 980b may be used to create synthesized speech for a second speech processing system 120b. In some cases, the TTS model 980 may generate the desired speech characteristics based on conditional data received or determined from the text data 915 and / or other input data 925. For example, the synthesized speech of the first speech processing system 120a may be different from the synthesized speech of the second speech processing system 120b.
[0263] The TTS system 480 can retrieve a TTS model 980 from the speech profile storage 985 and use it to process the input to generate synthesized speech based on the instructions received with the text data 915 and / or other input data 925. The TTS system 480 can provide any relevant conditional tags to the TTS model 980 to generate synthesized speech with desired speech characteristics. The TTS model 980 can generate spectrogram data 945 (e.g., frequency content data) representing the synthesized speech and send the spectrogram data to the vocoder 990 for conversion to an audio signal.
[0264] The TTS system 480 may generate other output data 955. The other output data 955 may include, for example, instructions or instructions for processing and / or outputting synthesized speech. For example, the text data 915 and / or other input data 925 may be received together with metadata (such as SSML tags) indicating that selected portions of the text data 915 should be louder or quieter. Thus, the other output data 955 may include a volume mark that instructs the vocoder 990 to increase or decrease the amplitude of the output speech audio data 995 at a time corresponding to the selected portion of the text data 915. Additionally or alternatively, the volume mark may instruct the playback device to increase or decrease the volume of the synthesized speech from the current volume level of the device, or to reduce the volume of other media output by the device (e.g., to convey an emergency message).
[0265] The vocoder 990 can convert the spectrogram data 945 generated by the TTS model 980 into an audio signal (e.g., an analog or digital time domain waveform) suitable for amplification and output as audio. The vocoder 990 can be, for example, a general neural vocoder based on a parallel WaveNet or related model. The vocoder 990 can take as input audio data in the form of a Mel spectrogram with, for example, 80 coefficients and a frequency range from 50 Hz to 12 kHz. The speech audio data 995 can be a time domain audio format (e.g., pulse code modulation (PCM), waveform audio format (WAV), μ-law, etc.), which can be easily converted to an analog signal for amplification and output by a speaker (such as speaker 112). The speech audio data 995 can be composed of, for example, 8-bit, 16-bit, or 24-bit audio with a sampling rate of 16 kHz, 24 kHz, 44.1 kHz, etc. In some implementations, other bits and / or sampling rates can be used.
[0266] Various machine learning techniques can be used to train and operate models to perform various steps described herein, such as user identification, emotion detection, image processing, dialogue management, etc. The model can be trained and operated according to various machine learning techniques. Such techniques may include, for example, neural networks (such as deep neural networks and / or recursive neural networks), inference engines, trained classifiers, etc. Examples of trained classifiers include support vector machines (SVMs), neural networks, decision trees, AdaBoost (abbreviation for "adaptive boost") and random forests combined with decision trees. Taking SVM as an example, SVM is a supervised learning model with an associated learning algorithm that analyzes data and identifies patterns in data and is commonly used in classification and regression analysis. Given a set of training instances, each training instance is marked as belonging to one of two categories, and the SVM training algorithm constructs a model that assigns new instances to one category or another category, making it a non-probabilistic binary linear classifier. More complex SVM models can be constructed with training sets that identify more than two categories, where SVM determines which category is most similar to the input data. The SVM model can be mapped so that instances of each category are separated by a clear gap. New instances are then mapped into the same space and predicted to which class they belong based on which side of the gap the instance lies. The classifier gives a "score" that indicates which class the data most closely matches. The score provides an indication of how well the data matches the class.
[0267] In order to apply machine learning techniques, the machine learning process itself needs to be trained. Training a machine learning component (such as one of the first model or the second model in this example) requires establishing "ground truth" for the training examples. In machine learning, the term "ground truth" refers to the accuracy of the classification of the training set of supervised learning techniques. Various techniques can be used to train the model, including backpropagation, statistical learning, supervised learning, semi-supervised learning, stochastic learning, or other known techniques.
[0268] Fig.10 is a block diagram conceptually illustrating a device 110 that may be used with the system. Fig.11 1 is a block diagram conceptually illustrating example components of a remote device, such as a natural language command processing system 120 and a skill processing component 125 that can assist in ASR processing, NLU processing, etc. The system (120 / 125) may include one or more servers. As used herein, "server" may refer to a traditional server as understood in a server / client computing structure, but may also refer to many different computing components that can assist in the operations discussed herein. For example, a server may include one or more physical computing components (such as rack-mounted servers) that are physically and / or connected to other devices / components via a network and are capable of performing computing operations. The server may also include one or more virtual machines that simulate a computer system and run on one device or across multiple devices. The server may also include other combinations of hardware, software, firmware, etc. to perform the operations discussed herein. The server may be configured to operate using one or more of a client-server model, a computer bureau model, grid computing technology, fog computing technology, a mainframe computer technology, a utility computing technology, a peer-to-peer model, a sandbox technology, or other computing technologies.
[0269] While the device 110 may be running locally to the user (e.g., within the same environment so that the device can receive the user's input and playback output), the server / system 120 may be located remotely from the device 110 because its operation may not require proximity to the user. The server / system 120 may be located in a completely different location from the device 110 (e.g., as part of a cloud computing system, etc.), or may be located in the same environment as the device 110 but physically separated (e.g., a home server or similar device located in the user's home or business but possibly in a closet, basement, attic, etc.). The processing system 120 may also be a version of the user device 110 that includes different (e.g., more) processing capabilities than other user devices 110 in the home / office. One benefit of the server / system 120 being in the user's home / business is that the data used to process the commands / return responses can be saved in the user's home, thereby reducing potential privacy perception issues.
[0270] A plurality of systems (120 / 125) may be included in the overall system 100 of the present disclosure, such as one or more natural language processing systems 120 for performing ASR processing, one or more natural language processing systems 120 for performing NLU processing, one or more skill systems 125, etc. In operation, each of these systems may include computer-readable and computer-executable instructions resident on a respective device (120 / 125), as will be discussed further below.
[0271] Each of these devices (110 / 120 / 125) may include one or more controllers / processors (1004 / 1104), which may each include a central processing unit (CPU) for processing data and computer-readable instructions and a memory (1006 / 1106) for storing data and instructions for the corresponding device. The memory (1006 / 1106) may individually include volatile random access memory (RAM), non-volatile read-only memory (ROM), non-volatile magnetoresistive memory (MRAM), and / or other types of memory. Each device (110 / 120 / 125) may also include a data storage component (1008 / 1108) for storing data and controller / processor executable instructions. Each data storage component (1008 / 1108) may individually include one or more non-volatile storage device types, such as magnetic storage devices, optical storage devices, solid-state storage devices, etc. Each device (110 / 120 / 125) may also be connected to a removable or external non-volatile memory and / or storage device (such as a removable memory card, memory key drive, network storage device, etc.) through a corresponding input / output device interface (1002 / 1102).
[0272] Computer instructions for operating each device (110 / 120 / 125) and its various components may be executed by the controller / processor (1004 / 1104) of the respective device, using the memory (1006 / 1106) as temporary "working" storage during operation. The computer instructions for a device may be stored in a non-volatile memory (1006 / 1106), storage device (1008 / 1108), or external device in a non-transitory manner. Alternatively, in addition to or in lieu of software, some or all of the executable instructions may be embedded in hardware or firmware on the respective device.
[0273] Each device / system (110 / 120 / 125) includes an input / output device interface (1002 / 1102). Various components may be connected via the input / output device interface (1002 / 1102), as will be discussed further below. In addition, each device (110 / 120 / 125) may include an address / data bus (1024 / 1124) for transferring data between components of the respective device. In addition to (or in lieu of) connecting to other components across the bus (1024 / 1124), each component within the device (110 / 120 / 125) may also be directly connected to other components.
[0274] refer to Fig.10 , the device 110 may include an input / output device interface 1002 connected to various components (such as an audio output component, such as a speaker 112, a wired headset or a wireless headset (not shown), or other components capable of outputting audio). The device 110 may also include an audio capture component. The audio capture component may be, for example, a microphone 114 or a microphone array, a wired headset or a wireless headset (not shown), etc. If a microphone array is included, the approximate distance to the origin point of the sound can be determined by acoustic positioning based on the time and amplitude differences between the sounds captured by different microphones in the array. The device 110 may additionally include a display 1016 for displaying content. The device 110 may also include a camera 1018.
[0275] Through antenna 1022, input / output device interface 1002 may be connected to one or more networks 199 via a wireless local area network (WLAN) (such as Wi-Fi) radio, Bluetooth, and / or a wireless network radio, such as a radio capable of communicating with a wireless communication network such as a Long Term Evolution (LTE) network, a WiMAX network, a 3G network, a 4G network, a 5G network, etc. Wired connections such as Ethernet may also be supported. Through network 199, the system may be distributed in a networked environment. The I / O device interface (1002 / 1102) may also include a communication component that allows data to be exchanged between devices (such as different physical servers or other components in a group of servers).
[0276] Components of the device 110, natural language command processing system 120, or skill processing component 125 may include their own dedicated processors, memories, and / or storage. Alternatively, one or more of the components of the device 110, natural language command processing system 120, or skill processing component 125 may utilize the I / O interface (1002 / 1102), processor (1004 / 1104), memory (1006 / 1106), and / or storage (1008 / 1108) of the device 110, natural language command processing system 120, or skill processing component 125, respectively. Thus, for the various components discussed herein, the ASR component XXA50 may have its own I / O interface, processor, memory, and / or storage; the NLU component XXA60 may have its own I / O interface, processor, memory, and / or storage; and so on.
[0277] As described above, multiple devices may be employed in a single system. In such a multi-device system, each device may include different components for performing different aspects of system processing. Multiple devices may include overlapping components. As described herein, the components of the device 110, the natural language command processing system 120, and the skill processing component 125 are illustrative and may be located as independent devices or may be included in whole or in part as components of a larger device or system. It will be appreciated that there may be many components on the system 120 and / or device 110. Unless otherwise expressly stated, the system version of such components may operate similarly to the device version of such components, so the description of one version (e.g., the system version or the local version) applies to the description of another version (e.g., the local version or the system version), and vice versa.
[0278] like Fig.12 As shown, multiple devices (110a to 110n, 120, 125) may include components of the system, and the devices may be connected via a network 199. The network 199 may include a local or private network, or may include a wide area network, such as the Internet. The device may be connected to the network 199 via a wired connection or a wireless connection. For example, a voice detection device 110a, a smart phone 110b, a smart watch 110c, a tablet computer 110d, a vehicle 110e, a language detection device with a display 110f, a display / smart TV 110g, a washing machine / dryer 110h, a refrigerator 110i, a microwave oven 110j, a headset 110b / 110n, etc. may be connected to the network 199 via a wireless service provider, via a Wi-Fi or cellular network connection, etc. Other devices are included as support devices for connecting to the network, such as a natural language command processing system 120, a skill processing system 125, and / or other devices. The support device may be connected to the network 199 via a wired connection or a wireless connection. A networked device may capture audio using one or more built-in or connected microphones or other audio capture devices, where processing is performed by an ASR component, NLU component, or other component (such as ASR component 450, NLU component 460, etc. of natural language command processing system 120) of the same device or another device connected via network 199.
[0279] The concepts disclosed herein may be applied in many different devices and computer systems including, for example, general purpose computing systems, speech processing systems, and distributed computing environments.
[0280] The contents of this article can also be understood in light of the following terms.
[0281] 1. A computer-implemented method comprising: capturing a first utterance by a first device, the first device being configured to operate with a plurality of speech processing systems including a first speech processing system and a second speech processing system; processing first audio data representing the first utterance to determine that the first utterance includes a first wake-up word that invokes the first speech processing system; sending the first audio data from a first component of the first device to the first speech processing system, the first component being configured to operate with respect to the first speech processing system; receiving first output data in response to the first audio data from the first speech processing system by the first component; sending the first output data from the first component to a second component of the first device, the second component being configured to operate with respect to the plurality of speech processing systems; processing the first output data by the second component to determine a first command to terminate a process involving the first device; processing state data corresponding to the first device by the second component to determine that the process is started in response to a command to invoke the second speech processing system; sending an instruction to terminate the process from the second component to a third component of the first device, the third component being configured to operate with respect to the second speech processing system; and terminating the process in response to the instruction.
[0282] 2. A computer-implemented method as described in clause 1, the method also includes, before capturing the first utterance: capturing a second utterance through the first device; processing second audio data representing the second utterance to determine that the second utterance includes a second wake-up word that calls the second voice processing system; sending the second audio data from the third component to the second voice processing system; receiving second output data from the second voice processing system indicating a second command to start the process; starting the process using the first device; and determining the status data through the second component, wherein the status data indicates that the process is started in response to a command to call the second voice processing system.
[0283] 3. The computer-implemented method of clause 2, further comprising: sending the state data from the second component to the first speech processing system before receiving the first output data.
[0284] 4. A computer-implemented method as described in clause 1, 2 or 3, the method also includes: determining, by the second component, that the process is active; and after stopping the process, sending a second indication from the third component to the second speech processing system that the process has been stopped by the first device.
[0285] 5. A computer-implemented method, comprising: capturing a first utterance through a first device, the first device being configured to operate with a plurality of speech processing systems including a first speech processing system and a second speech processing system; determining that the first utterance corresponds to a call to the first speech processing system; sending first audio data representing the first utterance from the first device to the first speech processing system; sending first status data representing an active first device process of the first device from a component of the first device to the first speech processing system, the component being configured to operate with respect to the plurality of speech processing systems; receiving, through the component of the first device, a first command to control the active first device process using the first device; performing a first operation to control the active first device process in response to the first command; and configuring updated status data corresponding to the first device, the updated status data reflecting the execution of the first operation.
[0286] 6. A computer-implemented method as described in clause 5, the method further comprising, before capturing the first utterance: capturing a second utterance through the first device; determining that the second utterance corresponds to a call to the second speech processing system; sending second audio data representing the second utterance from the first device to the second speech processing system; receiving, through the component of the first device, a second command to start the active first device process using the first device; and in response to the second command, performing a second operation to start the active first device process.
[0287] 7. The computer-implemented method of clause 6, further comprising, after receiving the first command: sending, from the first device to the second speech processing system, an indication corresponding to control of the active first device process.
[0288] 8. A computer-implemented method as described in clause 6 or 7, the method further comprising: before receiving the second command: determining a first user corresponding to the second utterance, determining that the first user corresponds to a first profile, and sending an indication of the first profile to the second speech processing system; and after receiving the second command and before receiving the first command: determining a second user corresponding to the first utterance, determining that the second user corresponds to the first profile, and sending an indication of the first profile to the first speech processing system.
[0289] 9. A computer-implemented method as described in clauses 6, 7 or 8, wherein: sending the second audio data to the second speech processing system uses a second component of the first device corresponding to the second speech processing system; and sending the first audio data to the first speech processing system uses a third component of the first device corresponding to the first speech processing system.
[0290] 10. The computer-implemented method of clause 5, 6, 7, 8, or 9, further comprising sending, from the first device to the first speech processing system, second data corresponding to at least one active hardware component of the first device.
[0291] 11. A computer-implemented method as described in clauses 5, 6, 7, 8, 9 or 10, wherein the first state data further corresponds to a second device process that the first device is capable of controlling, and the first state data indicates that the active first device process is active and the second device process is inactive.
[0292] 12. A computer-implemented method of clause 11, further comprising: determining, by the component, first priority data corresponding to the active first device process; determining, by the component, second priority data corresponding to the second device process; and sending the first priority data and the second priority data to the first voice processing system before receiving the first command.
[0293] 13. A device comprising: at least one microphone; a first component configured to operate with respect to multiple speech processing systems including a first speech processing system and a second speech processing system; at least one processor; and at least one memory comprising instructions that, when executed by the at least one processor, cause the device to perform the following operations: capture a first utterance using the at least one microphone; determine that the first utterance corresponds to a call to the first speech processing system; send first audio data representing the first utterance to the first speech processing system; send first status data representing an active first device process from the first component to the first speech processing system; receive a first command to control the active first device process through the first component; perform a first operation to control the active first device process in response to the first command; and configure updated status data, the updated status data reflecting the execution of the first operation.
[0294] 14. An apparatus as described in clause 13, wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the apparatus to perform the following operations: before capturing the first utterance: capturing a second utterance; determining that the second utterance corresponds to a call to the second speech processing system; sending second audio data representing the second utterance to the second speech processing system; receiving, through the first component, a second command to start the active first device process; and in response to the second command, performing a second operation to start the active first device process.
[0295] 15. An apparatus as described in clause 14, wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the apparatus to perform the following operations: after receiving the first command: sending an indication corresponding to control of the active first device process to the second voice processing system.
[0296] 16. An apparatus as described in clause 14 or 15, wherein the at least one memory also includes instructions that, when executed by the at least one processor, further cause the apparatus to perform the following operations: before receiving the second command: determining a first user corresponding to the second utterance, determining that the first user corresponds to a first profile, and sending an indication of the first profile to the second speech processing system; and after receiving the second command and before receiving the first command: determining a second user corresponding to the first utterance, determining that the second user corresponds to the first profile, and sending an indication of the first profile to the first speech processing system.
[0297] 17. An apparatus as described in clauses 14, 15 or 16, wherein: sending the second audio data to the second speech processing system uses a second component of the apparatus corresponding to the second speech processing system; and sending the first audio data to the first speech processing system uses a third component of the apparatus corresponding to the first speech processing system.
[0298] 18. An apparatus as described in clauses 13, 14, 15, 16 or 17, wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the apparatus to perform the following operations: send second data corresponding to at least one active hardware component of the apparatus to the first speech processing system.
[0299] 19. An apparatus as described in clause 13, 14, 15, 16, 17 or 18, wherein the first state data further corresponds to a second device process that the apparatus is capable of controlling, and the first state data indicates that the active first device process is active and the second device process is inactive.
[0300] 20. An apparatus as described in clause 19, wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the apparatus to perform the following operations: determine, through the first component, first priority data corresponding to the active first device process; determine, through the first component, second priority data corresponding to the second device process; and send the first priority data and the second priority data to the first voice processing system before receiving the first command.
[0301] The above aspects of the present disclosure are intended to be illustrative. They are selected to explain the principles and applications of the present disclosure and are not intended to be exhaustive or to limit the present disclosure. Those skilled in the art may make many modifications and variations to the disclosed aspects. Those of ordinary skill in the field of computer and speech processing should recognize that the components and process steps described herein may be interchangeable with other components or steps, or a combination of components or steps, and still achieve the benefits and advantages of the present disclosure. However, those skilled in the art should understand that the present disclosure may be practiced without some or all of the specific details disclosed herein. In addition, unless expressly stated to the contrary, features / operations / components, etc. from one embodiment discussed herein may be combined with features / operations / components, etc. from another embodiment discussed herein.
[0302] Various aspects of the disclosed system may be implemented as a computer method or article of manufacture, such as a memory device or a non-transitory computer-readable storage medium. The computer-readable storage medium may be computer-readable and may include instructions for causing a computer or other device to perform the processes described in the present disclosure. The computer-readable storage medium may be implemented by volatile computer memory, non-volatile computer memory, a hard drive, a solid-state memory, a flash drive, a removable disk, and / or other media. In addition, components of the system may be implemented as firmware or hardware.
[0303] Unless specifically stated otherwise, or otherwise understood within the context as used, conditional language used herein (such as, among other things, "can," "might," "may," "could," "may," "for example," etc.) is generally intended to convey that certain embodiments include certain features, elements, or steps, while other embodiments do not include certain features, elements, or steps. Thus, such conditional language is generally not intended to imply that one or more embodiments require features, elements, and / or steps in any way, or that one or more embodiments necessarily include logic for deciding whether to include or perform these features, elements, and / or steps in any particular embodiment with or without other input or prompting. The terms "comprising," "including," "having," and the like are synonymous and are used inclusively in an open manner and do not exclude additional elements, features, actions, operations, and the like. Moreover, the term "or" is used in its inclusive sense (and not in an exclusive sense) such that when used, for example, to connect elements of a list, the term "or" means one, some, or all of the elements in the list.
[0304] Unless specifically stated otherwise, disjunctive language, such as the phrase "at least one of X, Y, or Z," is understood in context as generally used to present that an item, term, etc. can be X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is generally not intended to, and should not, imply that certain embodiments require the presence of at least one of X, at least one of Y, or at least one of Z, respectively.
[0305] As used in this disclosure, unless specifically stated otherwise, the terms "a" or "an" may include one or more items. In addition, unless specifically stated otherwise, the phrase "based on" is intended to mean "based at least in part on."< / dialog>
Claims
1. A computer-implemented method, the method comprising: capturing a first utterance by a first device, the first device being configured to operate with a plurality of speech processing systems including a first speech processing system and a second speech processing system; determining that the first utterance corresponds to a call to the first speech processing system; sending, from the first device to the first speech processing system, first audio data representing the first utterance; sending, from a component of the first device to the first speech processing system, first status data representing an active first device process of the first device, the component being configured to operate with respect to the plurality of speech processing systems; receiving, by the component of the first device, a first command to control the active first device process using the first device; In response to the first command, performing a first operation to control the active first device process; as well as Updated state data corresponding to the first device is configured, and the updated state data reflects the performance of the first operation.
2. The computer-implemented method of claim 1 , further comprising, before capturing the first utterance: capturing a second utterance by the first device; determining that the second utterance corresponds to a call to the second speech processing system; sending, from the first device to the second speech processing system, second audio data representing the second utterance; receiving, by the component of the first device, a second command to initiate the active first device process using the first device; and In response to the second command, a second operation is performed to start the active first device process.
3. The computer-implemented method of claim 2, further comprising, after receiving the first command: An indication corresponding to control of the active first device process is sent from the first device to the second speech processing system.
4. The computer-implemented method of claim 2 or 3, further comprising: Before receiving the second command: determining a first user corresponding to the second utterance, determining that the first user corresponds to a first profile, and sending an indication of the first configuration file to the second speech processing system; and After receiving the second command and before receiving the first command: determining a second user corresponding to the first utterance, determining that the second user corresponds to the first profile, and An indication of the first configuration file is sent to the first speech processing system.
5. A computer-implemented method according to claim 2, 3 or 4, wherein: sending the second audio data to the second speech processing system using a second component of the first device corresponding to the second speech processing system; and Sending the first audio data to the first speech processing system uses a third component of the first device corresponding to the first speech processing system.
6. The computer-implemented method of claim 1, 2, 3, 4, or 5, further comprising: Second data corresponding to at least one active hardware component of the first device is sent from the first device to the first speech processing system.
7. A computer-implemented method according to claim 1, 2, 3, 4, 5 or 6, wherein the first state data further corresponds to a second device process that the first device can control, and the first state data indicates that the active first device process is active and the second device process is inactive.
8. The computer-implemented method of claim 7, further comprising: determining, by the component, first priority data corresponding to the active first device process; determining, by the component, second priority data corresponding to the second device process; as well as Prior to receiving the first command, the first priority data and the second priority data are sent to the first voice processing system.
9. An apparatus comprising: at least one microphone; a first component configured to operate with respect to a plurality of speech processing systems including a first speech processing system and a second speech processing system; at least one processor; as well as at least one memory comprising instructions that, when executed by the at least one processor, cause the apparatus to: capturing a first utterance using the at least one microphone; determining that the first utterance corresponds to a call to the first speech processing system; sending first audio data representing the first utterance to the first speech processing system; sending, from the first component to the first speech processing system, first status data representing an active first device process; receiving, by the first component, a first command to control the active first device process; In response to the first command, performing a first operation to control the active first device process; and Updated state data is configured, the updated state data reflecting the execution of the first operation.
10. The apparatus of claim 9, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the apparatus to perform the following operations before capturing the first utterance: Capturing the second discourse; determining that the second utterance corresponds to a call to the second speech processing system; sending second audio data representing the second utterance to the second speech processing system; receiving, by the first component, a second command to start the active first device process; and In response to the second command, a second operation is performed to start the active first device process.
11. The apparatus of claim 10, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the apparatus to perform the following operations after receiving the first command: An indication corresponding to control of the active first device process is sent to the second speech processing system.
12. The apparatus of claim 10 or 11, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the apparatus to: Before receiving the second command: determining a first user corresponding to the second utterance, determining that the first user corresponds to a first profile, and sending an indication of the first configuration file to the second speech processing system; and After receiving the second command and before receiving the first command: determining a second user corresponding to the first utterance, determining that the second user corresponds to the first profile, and An indication of the first configuration file is sent to the first speech processing system.
13. The device according to claim 10, 11 or 12, wherein: sending the second audio data to the second speech processing system using a second component of the device corresponding to the second speech processing system; and Sending the first audio data to the first speech processing system uses a third component of the device corresponding to the first speech processing system.
14. The apparatus of claim 9, 10, 11, 12 or 13, wherein the at least one memory further comprises instructions that when executed by the at least one processor further cause the apparatus to: Second data corresponding to at least one active hardware component of the device is sent to the first speech processing system.
15. An apparatus according to claim 9, 10, 11, 12, 13 or 14, wherein the first state data further corresponds to a second device process that the apparatus is capable of controlling, and the active first state data indicates that the active first device process is active and the second device process is inactive.
Citation Information
Cited By
Play generation method and related device
CN121901382A