System and method for routing content to associated output devices

By managing the association between voice-activated devices and output devices through a cloud-based backend system, and using STT and NLU to process user commands, the content routing problem in multi-device collaborative work is solved, realizing intelligent content routing and user-friendly interaction between devices.

CN116631391BActive Publication Date: 2026-03-17AMAZON TECH INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2017-06-26
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing electronic devices lack effective methods for handling collaborative work and content routing among multiple devices when receiving and outputting content, especially for information transfer and control between voice-activated devices and other output devices.

Method used

The system manages the association between voice-activated electronic devices, user accounts, and output devices through a cloud-based backend system. It utilizes speech-to-text (STT) and natural language understanding (NLU) functions to process user commands, determine the intent of the content, and route it to the appropriate output device, while taking into account device status and user preferences.

Benefits of technology

It enables collaborative work between voice-activated devices and output devices, improving the efficiency of content routing and user experience, ensuring accurate output of content on appropriate devices, and enhancing the convenience of user interaction and the intelligence of devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116631391B_ABST
    Figure CN116631391B_ABST
Patent Text Reader

Abstract

This document provides devices and methods for routing content. In some embodiments, a method for routing content includes: receiving audio data representing a command from a first electronic device; determining content associated with the command; sending responsive audio data to the first electronic device; and sending an instruction for outputting the content associated with the command to a second electronic device. In some embodiments, a method for routing content includes: determining a state of the second electronic device and, based on the state of the second electronic device, sending an instruction for outputting the content to a selected electronic device between the first electronic device and the second electronic device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Divisional application statement

[0002] This application is a divisional application of Chinese Patent Application No. 201780052537.4, which was filed on June 26, 2017 (PCT International Application No. PCT / US2017 / 039221) and entered the Chinese national phase on February 26, 2019. Background Technology

[0003] Voice-activated electronic devices, such as those for mobile devices, are relatively new but are becoming increasingly common. Individuals can interact with their electronic devices to perform a variety of basic functions, from making phone calls to streaming content. This article discusses improvements to electronic devices and the back-end systems that work with them. Attached Figure Description

[0004] Figure 1 This is an illustrative diagram of a system for routing content to associated output electronic devices according to various implementation schemes;

[0005] Figure 2 This is an example diagram of a system for routing content based on the state of associated output electronic devices, according to various implementation schemes;

[0006] Figure 3 It is based on various implementation plans. Figure 1 An example diagram of the system architecture;

[0007] Figure 4 This is another example of a table that includes categories of content types based on various implementation schemes;

[0008] Figure 5 This is an example diagram illustrating instances of resolving ambiguous requests for content based on various implementation schemes;

[0009] Figure 6 This is an illustrative diagram of a system for associating an output electronic device with a voice-activated electronic device according to various implementation schemes;

[0010] Figure 7 This is an illustrative diagram showing the linking of two exemplary devices according to various implementation schemes;

[0011] Figure 8 It is an illustrative flowchart of the process of sending content to associated devices according to various implementation schemes;

[0012] Figure 9A It is an exemplary flowchart of the process for routing content based on content type, according to various implementation schemes;

[0013] Figure 9B It is based on the succession of various implementation plans. Figure 9A An illustrative flowchart of the process, in which content is routed to associated devices based on content;

[0014] Figure 9C It is based on the succession of various implementation plans. Figure 9A An illustrative flowchart of the process in which content is routed to electronic devices based on content;

[0015] Figure 10 It is an illustrative flowchart of the process for receiving a request to change the output device according to various implementation schemes;

[0016] Figure 11A It is an illustrative flowchart of the process for routing content based on the state of associated devices, according to various implementation schemes;

[0017] Figure 11B It is based on the succession of various implementation plans. Figure 11A An illustrative flowchart of the process, where the state of the associated device is ready;

[0018] Figure 11C It is based on the succession of various implementation plans. Figure 11A An illustrative flowchart of the process, where the status of the associated devices is available; and

[0019] Figure 11D It is based on the succession of various implementation plans. Figure 11A An illustrative flowchart of the process, in which the state of the associated device is unavailable. Detailed Implementation

[0020] This disclosure (as stated below) generally relates to various embodiments of methods and apparatus relating to receiving commands (such as requests for content) at one device and outputting the requested content through another device.

[0021] In some implementations, an individual can speak to the voice-activated electronic device, such as requesting to listen to a weather forecast. The voice-activated electronic device can use one or more microphones or transducers to capture the audio signal of the spoken command, which can be converted into audio data representing the spoken words. The voice-activated electronic device can then send the audio data to a back-end system. In some implementations, the voice-activated electronic device can be associated with a user account, which is also associated with an output electronic device, such as a television and / or a streaming media device connected to the television. The association between the user account and both the voice-activated electronic device and the output electronic device can be stored in a cloud-based back-end system. The back-end system can identify this association by first identifying a device identifier associated with the voice-activated electronic device. The device identifier can then be used to determine the user account associated with that device identifier. Once the cloud-based back-end system identifies the user account associated with the voice-activated electronic device, it can then identify all products associated with the identified user account. In this example, the cloud-based back-end system can identify a television that is also associated with the identified user account.

[0022] Once the cloud-based backend system identifies a product associated with the voice-activated electronic device, or upon identifying such a product, the cloud-based backend system can then convert the audio data representing the utterance into text data by performing speech-to-text (STT) functionality on the audio data. Once the audio data is converted into text data, the cloud-based backend system will then determine the intent of the utterance by performing natural language understanding (NLU) functionality on the text data representing the audio data. NLU will determine the intent and meaning of the text data. For example, the cloud-based backend system may determine that the utterance includes a request to listen to a weather forecast on a target device. The NLU can then determine that the target device is an output device associated with the user account. Once the cloud-based backend system understands what the utterance requests, it will search for an appropriate response. For example, in response to the request for a weather forecast, the cloud-based backend system can find text data stating the weather forecast. Additionally, the backend system can determine that visual information in response to the utterance is available. In some implementations, once visual information in response to the utterance is found, the cloud-based backend system can determine that the target device is capable of displaying the visual information. Furthermore, if the user account is associated with a television, the backend system can further determine that, because the voice-activated electronic device is associated with the television, the response should include two responses. A first response may be sent to the voice-activated electronic device. A second response may be sent to the television. In some implementations, the second response may include both audio and visual responses.

[0023] In some implementations, after the cloud-based backend system determines that the response will be sent to both the voice-activated electronic system and the television, the cloud-based backend system can receive text data representing a response to the utterance for the voice-activated electronic system. The text data can be converted into audio data that performs text-to-speech (TTS) functionality on the responsive text data. After creating the audio data by performing TTS, the audio data can be sent to the voice-activated electronic system. Once received by the voice-activated electronic system, the audio data can be played by one or more speakers on the voice-activated electronic system. For example, the voice-activated electronic system could state: “The Seattle weather forecast is being displayed on your television.”

[0024] The cloud-based backend system can also receive video data representing responsive visual information. Before sending the video data to the television, the cloud-based backend system can recognize that second audio data should be generated to be sent to the television along with the video data. The cloud-based backend system can then receive text data for the television representing a response to the utterance. Similar to the text data for the voice-activated electronic device, the text data for the television can be converted into audio data that performs TTS functionality on the responsive text data for the television. After creating the audio data by performing TTS, the audio data and the video data can be sent to the television. Once received by the television, the audio data and the video data can be played by the television. For example, the television can state, "This is the Seattle forecast." In some embodiments, the backend system can also send the requested content to the output electronic device.

[0025] In some implementations, an individual may utter utterances to the voice-activated electronic device that could have two different meanings. For example, an individual might say, “Alexa, play Footloose.” In this case, “play Footloose” could have two different meanings. For example, “play” could refer to the action corresponding to both a movie and a song, among other content types. The cloud-based backend system will convert the audio data of the spoken utterance “Alexa, play Footloose” into text data by performing STT functionality on the audio data. Once the audio data is converted into text data, the cloud-based backend system will attempt to understand the intent of the utterance by applying NLU functionality to the text data representing the audio data. The NLU will attempt to understand the intent of the text data. In some implementations, the NLU will receive two confidence scores from two separate domains. Each of these confidence scores may exceed a predetermined threshold. If the NLU receives two confidence scores that exceed the predetermined threshold, then the cloud-based backend system can determine that more information is needed to send the correct response.

[0026] Once it is determined that more information is needed, the cloud-based backend system can receive query text data representing an intent question, which asks whether the utterance seeks a response from a first domain or a second domain. The cloud-based backend system can then convert the query text data into audio data by performing TTS functionality on the query text data. Before sending the audio data representing the query, the cloud-based backend system can generate an answer instruction for the voice-activated electronic device. In some embodiments, the answer instruction may instruct the voice-activated electronic device to listen for a response after playing the query audio. After generating the answer instruction, the cloud-based backend system can send the audio data representing the query to the voice-activated electronic device. The voice-activated electronic device can play the audio data on one or more of its speakers. For example, the voice-activated electronic device may play "Do you want to play the movie Footloose or the song Footloose?" Then, the cloud-based backend system can send the answer instruction to the voice-activated electronic device, thereby instructing the voice-activated electronic device to listen for a response to the query and send audio data representing the response to the query to the cloud-based backend system. In some implementations, once the voice-activated electronic device has played the query, it can send a response to the cloud-based backend system. This response is received by the cloud-based backend system as audio data. The audio data is converted into a text file by STT and analyzed by NLU. Based on the analyzed confirmation response, the cloud-based backend system is able to determine the intent of the original utterance.

[0027] For example, if the response is "movie," the cloud-based backend system can verify that the device associated with the user account and also with the voice electronic device can stream movies. If the device can stream movies, the cloud-based backend system can verify whether the user account can access movies. If the user account can access movies, the cloud-based backend system will generate a URL that allows the user to view the streaming movie. Once the URL is generated, the cloud-based backend system can send the URL to the device, causing the device to begin streaming the movie.

[0028] If the response is "song," the cloud-based backend system determines that a song should be played on the voice-activated electronic device. The cloud-based backend system can then generate a URL that allows the voice-activated electronic device to stream the song. Once the URL has been generated, the cloud-based backend system can send the generated URL to the voice-activated electronic device, causing the song to play on at least one speaker of the voice-activated electronic device. Once the song begins playing on the voice-activated electronic device, the individual may state another phrase to the voice-activated electronic device. This phrase might be a request to play the same song, but on a television. For example, the individual might state, "Play the song Footloose on the television." After the cloud-based backend system performs STT functionality on the received audio and NLU functionality on the text data representing the received audio data, the cloud-based backend system can recognize that the individual wants to play the same song on the television.

[0029] After recognizing that the utterance is a request to play the same song on the television, the cloud-based backend system can generate a stop command for the voice-activated user device. The cloud-based backend system then sends a stop command to the voice-activated electronic device. The generated command can then be sent to the voice-activated electronic device, causing it to stop streaming the song. After the voice-activated device stops streaming the song, the cloud-based backend system receives another URL that allows the television to stream the song. The song's URL can then be sent to the television so that playback can begin when the song is stopped on the voice-activated electronic device. In some implementations, the cloud-based backend system can generate a notification to the user that the song will be played on the television, and then send it to the voice-activated electronic device for playback. For example, the voice-activated electronic device might play "The song Footloose will be played on your television."

[0030] In some embodiments, the output electronics may include a streaming media device connected to a peripheral output device. The streaming media device may also control the peripheral output device. For example, in some embodiments, the peripheral output electronics may be a television connected to the streaming media device. The peripheral output device may not be directly connected to the back-end system. In other words, the back-end system may only be able to communicate with or control the peripheral output device through the streaming media device. In embodiments where the output electronics includes the peripheral output device, instructions sent by the output back-end system to the output electronics to output content can cause the streaming media device to control the peripheral electronics to output the content.

[0031] In some implementations, the cloud-based backend system can determine that associations are stored on the cloud-based backend system. This association may include stored input devices, output devices, and content preferences. In some implementations, once the voice-activated electronic device has sent audio data representing a first utterance, the cloud-based backend system can recognize that the voice-activated electronic device is an input device in the stored association. After recognizing this, the cloud-based backend system can then examine what the output device is and whether content preferences exist. For example, the stored association could be an association between the voice-activated electronic device and a television. The stored content preferences could be songs. If this is the case, then in some implementations, a request for a song from the voice-activated user device will cause the cloud-based backend system to send the requested song to the television.

[0032] In some implementations, the cloud-based backend system can determine whether the requested content should be output based on the state of the output electronic device. For example, in an implementation where the output electronic device is a television, the request to play content on the television may depend on whether the television is in an unavailable, available, or ready state. To determine the television's state, in some implementations, the cloud-based backend system can send a status request to the television. If the television does not send back a status response within a predetermined time, it can be considered unavailable. If the television sends back a status response within the predetermined time, the status response may include data indicating whether the television is in a ready or available mode.

[0033] In some implementations, the television may be unavailable when it is turned off. If the cloud-based backend system determines that the television is unavailable, it can receive a text message indicating a notification. The cloud-based backend system can then generate audio data representing the notification text by performing TTS functionality. This audio data can then be sent to and played by the voice-activated device. For example, the voice-activated device could play "Your television is unavailable." Once the voice-activated device has played the notification, the cloud-based backend system can receive requested content. This requested content can then be sent to the voice-activated device for playback.

[0034] In response to being notified that the television is unavailable, an individual can turn on the television, effectively putting it into a ready state. Once in a ready state, the cloud-based backend system can receive a status update from the television notifying it that the television is ready. The cloud-based backend system can then receive a text message indicating a prompt. The cloud-based backend system can then generate audio data representing the text by performing TTS functionality. Before sending the prompt, the cloud-based backend system can generate an answer command for the voice-activated electronic device. The audio data can then be sent to the voice-activated device for playback. For example, the voice-activated device can play "Do you want to play something on the television?" The cloud-based backend system can then send an answer command to the voice-activated electronic device. The answer command causes the voice-activated electronic device to respond and send audio data representing the response to the cloud-based backend system.

[0035] The cloud-based backend system can then receive a response to the request, indicating that the content should continue on the television. The cloud-based backend system can then generate a stop command to stop playback of the content on the voice-activated electronic device. The cloud-based backend system can then send the stop command to the voice-activated electronic device to stop playback. After the voice-activated device has stopped playback, the cloud-based backend system can receive the content again. The content can then be sent to the television so that the television can play it. In some embodiments, the cloud-based backend system can generate a notification to the user that content will be played on the television, and then send it to the voice-activated electronic device for playback. For example, the voice-activated electronic device might play "Content will be played on your television."

[0036] In some implementations, the television may be in a ready state when it is not performing other tasks and is ready to receive and play content. If the television is in a ready state, the cloud-based backend system can receive responsive text data and responsive video data. The cloud-based backend system can generate audio data by performing TTS functionality on the text data. The responsive audio data and responsive video data can then be sent to the television for playback. In some implementations, the cloud-based backend system can generate a notification to the user that content will be played on the television, and then send it to the voice-activated electronic device for playback. For example, the voice-activated electronic device might play "Content will continue on your television."

[0037] In some implementations, the television may be in an available state while performing other tasks. If the television is in an available state, the cloud-based backend system can generate instructions to change the television's state from available to ready. Once generated, the cloud-based backend system can send the instructions. Once the state has changed, the television can send an acknowledgment to the cloud-based backend system that the state has changed from available to ready. Once the television is in a ready state, the cloud-based backend system can receive responsive text data and responsive video data. The cloud-based backend system can generate audio data by performing TTS functionality on the text data. The responsive audio data and responsive video data can then be sent to the television for playback. In some implementations, the cloud-based backend system can generate a notification to the user that content will be played on the television, and then send it to the voice-activated electronic device for playback. For example, the voice-activated electronic device might play "Content will continue on your television." In some implementations, the cloud-based backend system can generate a notification to the user that content will be played on the television, and then send it to the voice-activated electronic device for playback. For example, the voice-activated electronic device plays "The content will continue on your TV".

[0038] Figure 1 This is an illustrative diagram of a system for routing content according to various embodiments. In one exemplary, non-limiting embodiment, voice-activated electronic device 10 can communicate with a backend system 100, which in turn can communicate with an output electronic device 300 associated with the voice-activated electronic device 10. Person 2 can speak Command 4 to the voice-activated electronic device 10 or within the room or space where the voice-activated electronic device 10 is located. As used herein, Command 4 can refer to any question, request, comment, and / or instruction that can be spoken to the voice-activated electronic device 10. For example, Person 2 could ask, “Alexa, what’s the weather forecast?”

[0039] In some implementations, the voice command begins with a wake-up word, which may also be referred to as a trigger expression, wake-up expression, or activation word. In response to the detection of the wake-up word, the voice-activated electronics 10 can be configured to detect and interpret any word following the detected wake-up word as an actionable input or command. In some implementations, the voice-activated electronics 10 can be activated by a phrase or set of words that the voice-activated electronics 10 can also be configured to detect. Therefore, the voice-activated electronics 10 can also detect and interpret any word following that phrase or set of words.

[0040] As used herein, the term "wake word" may correspond to "keyword" or "key phrase," "one activation word" or "multiple activation words," or "trigger," "trigger word," or "trigger expression." An exemplary wake word may be a name, such as the name "Alexa," however, those skilled in the art will recognize that any word (e.g., "Amazon") or a series of words (e.g., "Wake Up" or "Hello, Alexa") may be used alternatively as a wake word. Furthermore, the wake word may be set or configured by the individual operating the voice-activated electronic device 10, and in some embodiments, more than one wake word (e.g., two or more different wake words) may be available for activating the voice-activated electronic device. In yet another embodiment, the trigger for activating the voice-activated electronic device 10 may be any series of time-related sounds.

[0041] In some implementations, the trigger expression can be a nonverbal sound. For example, the sound of a door opening, an alarm sounding, glass breaking, a telephone ringing, or any other sound can alternatively be used to activate the activated device. In this particular scenario, the detection of a nonverbal sound by the activated device (which can alternatively be described as a sound-activated electronic device, substantially similar to voice-activated electronic device 10) can lead to some action or response. For example, if the sound of a door opening is detected (which is also the trigger for the sound-activated device), then that detected trigger can cause a burglar alarm to open.

[0042] Voice-activated electronic device 10 can use one or more microphones residing thereon to detect the said command 4. After detecting command 4, voice-activated electronic device 10 can send audio data representing command 4 to backend system 100. Voice-activated electronic device 10 can also send one or more additional associated data to backend system 100. Various types of associated data that can be included in the audio data include, but are not limited to: the time and / or date when voice-activated electronic device 10 detected command 4, the location of voice-activated electronic device 10 (e.g., GPS location), the IP address associated with voice-activated electronic device 10, the device type of voice-activated electronic device 10, or any other type of associated data, or any combination thereof. For example, when person 2 says command 4, voice-activated electronic device 10 can obtain the GPS location of voice-activated electronic device 10 to determine the location of person 2 and the time / date (e.g., hour, minute, second, day, month, year, etc.) when command 4 was made.

[0043] Audio data and associated data can be transmitted to backend system 100 via a network (such as the Internet) using any number of communication protocols. For example, Transmission Control Protocol and Internet Protocol (“TCP / IP”) (e.g., any of the protocols used in each TCP / IP layer), Hypertext Transfer Protocol (“HTTP”), and Wireless Application Protocol (“WAP”) are some of the various types of protocols that can be used to facilitate communication between voice-activated electronic device 10 and backend system 100. In some implementations, voice-activated electronic device 10 and backend system 100 can communicate with each other using HTTP via a web browser. Various other communication protocols can be used to facilitate communication between the voice-activated electronic device 10 and the back-end system 100, including but not limited to: Wi-Fi (e.g., 802.11 protocol), Bluetooth®, radio frequency systems (e.g., 900 MHz, 1.4 GHz, and 5.6 GHz communication systems), cellular networks (e.g., GSM, AMPS, GPRS, CDMA, EV-DO, EDGE, 3GSM, DECT, IS-136 / TDMA, iDen, LTE, or any other suitable cellular network protocol), infrared, BitTorrent, FTP, RTP, RTSP, SSH, and / or VoIP.

[0044] Backend system 100 may include one or more servers, each communicating with each other, voice-activated electronics 10, and / or output electronics 300. Backend system 100 and output electronics 300 may communicate with each other using any of the aforementioned communication protocols. Each server within backend system 100 may be associated with one or more databases or processors capable of storing, retrieving, processing, analyzing, and / or generating data to be provided to voice-activated electronics 10. For example, backend system 100 may include one or more servers each corresponding to a category. As an example, backend system 100 may include a "weather" category server comprising one or more databases of weather information (e.g., forecasts, radar images, allergy information, etc.). As another example, backend system 100 may include a "sports" category server comprising one or more databases containing various sports or athletic information (e.g., scores, teams, matches, etc.). As yet another example, backend system 100 may include a "transportation" category server comprising one or more databases containing transportation information for various geographic areas (e.g., street maps, traffic warnings, traffic conditions, direction information, etc.). In some implementations, backend system 100 may correspond to a series of servers located in a remote facility, and individuals may use one or more of the aforementioned communication protocols to store data on backend system 100 and / or communicate with backend system 100.

[0045] In some embodiments, the backend system 100 may include one or more servers capable of storing data structures 102 that associate voice-activated electronic device 10 with output electronic device 300. Data structure 102 may be, for example, a file, database entry, or other type of data structure capable of storing information indicating the association between voice-activated electronic device 10 and output electronic device 300. Data structure 102 may include, for example, device identification information for voice-activated electronic device 10 and output electronic device 300. Data structure 102 may also include additional information about voice-activated electronic device 10 and / or output electronic device 300. In some embodiments, data structure 102 may include the type of output electronic device 300 (e.g., television, streaming media device, speaker system, etc.). Data structure 102 may also include information about the status of output electronic device 300 (e.g., ready, available, unavailable). The backend system 100 can determine whether voice-activated electronic device 10 is associated with output electronic device 300 based on data structure 102.

[0046] Output electronics 300 can be one or more electronic devices of any type capable of outputting visual and / or auditory content. In some embodiments, output electronics 300 may include streaming media device 302 (media storage device) and peripheral video output device 304 (e.g., television or monitor) connected to streaming media device 302. The video output device can be any device capable of receiving and outputting content. Streaming media device 302 may be able to receive content from backend system 100 or other information sources and provide such content to video output device 304 according to a protocol compatible with video output device 304. In some embodiments, streaming media device 302 may provide content to video output device 304 according to the High Definition Multimedia Interface (HDMI) protocol. Streaming media device 302 may also be able to communicate with and control video output device 304. For example, streaming media device 302 may be able to communicate with video output device 304 to determine whether video output device 304 is turned on. Streaming media device 302 may also be able to communicate with video output device 304 to determine whether video output device 304 is set as an input source associated with streaming media device 302. Streaming device 302 can also control video output device 304 to perform functions such as turning it on or off, switching to a selected input source, adjusting the volume of video output device 304, or controlling other functions of video output device 304. In some embodiments, streaming device 302 can communicate with and control video output device 304 using the Consumer Electronics Control (CEC) protocol. The CEC protocol allows a device to control the HDMI functionality of another device connected to it via the HDMI protocol. Those skilled in the art will understand that streaming device 302 can also communicate with and control video output device 304 using other protocols. In some embodiments, the output electronics device 300 can be a functional video output device incorporated into streaming device 302 (e.g., a smart TV). Additionally, in some embodiments, the output electronics device 300 can be an audio output device, such as a speaker or speaker system (e.g., a base unit and several peripheral speakers connected to the base unit).

[0047] Returning to the reference backend system 100, once the backend system 100 receives audio data from the voice-activated electronic device 10, the backend system 100 can analyze the audio data, for example, by performing STT functionality on the audio data, to determine which words are included in the said command 4. Then, the backend system 100 can perform NLU functionality to determine the intent or meaning of the said command 4. The backend system 100 can further determine the response to the said command 4. In some embodiments, the backend system 100 can determine that the voice-activated electronic device 10 is associated with the output electronic device 300, and can also determine that the response to the said command 4 should include outputting content through the output electronic device 300. Additionally, the backend system 100 can determine that the response should include outputting a notification through the voice-activated electronic device 10 to inform the individual 2 that content will be output through the output electronic device 300. The backend system 100 can also determine that the response should include outputting a notification through the output electronic device 300 to inform the individual that content will be output through the output electronic device 300. The following... Figure 3 The backend system is described in more detail.

[0048] For example, in some implementations, the response to command 4 may include content, such as a weather forecast. Backend system 100 may first determine that output electronic device 300 is associated with voice-activated electronic device 10 by looking up the association between voice-activated electronic device 10 and output electronic device 300 stored in data structure 102. Then, backend system 100 may determine that content should be output through output electronic device 300. In some implementations, determining that output electronic device 300 is associated with voice-activated electronic device 10 is sufficient to determine that content should be output through output electronic device 300. However, in some implementations, backend system 100 may consider additional information (such as the status of output electronic device 300, content type, user preferences, or other additional information) when determining whether content should be output through output electronic device 300, as will be described in more detail. When determining that content should be output through output electronic device 300, backend system 100 may use text-to-speech (TTS) processing to generate first responsive audio data. The first responsive audio data may represent a first audio message 12 that notifies individual 2 that content will be output by output electronic device 300. The backend system 100 may send first responsive audio data to the voice-activated electronic device 10. In some embodiments, the backend system 100 may also send data representing an instruction to cause the first audio message 12 to be played on the voice-activated electronic device 10 upon receipt. For example, after receiving the first audio data and any associated instruction, the first audio message 12, such as “Show the weather forecast on your TV”, may be played on the voice-activated electronic device 10. The first audio message 12 may also incorporate information identifying the output electronic device 300 (e.g., “Your TV”, “Your speaker system”, etc.).

[0049] The backend system 100 can also use TTS processing to generate second responsive audio data. The second responsive audio data can represent a second audio message 14, which notifies individual 2 that content will be output by output electronics 300. After sending first responsive data to voice-activated electronics 10, the backend system 100 can send the second audio data to output electronics 300. In some embodiments, the backend system 100 can also send data representing an instruction to cause the second audio message 14 to be played on output electronics 300 upon receipt. For example, after receiving the second audio data and any associated instruction, a second audio message 14 such as "Here is the weather forecast" can be played on output electronics 300. Playing the first audio message 12 on voice-activated electronics 10 and then subsequently playing the second audio message 14 on output electronics 300 provides an enhanced experience for individual 2 by notifying individual 2 where content will be output and allowing individual 2 to identify the output electronics 300 where content will be output.

[0050] In some embodiments, after sending first responsive audio data to voice-activated electronic device 10 and second responsive audio data to output electronic device 300, backend system 100 may send an instruction to output electronic device 300 causing output electronic device 300 to output content in response to said command 4. Backend system 100 may also send content in response to said command 4 to output electronic device 300. For example, in some embodiments, backend system 100 may determine that the response to said command 4 should include content such as a weather forecast. Backend system 100 may retrieve content (e.g., a weather forecast) from one or more category servers (e.g., a "weather" category server) and send the content along with instructions for outputting the content to output electronic device 300. Upon receiving the content and instructions, output electronic device 300 may output the content (e.g., display a weather forecast). Although a weather forecast has been described as a type of content associated with embodiments of the disclosed concept, those skilled in the art will understand that content can include various types of visual and / or auditory content (e.g., movies, pictures, audiobooks, music, etc.).

[0051] In some implementations, backend system 100 may send instructions to output electronic device 300 causing output electronic device 300 to output content, and output electronic device 300 may obtain content from a source other than backend system 100. In some implementations, the content may already be stored on output electronic device 300, so backend system 100 does not need to send the content to output electronic device 300. Moreover, in some implementations, output electronic device 300 may be able to retrieve content from a cloud-based system other than backend system 100. For example, output electronic device 300 may be connected to a video or audio streaming service other than backend system 100. Backend system 100 may send instructions to output electronic device 300 causing it to retrieve selected content from a cloud-based system such as a video or audio streaming service and output the selected content. For example, backend system 100 may determine that command 4 includes a request to play a specific program. Backend system 100 may determine that content from a video streaming service is available for playback. For example, the user account associated with voice-activated electronic device 10 may include information instructing individual 2 to subscribe to a video streaming service. The backend system 100 can further determine whether the requested program is available through the video streaming service by communicating with the video streaming service or consulting other information sources (such as a database that identifies which content is available through the video streaming service). Finally, the backend system 100 can send an instruction to the output electronic device 300, causing the output electronic device 300 to request the program from the video streaming service and begin playing the requested program.

[0052] refer to Figure 2This diagram illustrates various embodiments for routing content based on the state of output electronic device 300. In some embodiments, the backend system 100 may consider the state of output electronic device 300 when determining whether a response to said command 4 should include outputting content via output electronic device 300. For example, output electronic device 300 may have a ready state, an available state, and an unavailable state. In the ready state, output electronic device 300 may be ready to output content. For example, in some embodiments, in the ready state, streaming device 302 is powered on, and the peripheral video output device is powered on and set to the input source associated with streaming device 302. In the available state, output electronic device 300 may be available for outputting content, but additional steps may be required to prepare output electronic device 300 for outputting content. For example, in some embodiments, output electronic device 300 may be in the available state when streaming device 302 is powered on, but peripheral video output device 304 is powered off or not set to the input source associated with streaming device 302. Before output electronics 300 is ready to output content, streaming media device 302 may need to power on video output device 304 or switch its input to an input source associated with streaming media device 302. In an unavailable state, output electronics 300 is unavailable for outputting content. For example, in some embodiments, output electronics 300 may be unavailable when streaming media device 302 is powered off. Furthermore, in some embodiments, output electronics 300 may be unavailable when peripheral video output device 304 is disconnected from streaming media device 302.

[0053] In some implementations, backend system 100 can determine the status of output electronic device 300. For example, backend system 100 communicates with output electronic device 300 by querying for the status of output electronic device 300. Output electronic device 300 can determine its status and can respond with information indicating its status. For example, streaming media device 302 can receive a query from backend system 100 and then communicate with peripheral video output device 304 to determine whether peripheral video output device 304 is connected, powered, and set to an input source associated with streaming media device 302. Streaming media device 302 can communicate with peripheral video output device 304 using the CEC protocol and determine whether peripheral video output device 304 is connected, powered, and set to an input source associated with streaming media device 302. For example, if streaming media device 302 determines that peripheral video output device 304 is connected, powered, and set to an input source associated with streaming media device 302, then streaming media device 302 can determine that the output electronic devices (e.g., streaming media device 302 and peripheral video output device 304) are in a ready state. The streaming media device 302 can transmit information indicating the status of the output electronic device 300 to the backend system 100, and the backend system 100 can store the information. In some embodiments, the backend system 100 can determine that the output electronic device 300 is unavailable based on the output electronic device 300's failure to respond to a query from the backend system 100. In some embodiments, the backend system 100 can store the determined status information of the output electronic device 300 in, for example, data structure 102.

[0054] Based on the determined state of the output electronic device 300, the backend system 100 can determine where to route the requested content. For example, in some embodiments, the backend system 100 can determine that the command 4 includes a request for content that should be output by the output electronic device 300 if the output electronic device 300 is in a ready state. However, if the output electronic device 300 is not in a ready state (e.g., the output electronic device is in an available or unavailable state), then the backend system 100 can send the requested content to the voice-activated electronic device 10. For example, if the requested content is a weather forecast, then the backend system 100 can retrieve the content (e.g., the weather forecast) from one or more category servers (e.g., a "weather" category server). The backend system 100 can use text-to-speech (TTS) processing to generate responsive audio data, and the responsive audio data can represent a first audio message 12 incorporated into the content. The backend system 100 can send the responsive audio data together with data representing an instruction to cause the first audio message 12 to be played on the voice-activated electronic device 10 when received. For example, after receiving the returned file 8, a first audio message 12, such as "Tomorrow's weather forecast is sunny and 70 degrees Celsius," can be played on the voice-activated electronic device 10. On the other hand, if the backend system 100 determines that the output electronic device 300 is ready, then the backend system 100 can continue to send first responsive audio data representing the first audio message 12, which notifies the user 2 that content will be output by the output electronic device 300. The backend system 100 can then subsequently send second responsive audio data notifying the user 2 that content will be output by the output electronic device 300, and then send instructions for outputting the content to the output electronic device 300.

[0055] Figure 3 It is based on various implementation plans. Figure 1The system architecture is illustrated in the diagram. In some embodiments, the voice-activated electronic device 10 can correspond to any type of electronic device capable of being activated in response to the detection of a specific sound. In some embodiments, the voice-activated electronic device 10 can recognize a command (e.g., an audio command, input) within the captured audio after detecting a specific sound (e.g., a wake word or trigger), and can perform one or more actions in response to the received command. Various types of electronic devices can include, but are not limited to: desktop computers, mobile computers (e.g., laptops, ultrabooks), mobile phones, smartphones, tablets, televisions, set-top boxes, smart TVs, watches, bracelets, displays, personal digital assistants (“PDAs”), smart furniture, smart home devices, smart vehicles, smart transportation devices, and / or smart accessories. In some embodiments, the voice-activated electronic device 10 can be structurally relatively simple or basic, such that it may not provide one or more mechanical input options (e.g., a keyboard, mouse, touchpad) or one or more touch input devices (e.g., a touchscreen, buttons). For example, the voice-activated electronic device 10 may be capable of receiving and outputting audio, and may include power, processing power, storage / memory capabilities, and communication capabilities.

[0056] The voice-activated electronic device 10 may include a very small number of input mechanisms, such as a power-on / power-off switch; however, in one embodiment, the primary functionality of the voice-activated electronic device 10 may be solely through audio input and audio output. For example, the voice-activated electronic device 10 may respond to a wake word (e.g., “Alexa” or “Amazon”) by continuously monitoring local audio. In response to detecting a wake word, the voice-activated electronic device 10 may establish a connection with the backend system 100, send audio data to the backend system 100, and wait for / receive a response from the backend system 100. However, in some embodiments, a non-voice-activated electronic device may also communicate with the backend system 100 (e.g., pressing or tapping the call device). For example, in one embodiment, the activated device corresponds to a manually activated electronic device, and the foregoing description may also apply to non-voice-activated electronic devices.

[0057] Voice-activated electronics 10 may include one or more processors 202, storage devices / memory 204, communication circuitry 206, one or more microphones 208 or other audio input devices (e.g., transducers), one or more speakers 210 or other audio output devices, and optional input / output (“I / O”) interfaces 212. However, one or more additional components may be included within voice-activated electronics 10, and / or one or more components may be omitted. For example, voice-activated electronics 10 may include a power supply or a bus connector. As another example, voice-activated electronics 10 may not include an I / O interface. Furthermore, while there are multiple instances where one or more components may be included within voice-activated electronics 10, only one example of each component is shown for simplicity.

[0058] One or more processors 202 may include any suitable processing circuitry capable of controlling the operation and functionality of the voice-activated electronic device 10 and facilitating communication between various components within the voice-activated electronic device 10. In some embodiments, one or more processors 202 may include a central processing unit (“CPU”), a graphics processing unit (“GPU”), one or more microprocessors, digital signal processors, or any other type of processor, or any combination thereof. In some embodiments, the functionality of one or more processors 202 may be performed by one or more hardware logic components, including but not limited to: field-programmable gate arrays (“FPGA”), application-specific integrated circuits (“ASIC”), application-specific standard products (“ASSP”), system-on-a-chip systems (“SOC”), and / or complex programmable logic devices (“CPLD”). Furthermore, each processor 202 may include its own local memory, which may store program modules, program data, and / or one or more operating systems. However, one or more processors 202 may run an operating system (“OS”) for the voice-activated electronic device 10, and / or one or more firmware applications, media applications, and / or applications residing thereon.

[0059] Storage device / memory 204 may include one or more types of storage media for storing data on voice-activated electronic device 10, such as any volatile or non-volatile memory or any removable or non-removable memory implemented in any suitable manner. For example, information may be stored using computer-readable instructions, data structures, and / or program modules. Various types of storage devices / memories may include, but are not limited to: hard disk drives, solid-state drives, flash memory, persistent memory (e.g., ROM), electrically erasable programmable read-only memory (“EEPROM”), CD-ROM, digital versatile optical disc (“DVD”) or other optical storage media, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, RAID storage systems, or any other storage type, or any combination thereof. Furthermore, storage device / memory 204 may be implemented as a computer-readable storage medium (“CRSM”), which may be any available physical medium accessible by one or more processors 202 to execute one or more instructions stored within storage device / memory 204. In some embodiments, one or more applications (e.g., games, music, videos, calendars, lists, etc.) may be run by one or more processors 202 and may be stored in memory 204.

[0060] In some embodiments, the storage device / memory 204 may include one or more modules and / or databases, such as a speech recognition module 214, a wake word list database 216, and a wake word detection module 218. The speech recognition module 214 may, for example, include an automatic speech recognition (“ASR”) component that identifies human speech in detected audio. The speech recognition module 214 may also include a natural language understanding (“NLU”) component that determines the user’s intent based on the detected audio. The speech recognition module 214 may also include a text-to-speech (“TTS”) component capable of converting text into speech for output by one or more speakers 210, and / or a speech-to-text (“STT”) component capable of converting received audio signals into text for transmission to the back-end system 100 for processing.

[0061] The wake-up word list database 216 may be a database locally stored on the voice-activated electronic device 10, including a current list of wake-up words for voice-activated electronic device 10, and one or more previously used or alternative wake-up words for voice-activated electronic device 10. In some embodiments, the individual 2 may set or configure a wake-up word for voice-activated electronic device 10. The wake-up word may be set directly on the voice-activated electronic device 10, or one or more wake-up words may be set by the individual through a backend system application communicating with backend system 100. For example, the individual 2 may use their mobile device with a backend system application running thereon to set the wake-up word. The specific wake-up word may then be transmitted from the mobile device to backend system 100, which may then send / notify the individual of the wake-up word selection to voice-activated electronic device 10. The selected activation may then be stored in database 216 of storage device / memory 204.

[0062] The wake-word detection module 218 may include an expression detector that analyzes audio signals generated by one or more microphones 208 to detect wake-words, which may generally be predefined words, phrases, or any other sounds, or any series of temporally related sounds. As an example, such an expression detector may be implemented using keyword detection techniques. A keyword detector may be a functional component or algorithm that evaluates audio signals to detect the presence of predefined words or expressions within the audio signals detected by one or more microphones 208. Instead of generating a transcript of the spoken words, the keyword detector generates a true / false output (e.g., logic 1 / 0) indicating whether a predefined word or expression is represented in the audio signal. In some embodiments, the expression detector may be configured to analyze the audio signal to generate a score indicating the probability that a wake-word is represented within the audio signals detected by one or more microphones 208. The expression detector may then compare that score to a threshold to determine whether the wake-word will be asserted as spoken.

[0063] In some implementations, the keyword detector may use a simplified ASR technique. For example, the expression detector may use a Hidden Markov Model (“HMM”) recognizer, which performs acoustic modeling of the audio signal and compares the HMM model of the audio signal with one or more reference HMM models that have been created by training on specific trigger expressions. The HMM model represents a word as a series of states. Overall, a portion of the audio signal is analyzed by comparing the HMM model of the audio signal with the HMM model of the trigger expression, thereby generating feature scores representing the similarity between the audio signal model and the trigger expression model.

[0064] In practice, the HMM recognizer can generate multiple feature scores corresponding to different features of the HMM model. The expression detector can use a support vector machine (“SVM”) classifier, which receives one or more feature scores generated by the HMM recognizer. The SVM classifier generates a confidence score indicating the probability that the audio signal contains a trigger expression. The confidence score is compared with a confidence threshold to make a final decision about whether a particular part of the audio signal represents a utterance that triggers an expression (e.g., a wake word). Upon asserting that the audio signal represents a utterance that triggers an expression, the voice activation electronics 10 can then begin transmitting the audio signal to the backend system 100 for detection and analysis of subsequent utterances made by the individual 2.

[0065] Communication circuitry 206 may include any circuitry that allows voice-activated electronic device 10 to communicate with one or more devices, servers, and / or systems. For example, communication circuitry 206 may facilitate communication between voice-activated electronic device 10 and backend system 100. Communication circuitry 206 may use any communication protocol, such as any of the previously mentioned exemplary communication protocols. In some embodiments, voice-activated electronic device 10 may include an antenna to facilitate wireless communication with networks using various wireless technologies, such as Wi-Fi, Bluetooth®, radio frequency, etc. In yet another embodiment, voice-activated electronic device 10 may include one or more Universal Serial Bus (“USB”) ports, one or more Ethernet or broadband ports, and / or any other type of hardwired access port, such that communication circuitry 206 allows voice-activated electronic device 10 to communicate with one or more communication networks.

[0066] The voice-activated electronic device 10 may also include one or more microphones 208 and / or transducers. The one or more microphones 208 can be any suitable component capable of detecting audio signals. For example, the one or more microphones 208 may include one or more sensors for generating electrical signals and circuitry capable of processing the generated electrical signals. In some embodiments, the one or more microphones 208 may include multiple microphones capable of detecting various frequency levels. As an illustrative example, the voice-activated electronic device 10 may include multiple microphones (e.g., four, seven, ten, etc.) placed at various locations to monitor / capture any audio output in the environment in which the voice-activated electronic device 10 is located. The various microphones 208 may include some microphones optimized for long-distance sound, while others may be optimized for sound occurring within the near-field range of the voice-activated electronic device 10.

[0067] The voice-activated electronic device 10 may also include one or more speakers 210. The one or more speakers 210 may correspond to any suitable mechanism for outputting audio signals. For example, the one or more speakers 210 may include one or more speaker units, transducers, speaker arrays, and / or transducer arrays capable of broadcasting audio signals and / or audio content to an area surrounding the voice-activated electronic device 10 where it may be located. In some embodiments, the one or more speakers 210 may include headphones or earpieces that can be wirelessly or hardwired to the voice-activated electronic device 10, said headphones or earpieces being able to broadcast audio directly to the individual 2.

[0068] In some implementations, the voice-activated electronics 10 can be hardwired or wirelessly connected to one or more speakers 210. For example, the voice-activated electronics 10 can cause one or more speakers 210 to output audio thereon. In this particular scenario, the voice-activated electronics 10 can receive audio to be output by the speakers 210, and the voice-activated electronics 10 can send the audio to the speakers 210 using one or more communication protocols. For example, the voice-activated electronics 10 and one or more speakers 210 can communicate with each other using a Bluetooth® connection or another near-field communication protocol. In some implementations, the voice-activated electronics 10 can communicate indirectly with one or more speakers 210.

[0069] In some embodiments, one or more microphones 208 may be used as input devices for receiving audio input, such as speech from individual 2. In the previously mentioned embodiments, the voice-activated electronics 10 may then further include one or more speakers 210 for outputting an auditory response. In this way, the voice-activated electronics 10 may function solely by voice or audio without the use of or requirement of any input mechanism or display.

[0070] In one exemplary embodiment, the voice-activated electronic device 10 includes an I / O interface 212. The input portion of the I / O interface 212 can correspond to any suitable mechanism for receiving input from a user of the voice-activated electronic device 10. For example, a camera, keyboard, mouse, joystick, or external controller can be used as the input mechanism for the I / O interface 212. The output portion of the I / O interface 212 can correspond to any suitable mechanism for generating output from the voice-activated electronic device 10. For example, one or more displays can be used as the output mechanism for the I / O interface 212. As another example, one or more lamps, light-emitting diodes (“LEDs”), or one or more other visual indicators can be used to output signals through the I / O interface 212 of the voice-activated electronic device 10. In some embodiments, the I / O interface 212 may include one or more vibration mechanisms or other tactile features to provide a tactile response from the voice-activated electronic device 10 to the individual 2. Those skilled in the art will recognize that in some embodiments, one or more features of the I / O interface 212 may be included in the voice-activated electronic device 10 in a purely voice-activated version. For example, one or more LEDs may be included on the voice-activated electronics 10 such that when one or more microphones 208 receive audio from the individual 2, the one or more LEDs light up, indicating that the voice-activated electronics 10 has received audio. In some embodiments, the I / O interface 212 may include a display screen and / or touch screen that may have any size and / or shape and may be located at any part of the voice-activated electronics 10. Various types of displays may include, but are not limited to: liquid crystal displays (“LCDs”), monochrome displays, color graphics adapter (“CGA”) displays, enhanced graphics adapter (“EGA”) displays, variable graphics array (“VGA”) displays, or any other type of display, or any combination thereof. Furthermore, in some embodiments, the touch screen may correspond to a display screen that includes a capacitive sensing panel capable of recognizing touch input thereon.

[0071] As previously mentioned, in some embodiments, backend system 100 may communicate with voice-activated electronic device 10. Backend system 100 includes various components and modules, including but not limited to: Automatic Speech Recognition (“ASR”) module 258, Natural Language Understanding (“NLU”) module 260, Skills module 262, Text-to-Speech (“TTS”) module 264, and User Account module 268. Speech-to-Text (“STT”) module 266 may be included in ASR module 258. In some embodiments, backend system 100 may also include computer-readable media, including but not limited to: flash memory, random access memory (“RAM”), and / or read-only memory (“ROM”). Backend system 100 may also include various modules storing software, hardware, logic, instructions, and / or commands for backend system 100, such as a speaker identification (“ID”) module, a user profile module, or any other module, or any combination thereof.

[0072] The backend system 100 may also include a content routing module 270. In one embodiment, the content routing module 270 may include one or more processors 252, storage devices / memory 254, and communication circuitry 256. In some embodiments, the one or more processors 252, storage devices / memory 254, and communication circuitry 256 may be substantially similar to the one or more processors 202, storage devices / memory 204, and communication circuitry 206 described in more detail above, and the above description of the latter may apply. Data structure 102 may be stored within the content routing module 270. The content routing module 270 may be configured to determine whether content should be output by the voice-activated electronic device 10 or by the output electronic device 300. The content routing module 270 may also store programs and / or instructions for facilitating the determination of whether content should be output by the voice-activated electronic device 10 or by the output electronic device 300.

[0073] The ASR module 258 can be configured to recognize human speech in detected audio (such as audio captured by the voice-activated electronics 10). In one embodiment, the ASR module 258 may include one or more processors 252, storage devices / memory 254, and communication circuitry 256. In some embodiments, the one or more processors 252, storage devices / memory 254, and communication circuitry 256 may be substantially similar to the one or more processors 202, storage devices / memory 204, and communication circuitry 206 described in more detail above, and the foregoing description of the latter may apply. The NLU module 260 can be configured to determine user intent based on detected audio received from the voice-activated electronics 10. The NLU module 260 may include one or more processors 252, storage devices / memory 254, and communication circuitry 256. In some embodiments, the ASR module 258 may include an STT module 266. The STT module 266 may employ various speech-to-text technologies. However, the techniques used to transcribe speech into text are well known in the art and do not need to be described in further detail herein, and any suitable computer-implemented speech-to-text technology (such as the SOFTSOUND® speech processing technology available from Autonomy, a company headquartered in Cambridge, England) can be used to convert one or more received audio signals into text.

[0074] Skill module 262 may, for example, correspond to various action-specific skills or servers capable of handling various task-specific actions. Skill module 262 may also correspond to first-party and / or third-party applications operable to perform different tasks or actions. For example, based on the context of audio received from voice-activated electronic device 10, backend system 100 may use an application or skill to retrieve or generate a response, which may then be transmitted back to voice-activated electronic device 10. Skill module 262 may include one or more processors 252, storage devices / memory 254, and communication circuitry 256. As an illustrative example, skill module 262 may correspond to one or more game servers for storing and processing information related to different games (e.g., "Simon Says," "karaoke," etc.). As another example, skill module 262 may include one or more weather servers for storing weather information and / or providing weather information to voice-activated electronic device 10.

[0075] The TTS module 264 can employ various text-to-speech technologies. Techniques for transcribing speech into text are well-known in the art and do not require further description herein; any suitable computer-implemented speech-to-text technology (such as the SOFTSOUND® speech processing technology available from Autonomy, headquartered in Cambridge, England) can be used to convert one or more received audio signals into text. The TTS module 264 may also include one or more processors 252, storage devices / memory 254, and communication circuitry 256. In some embodiments, one or more filters may be applied to the received audio data to reduce or minimize external noise.

[0076] User account module 268 may store one or more user profiles corresponding to users with registered accounts on backend system 100. For example, a parent may have a registered account on backend system 100, and each of the parent's children may register their own user profile under the parent's registered account. Information for each user profile (e.g., settings and / or preferences) may be stored in a user profile database. In some embodiments, user account module 268 may store voice signals, such as voice biometric information, for a particular user profile. This may allow the use of speaker recognition technology to match voice with voice and voice biometric data associated with a particular user profile. In some embodiments, user account module 268 may store telephone numbers assigned to a particular user profile. User account module 268 may also include one or more processors 252, storage devices / memory 254, and communication circuitry 256.

[0077] Those skilled in the art will recognize that although each of the ASR module 258, NLU module 260, skill module 262, TTS module 264, and user account module 268 includes instances of one or more processors 252, storage devices / memory 254, and communication circuitry 256, those instances of one or more processors 252, storage devices / memory 254, and communication circuitry 256 within each of the ASR module 258, NLU module 260, skill module 262, TTS module 264, and user account module 268 may differ. For example, the structure, function, and style of one or more processors 252 within the ASR module 258 may be substantially similar to the structure, function, and style of one or more processors 252 within the NLU module 260, but the actual one or more processors 252 need not be the same entity.

[0078] As previously mentioned, in some embodiments, the backend system 100 may also communicate with the output electronics 300. In some embodiments, the output electronics 300 may include a streaming media device 302 and a peripheral video output device 304 connected to the streaming media device 302. The streaming media device 302 may include one or more processors 306, storage devices / memory 308, and communication circuitry 310. As previously mentioned, the streaming media device 302 may communicate with the peripheral video output device 304. Additionally, the streaming media device 302 may communicate with cloud-based systems such as audio or video streaming services. Various types of output electronics include, but are not limited to: televisions, portable media players, cellular phones or smartphones, pocket personal computers, personal digital assistants (“PDAs”), desktop computers, laptop computers, tablet computers, and / or electronic accessory devices (such as smartwatches and bracelets).

[0079] The peripheral video output device 304 may include one or more processors 306, storage devices / memory 308, communication circuitry 310, a display 312, and a speaker 314. The display 312 may be a screen and / or touchscreen that can have any size and / or shape and can be located at any part of the voice-activated electronics 10. Various types of displays may include, but are not limited to: liquid crystal displays (“LCDs”), monochrome displays, color graphics adapter (“CGA”) displays, enhanced graphics adapter (“EGA”) displays, variable graphics array (“VGA”) displays, or any other type of display, or any combination thereof.

[0080] Those skilled in the art will understand that in some embodiments, the streaming media device 302 and the peripheral video output device 304 may be separate devices, or in some embodiments, they may be combined into a single device. For example, without departing from the scope of the disclosed concepts, the functionality of the streaming media device 302 may be integrated into the video output device 304. The streaming media device may be any device capable of communicating with the backend system 100. Various types of streaming media devices include, but are not limited to: Fire TV Sticks, Fire TV Sticks with voice remote control, televisions, portable media players, cellular phones or smartphones, pocket personal computers, personal digital assistants (“PDAs”), desktop computers, laptop computers, tablet computers, and / or electronic accessory devices (such as smartwatches and bracelets). Those skilled in the art will also understand that the streaming media device 302 and the peripheral video output device 304 are examples of output electronics 300. Output electronics 300 may be any type of electronic device or combination of devices capable of outputting auditory or visual content. For example, in some embodiments, output electronics 300 may include the streaming media device 302 and one or more connected peripheral audio devices (such as speakers).

[0081] Figure 4 Table 400 illustrates different content categories according to embodiments of the disclosed concepts. The backend system 100 may further consider the type of requested content when determining where to send the content. For example, Table 400 shows different content types categorized into different categories instructing the backend system 100 where to route the content. Table 400 may include a first category 402 corresponding to the type of content that should only be output by the output electronic device 300. Figure 4 In the example shown, the first category 402 includes videos and images. In some embodiments, the backend system 100 may determine that the command 4 includes a request for content from the first category 402, and the backend system 100 may send the requested content to the output electronic device 300. Alternatively, in some embodiments, the backend system may determine the status of the output electronic device 300 by querying it, for example, through a request for its status. In some embodiments, if the backend system 100 determines that the output electronic device 300 is in an available state rather than a ready state, the backend system 100 may first send an instruction to the output electronic device 300 causing it to change from an available state to a ready state (e.g., the instruction may cause the streaming media device 302 to use commands under the CEC protocol to open the peripheral video output device 304 and set the peripheral video output device 304 to the input source associated with the streaming media device 302), and then send an instruction to the output electronic device 300 causing it to output content.

[0082] In some implementations, backend system 100 may determine that the output electronic device 300 is unavailable (e.g., streaming device 302 is powered off or otherwise does not respond to queries requesting its status, or peripheral video output device 304 is not connected to streaming device 302). If backend system 100 determines that its output electronic device 300 is unavailable, then backend system 100 may generate responsive text data representing an audio message for notifying individual 2 that content cannot be played (e.g., "Content cannot be played because the associated TV is not connected"), and may send audio data to activate the output of electronic device 10 by voice.

[0083] Table 400 may also include a second category 404, which may include the types of content output by the output electronic device 300 or the voice-activated electronic device 10, depending on whether the output electronic device 300 is in a ready state. Figure 4 In the example shown, the second category 404 includes music, weather, and audiobooks. In some implementations, the backend system 100 may determine that the command 4 includes a request for content from the second category 404, and then further determine whether the output electronic device 300 is in a ready state by querying the output electronic device 300 to request its status. If the backend system 100 determines that the output electronic device 300 is in a ready state based on the response from the output electronic device 300, then the backend system 100 may send an instruction to the output electronic device 300 to output the content from the second category 404. However, if the backend system 100 determines that the output electronic device 300 is not in a ready state, then the backend system 100 may instead send the content from the second category 404 to the voice-activated electronic device 10. For example, if the requested content is a weather forecast and the backend system 100 determines that the weather forecast is a content type in the second category, then if the output electronic device 300 is in a ready state, the backend system 100 may simply send an instruction to output the requested weather forecast via the output electronic device 300. In some implementations, when the streaming device 302 is powered on but the peripheral video output device 304 is powered off, the output electronic device 300 will be in an available state but not in a ready state, and the backend system 100 can send weather forecasts to the voice-activated electronic device 10 instead of the output electronic device 300.

[0084] Finally, Table 400 may include a third category 406, which may include types of content that can be primarily routed to the voice-activated electronic device 10 due to the nature and format of the content. Figure 4In the example shown, the third category 406 includes content such as alarms and timers. This content can be essentially audio, and in particular, can be of a type where quality is less important. For example, the primary purpose of an alarm is to make it audible, at which point it is typically turned off. Therefore, this content can more appropriately be provided via a voice-activated electronic device 10 that can be physically closer to the individual user. In those embodiments, the backend system 100 can determine that the command 4 includes a request for content from the third category 406, and can then send the requested content to the voice-activated electronic device 10. Even if the voice-activated electronic device 10 is associated with the output electronic device 300 and the output electronic device 300 is in a ready state, the content from the third category 406 should generally be sent to the voice-activated electronic device 10.

[0085] although Figure 4 Table 400 illustrates some examples of content types in Category 1 402, Category 2 404, and Category 3 406, but those skilled in the art will understand that other or different content types may be included in Table 400. Furthermore, the content types in Table 400 and their division between categories are merely examples, and those skilled in the art will understand that content types and their division between categories may be varied without departing from the scope of the disclosed concepts. Figure 4 The examples shown are different. Furthermore, the content type and its classification within categories can be set and changed by the user of the voice-activated electronic device 10. Figure 4 The information included therein can be stored on the backend system 100, for example, in the content routing module 270.

[0086] In some implementations, the state of the output electronic device 300 may change while outputting content. For example, while content is being output through the peripheral video output device 304, individual 2 may turn off the peripheral video output device 304, causing the state of the output electronic device 300 to change from ready to available. In some implementations, the output electronic device 300 can monitor any changes in its state and transmit said changes to the backend system 100. For example, the streaming media device 302 can use the CEC protocol to periodically monitor whether the peripheral video output device 304 has been turned off or its input source has been changed to determine whether the state of the output electronic device 300 has changed from ready to available. The streaming media device 302 can then transmit the state change to the backend system 100. In some implementations, the backend system 100 can periodically query the output electronic device 300 to request its state. Based on the response from the output electronic device 300, the backend system 100 can determine whether the state of the output electronic device 300 has changed. In some implementations, the backend system 100 can determine that the state of the output electronic device 300 has changed from a ready state to an available or unavailable state, and can send a stop command to the output electronic device 300 to stop outputting content based on the state change. For example, the command could cause the streaming media device 302 to stop sending content to the peripheral video output device 304. In some implementations, when the state of the output electronic device 300 changes from a ready state to an available or unavailable state, the backend system 100 can subsequently begin sending content to the voice-activated electronic device 10.

[0087] Similarly, in some implementations, when content is being output via voice activation of electronic device 10, the state of output electronic device 300 may change from an unavailable or available state to a ready state. Backend system 100 can determine that the state of output electronic device 300 has changed from unavailable or available to ready, and can begin sending content to output electronic device 300 instead of sending it to voice activation of electronic device 10. In some implementations, upon determining that the state of output electronic device 300 has changed from unavailable or available to ready, backend system 100 can generate a prompt asking individual 2 whether they want to output content via output electronic device 300 instead of voice activation of electronic device 10. In some implementations, the prompt can be displayed as a user interface on output electronic device 300. Individual 2 can interact with the user interface to indicate whether content should be output via output electronic device 300. Furthermore, in some implementations, the prompt can be output as audio via voice activation of electronic device 10. Individual 2 can provide a spoken response to indicate whether content should be output via output electronic device 300. The voice-activated electronic device 10 can send audio data representing the response of the individual 2 to the back-end system 100, and the back-end system 100 can determine the nature of the individual 2's response and route the content accordingly (e.g., the back-end system 100 can send an instruction to output the content via the output electronic device 300 in response to the individual 2 indicating that he / she wants to send the content).

[0088] In some implementations, while content is being output via voice activation of electronic device 10, the state of output electronic device 300 may change from an unavailable state to an available state. Backend system 100 can determine that the state of output electronic device 300 has changed from unavailable to available and can generate an audio prompt asking individual 2 whether they want to output content via output electronic device 300 instead of voice activation of electronic device 10. Backend system 100 can send the audio prompt to voice activation electronic device 10 as audio output. Individual 2 can provide a spoken response indicating whether content should be output via output electronic device 300. Voice activation electronic device 10 can send audio data representing individual 2's response to backend system 100, and backend system 100 can determine the nature of individual 2's response and route content accordingly. If the response indicates that individual 2 wants to send content to output electronic device 300, then backend system 100 can send an instruction to output electronic device 300 to change it to a ready state and output content. Backend system 100 can also send an instruction to voice activation electronic device 10 to stop outputting content.

[0089] Furthermore, in some implementations, individual 2 can direct content to voice-activated electronic device 10 or output electronic device 300 by specifying a target device in said command 4. For example, individual 2 can say to voice-activated electronic device 10, "Alexa, play my music playlist on my TV." Backend system 100 can use STT and NLU processing to determine that individual 2 has specified a target device for the content and can send the requested content to output electronic device 300. For example, the type of output electronic device 300 (e.g., TV) can be stored in data structure 102 in content routing module 270. NLU module 260 can determine that a target device is specified in command 4 and can query content routing module 270 to request information about whether voice-activated electronic device 10 is associated with output electronic device 300 and information about the type of output electronic device 300. NLU module 260 can use NLU functionality and information about the type of output electronic device 300 to determine whether the probability that individual 2 has requested output electronic device 300 is higher than a predetermined threshold probability. For example, if command 4 includes a request to play content on "My TV" and output electronic device 300 is a television, then output electronic device 300 is likely the requested device. However, if command 4 includes a request to play content on "My Speaker System" and output electronic device 300 is a television, then output electronic device 300 is unlikely to be the requested device. If backend system 100 cannot determine the requested target device, then backend system 100 may generate an audio prompt to request clarification from the individual or to notify the individual that the requested target device cannot be found. Backend system 100 may send the audio prompt to voice-activated electronic device 10 as audio output to individual 2. Individual 2 may similarly request that content be sent to voice-activated electronic device 10 by specifying voice-activated electronic device 10 as the target device. In some embodiments, backend system 100 sends content to the explicitly requested target device, even if the content type may fall into a category that would normally be sent to a different device.

[0090] Figure 5This is an illustrative diagram illustrating an instance of resolving ambiguous requests for content. In some embodiments of the disclosed concept, the command 4 may include an ambiguous request for content. For example, individual 2 may say, “Alexa, play Footloose.” The request may correspond to either the movie Footloose or the movie soundtrack Footloose. To resolve ambiguity, backend system 100 may generate responsive audio data representing a first audio message 12 (auditory message) clarifying the request for content. For example, the first audio message 12 may be “Movie or soundtrack?” Individual 2 may provide a spoken response to voice-activated electronic device 10. Backend system 100 may then use STT to analyze the spoken response to determine which words were spoken, followed by NLU processing to determine the meaning of the spoken words and thus determine which specific content individual 2 is referring to. In determining which specific content individual 2 is referring to, backend system 100 may route the selected content to the appropriate one of output electronic device 300 or voice-activated electronic device 10.

[0091] In some implementations, backend system 100 may use other methods to resolve ambiguous requests for content. For example, if the command 4 includes a request to play content on a specific device, it can help determine which content individual 2 has requested. For instance, if the command 4 is “Alexa, play Book Thief on my TV,” then backend system 100 can determine that individual 2 is requesting the movie *The Book Thief*, not the book *The Book Thief*. In some implementations, backend system 100 may also use other information included in the command 4 to help determine which content individual 2 has requested. For instance, if the command 4 is “Alexa, read Book Thief,” then backend system 100 can determine that individual 2 is requesting the book *The Book Thief*, not the movie *The Book Thief*, based on the use of the term “read.”

[0092] Figure 6This is an illustrative diagram of a system that can be used to establish an association between a voice-activated electronic device 10 and an output electronic device 300. In some embodiments, the settings of the voice-activated electronic device 10 can be accessed through a user electronic device 500. The user electronic device 500 can be, for example, a mobile phone, computer, tablet computer, or other type of electronic device. The user electronic device 500 can communicate with the back-end system 100 and may include an application or other program that allows the user to establish an association between the voice-activated electronic device 10 and the output electronic device 300. For example, the user can identify and / or select the output electronic device 300 as a device associated with the voice-activated electronic device 10. After the association between the voice-activated electronic device 10 and the output electronic device 300 has been established through the user electronic device 500, the association can be stored on the back-end system 100 in, for example, a data structure 102 (see example...). Figure 1 In some implementations, information about the association (such as information identifying the output electronic device 300) may be stored in data structure 102, which may be stored in content routing module 207 (see, for example...). Figure 3 Additional metadata (e.g., flags) indicating that the voice-activated electronic device 10 is associated with another device can be stored, for example, in the user account module 268. This additional metadata stored in the user account module 268 can trigger the backend system 100 to consult the content routing module 270 to identify the output electronic device 300 and determine whether the requested content should be output by the output electronic device 300.

[0093] User electronic device 500 can be used to change settings associated with how backend system 100 determines where content should be routed. For example, user electronic device 500 can be used to specify the types of content that should always be played on output electronic device 300, and the types of content that should be played on output electronic device 300 if it is active. Those skilled in the art will understand that various other settings associated with voice-activated electronic device 10, backend system 100, and output electronic device 300 can be set via user electronic device 500. In some embodiments, it is output electronic device 300, rather than user electronic device 500, that can be used to associate voice-activated electronic device 10 with output electronic device 300.

[0094] Once the output electronic device 300 is associated with the voice-activated electronic device 10, the association can be terminated or re-established via a command received by the voice-activated electronic device 10. For example, a command such as "Disconnect my TV" can be used to terminate the association between the voice-activated electronic device 10 and the output electronic device 300, and a command such as "Reconnect my TV" can be used to re-establish the association between the voice-activated electronic device 10 and the output electronic device 300.

[0095] Figure 7 This is an illustrative diagram showing the linking of two exemplary devices according to various embodiments. In some embodiments, electronic device 702 can correspond to any electronic device or system. Various types of electronic devices can include, but are not limited to: desktop computers, mobile computers (e.g., laptops, ultrabooks), mobile phones, smartphones, tablets, televisions, set-top boxes, smart TVs, watches, bracelets, displays, personal digital assistants (“PDAs”), smart furniture, smart home devices, smart vehicles, smart transportation devices, and / or smart accessories. In some embodiments, electronic device 702 can be structurally relatively simple or basic, such that it may not provide one or more mechanical input options (e.g., keyboard, mouse, touchpad) or one or more touch input devices (e.g., touchscreen, buttons). However, in some embodiments, electronic device 702 may also correspond to a network of devices.

[0096] Electronic device 702 may have a display screen 704. The display screen 704 can display content on electronic device 702. In some embodiments, electronic device 702 may have one or more processors, memory, communication circuitry, and input / output interfaces. The one or more processors of electronic device 702 may be similar to... Figure 3 One or more processors 202, and the same description applies. The memory of electronic device 702 may be similar to... Figure 3 The storage device / memory 204, and the same description applies. The communication circuit of the electronic device 702 can be similar to... Figure 3 The communication circuit 206, and the same description applies. The input / output interface of the electronic device 702 can be similar to... Figure 3 The input / output interface 212, and the same description applies. Additionally, the electronic device 702 may have one or more microphones. The one or more microphones of the electronic device 702 may be similar to... Figure 3 One or more microphones 208, and the same description applies. Furthermore, the electronic device 702 may have one or more speakers. The one or more speakers of the electronic device 702 may be similar to... Figure 3 The speaker 210, and the same description applies.

[0097] In one exemplary implementation, an individual may want to link two or more devices together by selecting a device to receive a command and another device to respond to the output of the received command. While only one device is shown for each option (receive command and output response), those skilled in the art will recognize that any number of devices can be linked. In some implementations, an input device 706 can be selected. To select an input device 706, an electronic device 702 can search for devices capable of receiving input. In some implementations, the electronic device 702 can use HTTP to search for suitable devices via a web browser. Various other communication protocols can be used to facilitate communication between the voice-activated electronic device 10 and the backend system 100, including but not limited to: Wi-Fi (e.g., 802.11 protocol), Bluetooth®, radio frequency systems (e.g., 900 MHz, 1.4 GHz, and 5.6 GHz communication systems), cellular networks (e.g., GSM, AMPS, GPRS, CDMA, EV-DO, EDGE, 3GSM, DECT, IS-136 / TDMA, iDen, LTE, or any other suitable cellular network protocol), infrared, BitTorrent, FTP, RTP, RTSP, SSH, and / or VoIP. Once the electronic device 702 has located the appropriate input device, it can list the device on the display screen 704 for the individual to select from. Figure 7 In the example shown, the selected device is the first device 712. Once the first device 712 is selected as the input device, the electronic device 702 can store the identifier of the first device 712. In some embodiments, the first device 712 may be similar to the voice-activated electronic device 10, and the same description applies.

[0098] To select an output device 708, electronic device 702 can search for devices capable of outputting content. Similar to searching for an input device 706, electronic device 702 can use HTTP to search for suitable devices via a web browser. Various additional communication protocols can be used to facilitate communication between voice-activated electronic device 10 and backend system 100, including but not limited to: Wi-Fi (e.g., 802.11 protocol), Bluetooth®, radio frequency systems (e.g., 900 MHz, 1.4 GHz, and 5.6 GHz communication systems), cellular networks (e.g., GSM, AMPS, GPRS, CDMA, EV-DO, EDGE, 3GSM, DECT, IS-136 / TDMA, iDen, LTE, or any other suitable cellular network protocol), infrared, BitTorrent, FTP, RTP, RTSP, SSH, and / or VoIP. Once electronic device 702 has located a suitable output device, it can list the devices on display screen 704 for the individual to choose from. The listed devices can be based on the selected content option 710A. The following provides a more detailed description of content option 710A. Figure 7 In the example shown, the selected output device is the second device 714. Once the second device 714 is selected as the output device, the electronic device 702 can store the identifier of the second device 714.

[0099] An individual can also select the type of content to be sent to output device 708. In some embodiments, the individual can select from a drop-down menu. Content option 710A can include various options. In some embodiments, the first option can be an image file 710B. If selected, this content option can send any image file requested by the first device 712 to the second device 714. Image data can include any content containing visual information, including but not limited to videos, movies, photos, presentations, or any other visual display. For example, if an individual states to the first device 712, “Alexa, play a movie,” then the second device 714 will play the movie. In some embodiments, the second option can be an audio file 710C. Audio files can include any type of content containing audio data. If selected, this content option can send any audio file requested by the first device 712 to the second device 714. For example, if an individual states to the first device 712, “Alexa, play a song,” then the second device 714 will play the song. The third option, also known as more options 710D, can be any type of content. More options 710D can be specific to a particular request. For example, more options 710D could be a weather forecast. In this implementation, if a person says to the first device 712, "Alexa, give me the weather forecast," then the weather forecast will be output on the second device 714. As another example, more options 710D could be news information. In this implementation, if a person says to the first device 712, "Alexa, tell me the news," then the news will be output on the second device 714. In some implementations, multiple options can be selected. For example, an image file 710B and an audio file 710C can be selected. In this example, if the first device 712 receives a request for an audio file or an image file, then the content will be sent to the second device 714. Although Figure 7 Only a few types of content are shown, but those skilled in the art will recognize that this is for illustrative purposes only and that any type or any number of types can be selected in content option 710A.

[0100] In some implementations, electronic device 702 can communicate with backend system 100. If so, electronic device 702 can send a first device 712 identifier to backend system 100. Additionally, electronic device 702 can send a second device 714 identifier to the backend system. Furthermore, electronic device 702 can send a content option 710A identifier to the backend system. The identifiers can be used... Figure 3The user account module 268 is used for storage, and the same description applies. Then, the backend system 100 can store the link between the first device 712 and the second device 714, so that when the first device 712 requests the type of content selected in content option 710A, the requested content is sent to the second device 714. If the user does not select a content type under content option 710A, the backend system can send the content requested by the first device 712 to the second device 714, and the content can be output by the second device 714.

[0101] In some implementations, multiple input devices may exist for output device 708. In another implementation, multiple output devices may exist for input device 706. In yet another implementation, a primary output device and a secondary output device may exist for input device 706. In this implementation, a request for content that is to be routed from input device 706 to output device 708 will be routed to the primary output device. If the primary output device is unable to receive the content, then the content can be routed to the secondary output device.

[0102] Figure 8 This is an exemplary flowchart of a process 1000 for sending content to an associated device according to various embodiments. Process 1000 can be implemented, for example, in a backend system 100, and the same description applies thereto. In some embodiments, process 1000 may begin at step 1002. At step 1002, the backend system 100 may receive first audio data from a first electronic device. In some embodiments, the first electronic device of process 1000 may be... Figures 1 to 3 and Figures 5 to 6 The voice-activated electronic device 10, and the same description applies. The first audio data may represent a spoken command 4 from person 2 and may include requests for content, such as a request for a weather forecast. For example, if person 2 states, "Alexa, what's the weather forecast?", the voice-activated electronic device may record the stated wording and send the audio data to the backend system. The voice-activated electronic device may use one or more microphones on the voice-activated electronic device to receive the first audio data. The one or more microphones on the voice-activated device may be similar to... Figure 3 One or more microphones 208, and the same description applies.

[0103] At step 1004, the backend system 100 may determine that a user account exists associated with the first electronic device. In some embodiments, the backend system may receive an identifier associated with the voice-activated electronic device. This data may be presented in the form of a customer identifier, product number, IP address, GPS location, or any other suitable method of identifying the voice-activated electronic device. The backend system may then search for and identify the user account associated with the identifier. The user account may be any suitable number or identifier capable of identifying a user associated with the voice-activated electronic device.

[0104] In some implementations, once the backend system has identified the first user account associated with the first electronic device, it can find the stored associations between the two electronic devices. In some implementations, the backend system can discover that the first electronic device is an input device in the stored associations between the two electronic devices. The response to the received first audio data can be routed based on the stored associations. The following... Figure 9A The following is a further description of routing content based on stored associations.

[0105] At step 1006, the backend system can generate first text data representing the first audio data received from the voice-activated electronic device. The text data can be generated by performing STT functionality on the received first audio data. STT functionality can be used to determine individual words within the received first audio data. The STT functionality of process 1000 can be achieved by using... Figure 3 The ASR module 258 shown is used to complete this. Figure 3 The same disclosure applies here. More specifically, step 1006 can be accomplished using the STT module 266 within the ASR module 258, and the same disclosure applies. Continuing with the example, once the backend system receives audio data stating "Alexa, what's the weather forecast?", the ASR module 258 will perform the STT functionality on the audio data. This will create text data representing "Alexa, what's the weather forecast?".

[0106] At step 1008, the backend system 100 can determine the intent of the first text data. After the backend system has generated the first text data representing the first audio data, it will send the text data to the NLU for processing. The NLU receives the first text data to determine the intent of the first text data. The NLU described herein can be used... Figure 3 The natural language understanding function 260 is used to complete this task. Figure 3The same public disclosure applies here. Continuing with the example, the NLU receives text data representing the statement "Alexa, what's the weather forecast?". The NLU can recognize "Alexa" as a wake word and therefore is irrelevant to determining the intent of the received audio data. The NLU can then break down and analyze the wording or utterance "what's the weather forecast?". First, the NLU can analyze the verb "what's" in the utterance. This allows the NLU to better understand the intent of the utterance. Next, the NLU can split the remaining wording "the weather forecast is" into "weather forecast" and "is". This also allows the NLU to better understand the intent of the utterance. In the case of wording splitting and analysis, the NLU can then search the backend system against a list of possible requests, thereby assigning a confidence score to each request. As used herein, a confidence score can be any identifier that can be assigned, thereby allowing the system to rank possible data matches. The confidence score can then be compared with a predetermined threshold to determine whether possible intents match. The following is in the context of... Figure 8 The description provides a more detailed explanation of the confidence score and predetermined threshold. In step 1008, the NLU can determine that the intent of the first text data is to find a weather forecast. Because no location is stated in the first audio data, the NLU can also determine that the weather forecast should be within the geographic area of ​​the voice-activated electronic device. The geographic area of ​​the voice-activated electronic device can be derived from the data sent by the voice-activated electronic device.

[0107] If the NLU fails to find a request that matches or exceeds a predetermined threshold in the backend system's database, the backend system can generate apology text data. The backend system can then receive audio data representing the apology text data by performing STT functionality on it. The backend system can then send the audio data to a voice-activated electronic device. The voice-activated electronic device will then play the audio data on one or more of its speakers. For example, if the NLU cannot find a match suitable for the request, the voice-activated device might say, "Sorry, I didn't understand the request."

[0108] Alternatively, if the NLU finds more than one suitable match, the backend system can generate a confirmation. This confirmation helps the NLU make a decision among more than one suitable match. The following section discusses... Figure 9A , Figure 9B and Figure 9C This situation is described in more detail in the description.

[0109] At step 1010, the backend system determines that the second electronic device is also associated with the user account. After identifying the user account associated with the identifier, the backend system can then search for any other devices associated with the user account. The device can be any device capable of communicating with the cloud-based backend system. The device can be, but is not limited to, a television, computer, laptop, personal digital assistant (PDA), and any device that can connect to the internet or connect to another device via Bluetooth. While some devices have been listed, those skilled in the art will recognize that any device capable of connecting to another device can be used. Furthermore, the second electronic device associated with the user account in process 1000 can be… Figures 1 to 3 and Figures 5 to 6 The output electronic device 300. The disclosure of the output electronic device 300 also applies to the devices associated with the voice-activated electronic device herein.

[0110] At step 1012, the backend system determines that the response to the first audio data will be both an audio and a visual response. Continuing the example, when the NLU determines that the intent of the first text data is to determine a weather forecast, the backend system can access a weather category server. The weather category server can be similar to... Figure 3 The category server / skill module 262 or therein, and the same description applies here. The weather category server may have weather information related to the location of the voice-activated electronic device. Different data categories may exist in the memory of the weather category server. The different data categories described herein may be similar to... Figure 4 The categories shown, and the same descriptions apply. The weather category server's memory may contain text data representing responses to weather requests. Additionally, the weather category memory may contain video data in response to weather requests. If the weather category memory finds video data in response to a weather request, the backend system can check if the associated device is capable of displaying the video data. Because the backend system has already determined that the device is associated with an electronic device, it can look for both audio data and video data in response to first audio data from a voice-activated electronic device. For example, the audio data in response to a weather forecast request might contain statements about the day's high and low temperatures along with various other weather conditions. Visual data could be a five-day weather forecast that can be displayed on the screen of the associated device.

[0111] At step 1014, the backend system determines that a response will be sent to the first electronic device. Once responsive audio data is found, the backend system determines that a response to the first audio data will be sent to the first electronic device. In some embodiments, this determination is made so that the backend system is ready to send audio data to the first electronic device. Continuing with the example, the backend system now determines that a response to a weather forecast request will be sent to the first electronic device.

[0112] At step 1016, the backend system determines that a response will be sent to the second electronic device. If the backend system determines that the second electronic device can display visual data, then in some embodiments, at step 1016, the backend system determines that both an audio response and a video response will be sent to the second electronic device. In some embodiments, because the visual response will be played on the second electronic device, the audio response sent to the voice-activated electronic device may simply be a signal that the visual response will be displayed on the second electronic device. Continuing with the example, the backend system now determines that a response to a weather forecast request will be displayed on the associated device. Furthermore, the response from the voice-activated device may state, "Displaying your weather forecast on your TV."

[0113] At step 1018, the backend system receives second text data representing the first audio response. In some embodiments, the text data received by the backend system will come from a category server or a skill server. The category server or skill server can be connected to... Figure 3 The category server / skill module 262 is the same as or within it, and the same description applies. In some embodiments, the second text data may be a complete response to the first audio data received from the voice-activated electronic device. For example, in response to a request for news, the text data may contain the day's news. If visual data and a second electronic device capable of displaying the visual data are present, then in some embodiments, the second text may indicate where the response will be played. For example, in response to a news request, the text data may indicate that the response will be displayed on a television. Continuing with the weather forecast example, the weather category server will send text data representing the response to the backend system. The text data may contain text indicating that a weather forecast will be played on a television.

[0114] At step 1020, the backend system generates second audio data representing the second text data. Once text data has been received from the category server or skill server, that text data is converted into audio data. This data is converted into audio data by performing TTS functionality on the text data. The TTS functionality can be similar to... Figure 3The same description applies to the TTS module 264. Continuing with the weather forecast example, if the text data received by the weather category server includes a full audio response, then the audio data could state, "The Seattle weather forecast has a high of 72 degrees and a low of 55 degrees, with a possibility of showers." If the text data is merely a representation of visual data that will be played on a second electronic device, then the audio data could state, "The Seattle weather forecast is on your TV."

[0115] At step 1022, the backend system sends the second audio data to the first electronic device. The second audio data, created by performing TTS functionality on the second text data, is transmitted to the first electronic device. Once the second audio data is sent to the first electronic device, it is output by one or more speakers on the first electronic device. The one or more speakers are similar to... Figure 3 One or more speakers 210, and the same description applies. Continuing with the weather forecast example, if the audio data includes a full audio response, then the first electronic device can state, "The Seattle weather forecast has a high of 72 degrees and a low of 55 degrees, and there is a possibility of showers." If the audio data is merely a representation of visual data that will be played on a second electronic device, then the first electronic device can state, "The Seattle weather forecast is on your television."

[0116] At step 1024, the backend system receives third text data representing the second audio response. In some embodiments, the third text data received by the backend system will come from a category server or a skill server. The category server or skill server may be connected to... Figure 3 The category server / skill module 262 is the same as or within it, and the same description applies. In some embodiments, the third text data may be a complete response to the first audio data received from the first electronic device. For example, in response to a commuting request, the text data may contain a traffic report. If visual data is available and the second electronic device is capable of outputting visual data, then in some embodiments, the third text may be an indication of a response that will be displayed on the second electronic device. For example, in response to a news request, the text data may indicate that a response will be displayed on a television. Continuing with the weather forecast example, similar to the second text data, the backend system will receive third text data from the weather category server representing a response to the first audio data. The third text data may have text indicating that a weather forecast is being displayed on a television.

[0117] At step 1026, the backend system receives third audio data representing third text data. Once the third text data has been received from the category server or skill server, it is converted into audio data. This conversion is achieved by performing TTS functionality on the third text data. The TTS functionality can be similar to... Figure 3 The same description applies to the TTS module 264. Continuing with the weather forecast example, if the third text data received by the weather category server contains a full audio response, then the audio data can state, "The Seattle weather forecast has a high of 72 degrees and a low of 55 degrees, with a possibility of showers." If the third text data is merely a representation of visual data that will be played on a second electronic device, then the audio data can state, "This is the Seattle weather forecast."

[0118] At step 1028, the backend system receives image data representing the video response. As described herein, the image data can be any visual information, including but not limited to movies, videos, photographs, and presentations. Once the backend system determines that the second electronic device is capable of displaying video content, the backend system will then send visual data from the category server or skill server in response to the first audio data to the second electronic device. The category server or skill server may be connected to... Figure 3 The category server / skill module 262 is the same as or within it, and the same description applies. Video content can be part of a category of data within the category server or skill server. This category can be similar to... Figure 4 The video is played only on output device 402, and the same description applies. In some embodiments, output device may refer to a device capable of outputting video data, such as the associated device of process 1000 or... Figure 3 The output electronic device 300. The second electronic device of process 1000 can be similar to... Figure 3 The output electronic device 300, and the same description applies. Continuing with the weather forecast example, the backend system can receive visual data from the weather forecast.

[0119] At step 1030, the backend system sends third audio data to the second electronic device. The third audio data, created by performing TTS functionality on the third text data, is sent to the associated device. Once the third audio data is sent to the second electronic device, it is output by one or more speakers on the second electronic device. One or more speakers are similar to... Figure 3The speaker 314, and the same description applies. Continuing with the weather forecast example, if the audio data includes a full audio response, then the second electronic device can state, "The Seattle weather forecast has a high of 72 degrees and a low of 55 degrees, and there is a possibility of showers." If the audio data is merely a representation of visual data that will be played on the second electronic device, then one or more speakers of the second electronic device can play, "This is the Seattle weather forecast."

[0120] In step 1032, the backend system sends video data to the second electronic device. The video data received from the category server or skill server is transferred to the second electronic device. The second electronic device can then play or display the video data on the display screen of the associated device. The display screen of the second electronic device can be similar to... Figure 3 The same description applies to the display 312.

[0121] Figure 9A This is an exemplary flowchart of a process 1100 for routing content based on content type, according to various embodiments. Like process 1000, process 1100 can be implemented in, for example, backend system 100, and the same description applies thereto. In some embodiments, process 1100 may begin at step 1102. At step 1102, backend system 100 may receive first audio data from a voice-activated electronics device. Step 1102 may be similar to step 1002 of process 1000, and the same description applies. In some embodiments, the voice-activated electronics device of process 1100 may be… Figures 1 to 3 and Figures 5 to 6 The voice-activated electronic device 10, and the same description applies. The first audio data may represent the words spoken by the individual and may include requests. For example, if the individual states "Alexa, play Footloose," the voice-activated electronic device may record the stated words and send the audio data to the backend system. The voice-activated electronic device may use one or more microphones on the voice-activated electronic device to receive the first audio data. The one or more microphones on the voice-activated device may be similar to... Figure 3 One or more microphones 208, and the same description applies.

[0122] At step 1104, the backend system 100 may determine that a user account exists associated with the first electronic device. In some embodiments, step 1104 may be substantially similar to step 1004 of process 1000, and the same description applies. In some embodiments, the backend system may receive an identifier associated with the voice-activated electronic device. The backend system may then search for and identify the user account associated with the identifier. The user account may be any suitable number or identifier capable of identifying a user associated with the voice-activated electronic device.

[0123] At step 1106, the backend system can determine that the second electronic device is also associated with the user account. After identifying the user account associated with the identifier, the backend system can then search for any other devices associated with the user account. In some embodiments, the backend system can discover that the second electronic device is also associated with the user account. The second electronic device can be similar to the second electronic device in process 1000, and the same description applies. Additionally, the second electronic device associated with the user account described in process 1100 can be... Figures 1 to 3 and Figures 5 to 6 The output electronic device 300. The disclosure of the output electronic device 300 also applies to the devices associated with the voice-activated electronic device in process 1100. For example, a second electronic device also associated with a user account could be a television.

[0124] At step 1108, the backend system may generate first text data representing the first audio data received from the voice-activated electronic device. Step 1106 may be similar to step 1006 of process 1000, and the same description applies. The text data can be generated by performing STT functionality on the received first audio data. The STT functionality of process 1100 can be achieved by using... Figure 3 The ASR module 258 shown is used to complete this. Figure 3 ASR module 258 and Figure 3 The disclosure of STT module 266 applies here. Continuing with the example described, once the backend system receives audio data stating "Alexa, play Footloose," ASR module 258 can perform STT functionality on the audio data. This will create text data representing the received audio.

[0125] At step 1110, the backend system receives a first confidence score from the first domain. At step 1110, the backend system can utilize NLU functionality in a manner similar to step 1008 of process 1000, and the same description applies. The NLU described in process 1100 can be similar to... Figure 3The same description applies to Natural Language Understanding 260. The first domain can refer to any one or more servers located within or connected to the backend system. A domain can be similar to... Figure 3 The category server / skill module 262 or therein, and the same description applies. Similar to step 1008 of process 1000, the NLU can split the phrasing “Alexa, play Footloose”. The NLU can identify the wake word and focus on the verb and noun in the phrasing “Alexa, play Footloose”. The verb in the phrasing will be “play”. Play can refer to songs and movies, among other content types. While only songs and movies are disclosed herein, those skilled in the art will recognize that any content can be used. While playing may not narrow down the correct response, the NLU will pay attention to the noun to see if it is possible to narrow down the probability. The NLU can analyze “Footloose” by searching the category server or skill server to determine what content “Footloose” might be associated with. The song category server can send a confidence score to the NLU, indicating that the first audio request to play the song footloose is highly likely. The assigned confidence score can be a function of the probability that the response from the domain is the correct response. When the NLU searches for a match, it assigns a value to the possible response.

[0126] Because NLU can determine that the utterance "play Footloose" highly likely refers to playing the song Footloose on the first electronic device, it means the confidence score for the song Footloose may exceed a predetermined threshold. To determine which response is correct, a predetermined threshold can be set. This threshold ensures that incorrect responses to the utterance are not sent back to the voice-activated electronic device. Additionally, the predetermined threshold helps ensure that multiple irrelevant responses are not selected. This helps to obtain responses to utterances received by the first electronic device more quickly and accurately.

[0127] At step 1112, the backend system receives a second confidence score from the second domain. At step 1112, the backend system may use NLU functionality in a manner similar to step 1008 of process 1000, and the same description applies. The second domain may refer to any one or more servers located within or connected to the backend system. The second domain may be substantially similar to the first domain of step 1110, and the same description applies. The backend system may receive a second intent with a confidence score greater than a predetermined threshold. Continuing with the Footloose instance, the NLU may also receive a confidence score from the domain indicating that the first audio has requested a video response. If the backend system has confirmed that the second electronic device is capable of outputting video data, then the NLU may search only for video data. If the backend system determines that the second electronic device is capable of outputting video data, then the NLU may then receive confidence scores for both audio and video content from the category server or skill server. Because "play Footloose" can also refer to playing the Footloose movie, the NLU may receive a confidence score exceeding a predetermined threshold indicating that the utterance suggests a high probability of playing the Footloose movie on the associated device. NLU can also determine that trailers for the movie Footloose and other responses may have confidence scores greater than a predetermined threshold. However, for simplicity, only two different types of intent are shown in this example. Because in this implementation, NLU considers it highly likely that the utterance is requesting the song Footloose and the movie Footloose, NLU determines that more information is needed to accurately respond to the first audio data.

[0128] At step 1114, the backend system receives the query text data. Since it has already received two confidence scores indicating that either response is likely correct, the backend system can determine that more information is necessary. If the backend system determines that more information is necessary, it can then generate the query text. This query text can represent a question asking which response is correct. For example, the query text could represent a question asking whether "play Footloose" refers to the song Footloose or the movie Footloose.

[0129] In step 1116, the backend system generates query audio data representing the query text data. Once the backend system has received the query text data, it converts the query text data into audio data. The query text data is converted into audio data by performing TTS functionality on the query text data. TTS functionality can be similar to... Figure 3 The same description applies to the TTS module 264. For example, querying audio data could be stated as: "Do you mean the song Footloose or the movie Footloose?"

[0130] At step 1118, the backend system generates an answer instruction. Before sending the query, the backend system can generate an answer instruction for the first electronic device. The answer instruction can instruct the first electronic device to record the response to the query and send that response to the backend system. In some implementations, the answer instruction instructs the first electronic device to record without waiting for a wake word.

[0131] At step 1120, the backend system sends the query audio data to the first electronic device. The query audio data can be transmitted to the first electronic device by performing TTS functionality on the query text data. Once the query audio data is sent to the first electronic device, it is output by one or more speakers on the first electronic device. One or more speakers are similar to... Figure 3 One or more speakers 210, and the same description applies. For example, the first electronic device could play: "Do you mean the song Footloose or the movie Footloose?"

[0132] At step 1122, the backend system sends an answer instruction to the first electronic device. After sending the query, the backend system can send an answer instruction to the first electronic device. The answer instruction can instruct the first electronic device to record a response to the query and send audio data representing the response to the backend system.

[0133] At step 1124, the backend system receives second audio data from the first electronic device. In some embodiments, the first electronic device may receive second audio data representing a response to a query for audio data. For example, the second audio could be a "movie". As another example, the second audio could be a "song".

[0134] At step 1126, the backend system generates second text data representing the second audio data. Once the second audio data is received, it can then be converted into text data by performing the STT functionality on the second audio data. This can be similar to steps 1006 of process 1000 and 1106 of process 1100, and the same description applies. The STT functionality can be achieved by using... Figure 3 The ASR module 258 shown is used to complete this. Figure 3 ASR module 258 and Figure 3 The disclosure of STT module 266 applies here. Continuing with the example described, once the backend system receives audio data stating "movie" or "song," ASR module 258 can perform STT functionality on the audio data. This will create text data representing the received audio.

[0135] At step 1128, the backend system determines the intent of the second text data. Once the backend system generates the second text data, NLU can then analyze it. The described NLU can be similar to... Figure 3 Natural Language Understanding 260, and the same description applies. In this instance, NLU might simply be looking for nouns. Because NLU has already determined the verb is "play," but there's uncertainty about whether "play" refers to the movie Footloose or the song Footloose. Therefore, NLU can only analyze the nouns from the second text data representing the second audio data. At step 1130, NLU determines whether it's playing audio data or playing video data. NLU can determine that the second audio data requested the movie Footloose. This situation is... Figure 9B The process continues. NLU can also determine that the second audio data request was for the song Footloose. This situation occurs in... Figure 9C The process continues. Furthermore, the NLU can determine whether the second audio data is unresponsive or provides a negative response to the query. In this case, the NLU can signal to the backend system that the process should be completely stopped.

[0136] Figure 9B It is based on the succession of various implementation plans. Figure 9A An illustrative flowchart of the process is provided, in which content is routed to the associated device based on the content. At step 1132B, the backend system determines that the second electronic device will output video content. If the NLU determines that the intent of the second audio data is to play video data, then the NLU can determine that the target device for the video data is the second electronic device. In some embodiments, this may occur because the backend system determines that the second electronic device is capable of outputting video data. In some embodiments, the backend system may make this determination based on whether the first electronic device is capable of outputting video data. In some embodiments, the NLU may determine that the target device is the second electronic device because the first electronic device cannot output video data and the second electronic device can output video data. The capabilities of the first and second electronic devices may be stored on the backend system. Alternatively, capabilities can be determined by generating a request for capability information and sending it to the first and / or second electronic devices. The electronic devices may send a response to the information request, the response indicating what type of data the electronic device can output. In some embodiments, the backend system may send test content to the first and second electronic devices. In some embodiments, the test content may be specifically used to determine the capabilities of the first and second electronic devices and may not be output by the first and second electronic devices. For example, if the second electronic device is a television, then the information request or test content may indicate that the second electronic device can output video data. If the first electronic device is Figure 1If the voice-activated electronic device 10 responds to an information request or test content, it can indicate that the first electronic device cannot output video data. Furthermore, if the second audio data indicates a target device, the NLU can identify the target device. For example, the second audio data could state a response such as "Play a movie on my TV."

[0137] At step 1134B, the backend system determines that the user account has access to the video content. Once the backend system determines that the second electronic device will play the video content, it can search for the requested video content in categories accessible to the user account. The user account may be associated with an account that has access to multiple movies and songs. If the user account has access to multiple movies, then the user account will look for the requested video content among the accessible movies. In some embodiments, the user account will have access to the requested video content. In some embodiments, the user account will not have access to the requested movie. If the user account does not have access to the requested content, then the backend system can search for a preview of the requested content. Furthermore, if the user account does not have access to the requested content, then the backend system can receive a notification message stating that the content is unavailable. This notification message can then be converted into audio data by performing TTS functionality on the notification message. The audio data can then be sent to the first or second electronic device for output on one or more speakers on the first or second electronic device.

[0138] At step 1136B, the backend system generates a URL that allows the second electronic device to stream video content. The backend system can generate the URL once it determines that the user account has access to the video content. This URL allows the second electronic device to stream the video content requested by the first audio data and acknowledged by the second audio data. In some embodiments, once the backend system generates the URL, it can generate text representing an acknowledgment message. The acknowledgment message can signal to the first electronic device that it understands the second audio data. This text is then converted into audio by performing TTS functionality. The acknowledgment message can then be sent to the first electronic device. The first electronic device can then output the acknowledgment message using one or more speakers. For example, the first electronic device can state "Ok".

[0139] At step 1138B, the backend system sends the URL to the second electronic device. Then, the URL generated by the backend system can be sent from the backend system to the second electronic device. The video data can then be played by the speaker on the second electronic device and displayed on its screen. The speaker of the second electronic device can be similar to... Figure 3 The speaker 314, and the same description applies. The display screen of the second electronic device may be similar to... Figure 3The display 312, and the same description applies. In some embodiments, the backend system may generate text representing an acknowledgment message. This text is then converted into audio by performing TTS functionality. The acknowledgment message may be sent to a first electronic device. The first electronic device may then output the acknowledgment message using one or more speakers. For example, the first electronic device may state: “Your movie is starting on your television.”

[0140] Figure 9C It is based on the succession of various implementation plans. Figure 9A An illustrative flowchart of the process is provided, where content is routed to an electronic device based on its content. At step 1132C, the backend system determines that the first electronic device will output audio content. If the NLU determines that the intent of the second audio data is to play audio data, then the NLU can determine that the target device for the audio data is the first electronic device. This determination can be compared with... Figure 9B Step 1132B occurs in a similar manner, and the same description applies. In some embodiments, both the first electronic device and the second electronic device may be capable of outputting audio data. In this case, the back-end system can determine that the first electronic device is the target device because it is the default device for playing audio data. Furthermore, the back-end system may want more information to determine between devices. This information can be found in a similar manner to steps 1114 to 1128 of process 1100. In some embodiments, if the second audio data indicates a target device, then the NLU can determine the target device.

[0141] At step 1134C, the backend system determines that the user account has access to audio content. Once the backend system determines that the first electronic device will play a song, it can search for the requested audio content within the categories accessible to the user account. The user account may be associated with an account that has access to multiple movies and songs. If the user account has access to multiple songs, it will search for the requested audio content among those accessible songs. In some embodiments, the user account will have access to the requested song. In some embodiments, the user account will not have access to the requested song. If the user account does not have access to the requested content, the backend system can search for a preview of the requested content. Furthermore, if the user account does not have access to the requested content, the backend system can receive a notification message stating that the content is unavailable. This notification message can then be converted into audio data by performing TTS functionality on it. The audio data can then be sent to the first or second electronic device for output on one or more speakers on the first or second electronic device.

[0142] At step 1136C, the backend system generates a URL that allows the first electronic device to stream audio content. The backend system can generate the URL once it determines that the user account has access to the audio content. This URL allows the first electronic device to stream audio content requested by first audio data and acknowledged by second audio data. In some embodiments, once the backend system generates the URL, it can generate text representing an acknowledgment message. The acknowledgment message can signal to the first electronic device that it understands the second audio data. This text is then converted into audio by performing TTS functionality. The acknowledgment message can then be sent to the first electronic device. The first electronic device can then output the acknowledgment message using one or more speakers. For example, the first electronic device can state "Ok".

[0143] At step 1138C, the backend system sends the URL to the first electronic device. The URL generated by the backend system can then be sent from the backend system to the first electronic device. The first electronic device can then activate one or more microphones on the electronic device via voice to play or stream audio data. The one or more microphones on the first electronic device can be similar to... Figure 3 One or more microphones 208, and the same description applies. In some embodiments, the backend system may generate text representing an acknowledgment message. This text is then converted into audio by performing TTS functionality. The acknowledgment message may be sent to a first electronic device. The first electronic device may then output the acknowledgment message using one or more speakers. For example, a voice-activated electronic device may state: “Play the song Footloose.”

[0144] Figure 10 This is an exemplary flowchart of a process 1200 for receiving a request to change an output device, according to various embodiments; like process 1100, process 1200 can be implemented in, for example, backend system 100, and the same description applies thereto. In some embodiments, process 1200 may begin at step 1202. At step 1202, backend system 100 may receive first audio data from a first electronic device. Step 1202 may be similar to step 1002 of process 1000, and the same description applies. In some embodiments, the first electronic device of process 1200 may be... Figures 1 to 3 and Figures 5 to 6The voice-activated electronic device 10, and the same description applies. The first audio data may represent the words spoken by the individual and may include requests. For example, if the individual states "Alexa, play Content," the voice-activated electronic device may record the stated words and send the audio data to the backend system. The voice-activated electronic device may use one or more microphones on the voice-activated electronic device to receive the first audio data. The one or more microphones on the voice-activated device may be similar to... Figure 3 One or more microphones 208, and the same description applies.

[0145] At step 1204, the backend system 100 determines that a user account exists associated with the first electronic device. Step 1204 may be similar to step 1004 of process 1000, and the same description applies. In some embodiments, such as in step 1004 of process 1000, the backend system may receive an identifier associated with the first electronic device. Once the identifier is received, the backend system can then identify the user account associated with the identifier.

[0146] At step 1206, the backend system generates first text data representing the first audio data. Step 1206 can be similar to step 1006 of process 1000, and the same description applies. The text data can be generated by performing STT functionality on the received first audio data. The STT functionality of process 1200 can be achieved by using... Figure 3 The ASR module 258 shown is used to complete this. Figure 3 ASR module 258 and Figure 3 The disclosure of STT module 266 applies here. Continuing with the example described, once the backend system receives audio data stating "Alexa, play Song," ASR module 258 can perform STT functionality on the audio data. This will create text data representing the received audio.

[0147] At step 1208, the backend system 100 determines the intent of the first text data. Similar to step 1008, after the backend system has generated the first text data representing the first audio data, it will send the text data to the NLU for processing. The NLU processing in step 1208 can be similar to the NLU processing in step 1008 of process 1000, and the same description applies. The NLU receives the first text data to determine the intent of the first text data. The NLU described herein can be used... Figure 3 The natural language understanding function 260 is used to complete this task. Figure 3The same disclosed content applies here. Continuing with the example described, the NLU receives text data representing audio data stating "Alexa, play Content". After recognizing the wake word, the NLU can break down and analyze the utterance "play Content". To allow the NLU to better understand the intent, the NLU will split off the verb "play" and analyze it. As with process 1100, play can refer to many types of content, such as songs or movies. In this embodiment, the noun is deterministic because it is "Content". The word content used in this embodiment refers to specific content that the NLU will understand. However, those skilled in the art will understand that if the term "Content" refers to the titles of two different songs, then a process similar to process 1100 can be used to narrow down the selection. Next, the NLU can then search the backend system against a list of possible requests, thereby configuring a confidence score for each request. The above is in the context of... Figure 8 The confidence score and predetermined threshold are explained in more detail in the description.

[0148] At step 1210, the backend system receives content in response to the first audio data. The content can be any content playable on the first electronic device. If the NLU determines that the first audio data signals that "Content" should be played, the backend system can receive content from a specific content category. This content category can be similar to... Figure 3 The category server / skill module 262 or therein, and the same description applies. If the backend system is unsure whether it has retrieved the correct content, it can generate text data representing a confirmation message. This text data can be converted into audio data using TTS functionality. Once the backend system receives the audio data, it can send it to a first electronic device, causing one or more speakers on the voice-activated electronic device to play an audio message. This message could be, for example, “Do you mean Content?” Once this confirmation message is sent, the content electronic device can receive responsive audio data. This audio data can be sent to the backend system, where STT functionality will be used to convert the audio data into text data. The NLU will then analyze the text data to determine if the backend system has the correct content. If the NLU determines that the response indicates the backend system does not have the correct content, the backend system can stop the process.

[0149] At step 1212, the backend system sends content to the first electronic device. Continuing the example, content data received from the category server or skill server is transferred to the first electronic device. The content can then be played by the first electronic device via one or more microphones on the first electronic device. The one or more microphones on the first electronic device can be similar to... Figure 3One or more microphones 208, and the same description applies. In some embodiments, the backend system may generate text representing an acknowledgment message. This text is then converted into audio by performing TTS functionality. The acknowledgment message may be sent to a first electronic device. The first electronic device may then output the acknowledgment message using one or more speakers. For example, the first electronic device may state: “Play Content.”

[0150] At step 1214, the backend system receives second audio data from the first electronic device. Step 1214 may be similar to step 1002 of process 1000, and the same description applies. The second audio data may represent spoken words and may include requests. For example, if an individual states, “Alexa, play Content on my TV,” then the first electronic device may record the stated words and send the audio data to the backend system. The second audio data may be recorded by the first electronic device using one or more microphones on the first electronic device.

[0151] At step 1216, the backend system generates second text data representing the second audio data. Step 1216 can be similar to step 1206 and step 1006 of process 1000, and the same description applies. The second text data can be received by performing STT functionality on the received second audio data. The STT functionality of process 1200 can be achieved by using... Figure 3 The ASR module 258 shown is used to complete this. Figure 3 ASR module 258 and Figure 3 The disclosure of STT module 266 applies here. Continuing with the example described, once the backend system receives the audio data stating "Alexa, play Song", ASR module 258 can perform STT.

[0152] At step 1218, the backend system determines that the second electronic device is also associated with the user account. After identifying the user account associated with the identifier, the backend system can then search for any other devices associated with the user account. Associated devices can be similar to those in process 1000, and the same description applies. Additionally, the second electronic device with a user account described in process 1200 can be... Figures 1 to 3 and Figures 5 to 6 The output electronic device 300. The disclosure of the output electronic device 300 also applies to the second electronic device in process 1200.

[0153] At step 1220, the backend system determines that the intent of the second text data is to request content on a second electronic device. The NLU can analyze the second text data representing the second audio data and determine the target device in the second text data. In some embodiments, the target device may be a second electronic device. The determination of the target device may be similar to steps 1132B and 1132C, and the same description applies.

[0154] At step 1222, the backend system determines that the second content and the first content are the same. Similar to step 1008, after the backend system has generated second text data representing the second audio data, the text data will be transferred to the NLU for processing. The NLU processing in step 1218 can be similar to the NLU processing from both step 1208 and step 1008 of process 1000, and the same description applies. The NLU receives the first text data to determine the intent of the first text data. After undergoing a process similar to step 1208, the NLU can split the verbs and nouns of the second text data and analyze them. Then, the NLU can search the backend system against a list of possible requests to assign a confidence score to each request. The above is in the context of... Figure 8 The description explains the confidence score and predetermined threshold in more detail. The NLU can then compare the first and second requested content to determine if a perfect match exists. This is done by comparing the analyzed second text with the analyzed first text and creating a confidence score. If the confidence score exceeds the predetermined threshold, the backend system can determine that the content is identical. For example, the NLU can determine that the second text data's mention of "Content" is the same as that mentioned in the first text data. In some implementations, the backend system can determine that confirmation is necessary. If so, the backend system can generate confirmation text. This confirmation text might represent a question asking whether the song should be transmitted to a second electronic device. For example, the confirmation text might represent a question asking "Do you want to play Content on your TV?" This confirmation text data is then converted into confirmation audio data by performing TTS functionality on the confirmation text data. Once received by the backend system, the confirmation audio can be sent to the first electronic device, causing it to be played by one or more speakers on the first electronic device. For example, the first electronic device might state: "Do you want to play Content on your TV?"

[0155] The first electronic device can then receive responsive audio in response to the acknowledgment audio. The responsive audio can then be transmitted to the backend system. As in step 1206, the responsive audio will then be converted into text by performing STT functionality on the responsive audio. Once the backend system receives the text representing the responsive audio, it will send the text to the NLU for analysis. The NLU will determine whether the response is positive or negative. If the response is positive, the process continues to step 1224 below. A positive response could be, for example, "yes". If the response is negative, the process can stop and the content can be played on the first electronic device. A negative response could be, for example, "no".

[0156] At step 1224, the backend system determines to generate a stop command. The stop command can be used to stop the content being played by the first electronic device. The stop command instructs the first electronic device to stop playing the content.

[0157] At step 1226, the backend system sends a stop command to the first electronic device. Once the backend system has generated the stop command, it can then send it to the first device to stop content playback. The voice-activated electronic device will receive the command and stop the content. In some embodiments, the backend system may generate text representing a notification message. The purpose of the notification message may be to inform an individual that content playback will continue on the associated device. The notification text will be converted into notification audio by performing TTS functionality on the notification text. Once the backend system has received the notification audio, it transmits the notification audio to the first electronic device, causing it to be played by one or more speakers on the first electronic device. For example, the voice-activated electronic device may state: "Content will be played on your TV." In some embodiments, the notification audio may be played by one or more speakers on a second electronic device. In this embodiment, instead of sending it to the first electronic device, or in addition to sending it to the first electronic device, the notification audio will be sent to the second electronic device.

[0158] At step 1228, the backend system receives responsive content to the second audio data. Similar to step 1210, the content can be anything playable on the second electronic device. In some embodiments, the content can be a Content. Those skilled in the art will recognize that the use of Content is merely exemplary. The backend system can receive the same content played on the first electronic device.

[0159] At step 1230, the backend system sends the second content to the second electronic device. The second electronic device can then play the second content data via one or more microphones on it. The one or more microphones on the second electronic device can be similar to... Figure 3 The speaker 314, and the same description applies. In some embodiments, the backend system may generate text representing an acknowledgment message. This text is then converted into audio by performing TTS functionality. The acknowledgment message may be sent to a second electronic device. The second electronic device may then output the acknowledgment message using one or more speakers. For example, a voice-activated electronic device may state, “Play Content.” In some embodiments, the acknowledgment audio may be played by one or more speakers on the first electronic device. In this embodiment, instead of sending it to the second electronic device, the acknowledgment audio will be sent to the first electronic device.

[0160] Figure 11A This is an exemplary flowchart of process 1300 for routing content based on the state of associated devices, according to various embodiments. Like process 1300, process 1300 can be implemented in, for example, backend system 100, and the same description applies thereto. Those skilled in the art will recognize that in some embodiments, steps within process 1300 can be rearranged or omitted. In some embodiments, process 1300 may begin at step 1302. At step 1302, backend system 100 may receive first audio data from a first electronic device. Step 1202 may be similar to step 1002 of process 1000, and the same description applies. In some embodiments, the first electronic device of process 1300 may be… Figures 1 to 3 and Figures 5 to 6 The voice-activated electronic device 10, and the same description applies. The first audio data may represent spoken words and may include requests. For example, if a person states, "Alexa, play a song on the TV," then the first electronic device may record the stated words and send the audio data to a backend system. The first audio data may be received by the first electronic device using one or more microphones on the first electronic device. The one or more microphones on the first electronic device may be similar to... Figure 3 One or more microphones 208, and the same description applies.

[0161] At step 1304, the backend system determines that a user account exists associated with the first electronic device. Step 1304 may be similar to step 1004 of process 1000, and the same description applies. In some embodiments, such as in step 1004 of process 1000, the backend system may receive an identifier associated with a voice-activated electronic device. Once the identifier is received, the backend system can then identify the user account associated with the identifier. After identifying the user account associated with the identifier, the backend system can then search for any other devices associated with the user account.

[0162] At step 1306, the backend system can determine that the first audio data originates from an input device within the stored association. The stored association can be stored on a user account. In some implementations, the stored association can exist on the backend system. The stored association can be similar to... Figure 7 The associations shown are identical, and the same description applies. For example, a backend system can determine that the first audio data comes from a voice-activated electronic device. Once the source of the audio is determined, the backend system can identify the voice-activated electronic device as an input device within a stored association. A stored association could be an association between a voice-activated electronic device and a television. In one instance, a stored association could have a voice-activated electronic device as an input device and a television as an output device. Furthermore, there may be stored content preferences. For example, a stored content preference could be a preference for a song. If this is the case, then a request for a song from the voice-activated electronic device will be output to the television.

[0163] At step 1308, the backend system can determine the content preferences and output devices in the stored associations. In some embodiments, an association may have an input device and an output device. Once it is determined that the first audio comes from the input device in the stored association, the backend system can determine what the output device is. Furthermore, the backend system can determine what the content preferences are (if any). For example, the stored association could be an association between a voice-activated electronic device and a television. The voice-activated electronic device can be the input device. The television can be the output device. Additionally, there may be content preferences. If so, then the content preferences can determine whether audio data received from the input device triggers the association. For example, if the content preference is an audiobook (which would mean that any audio received from the voice-activated electronic device at any time requests an audiobook), then the audiobook will be played on the television. In some embodiments, this step can be omitted.

[0164] At step 1310, the backend system can generate first text data representing the first audio data. Step 1310 can be similar to step 1006 of process 1000, and the same description applies. The text data can be generated by performing STT functionality on the received first audio data. The STT functionality of process 1300 can be achieved by using... Figure 3 The ASR module 258 shown is used to complete this. Figure 3 ASR module 258 and Figure 3 The disclosure of STT module 266 applies here. Continuing with the example described, once the backend system receives audio data stating "Alexa, play Song on TV," ASR module 258 can perform STT functionality on the audio data. This will create text data representing the received audio.

[0165] At step 1312, the backend system 100 can determine the intent of the first text data. Similar to step 1008, after the backend system has generated the first text data representing the first audio data, the text data will be sent to the NLU for processing. The NLU processing in step 1308 can be similar to the NLU processing in step 1008 of process 1000, and the same description applies. The NLU receives the first text data to determine the intent of the first text data. The NLU described herein can be used... Figure 3 The natural language understanding function 260 is used to complete this task. Figure 3 The same publicly available information applies here. Continuing with the example described, the NLU receives text data representing audio data stating "Alexa, play Song". After recognizing the wake word, the NLU can split the utterances "play Song" and "on TV" and analyze them. To allow the NLU to better understand the intent, the NLU will split the verb "play" and analyze it. As in process 1100, playing can refer to many content types, such as songs or movies. In this implementation, the noun is deterministic because it is "Song". The word "song" used in this implementation refers to a specific movie that the NLU will understand. If "Song" is not deterministic, then the backend system can use a process similar to process 1100 to narrow down the intent of the utterance. Next, the NLU can then search the backend system against a list of possible requests, thereby configuring a confidence score for each request. The above is in the context of... Figure 8 The description explains the confidence score and predetermined threshold in more detail. NLU can determine that the intent of the first text data is a request to play the song.

[0166] At step 1314, the backend system can determine that the type of the requested content is the same type as the content stored in the association. In some embodiments, the stored associations may have content preferences. Once the backend system determines that the received audio comes from an input device within the association, it can check whether a content preference exists. If a content preference exists, the backend system can attempt to match the requested content type with the stored content preference. For example, when the NLU has determined the intent of the first text data, the NLU can determine the type of the requested content. After determining the type of the requested content, the NLU can attempt to match the type of the requested content with the stored content preference within the association. If the requested content type matches the stored preference, the NLU will know where to send the content. For example, if the stored preference is a song, the NLU will attempt to match the requested content type with a song. Continuing the example above, since the requested content is a song, the NLU will know that the target device will be the output device in the association. In this embodiment, since the output device is a television, the song requested by the input device will be played on the television. In some embodiments, this step can be omitted.

[0167] In some implementations, the type of requested content will not match the content preference. If this is the case, the backend system can operate in a manner similar to processes 1000 and 1100. In some implementations, there is no content preference. If this is the case, content can be routed to the output device based on whether the output device can output the requested content. If the output device cannot output the requested content, then the input device or any other associated device can output the requested content.

[0168] At step 1316, the backend system determines whether the second electronic device is ready, available, or unavailable. In some implementations, "ready, available, or unavailable" may be referred to as a functional state. Once the backend system determines that content will be routed to the second electronic device due to the association, it can determine whether the second electronic device can receive the content. In some implementations, if an association exists, the state of the second electronic device can be stored in the association. Furthermore, in some implementations, the backend system can send a status request to the second electronic device. The status request can originate from... Figure 3The content routing module 270, and the same description applies. A status request sent by the backend system can determine the status of the associated device. The status of the associated device helps determine whether content can be routed to the associated device. In some embodiments, three statuses exist: unavailable, ready, and available. Although only three statuses are disclosed, those skilled in the art will recognize that any number of statuses can be used and can effectively determine whether content can be routed to the associated device. The backend system can send a simulation test or any other suitable means to determine whether a second electronic device can output the requested content. The following... Figures 11B to 11D The description provides a more detailed account of each state.

[0169] Figure 11B It is based on the succession of various implementation plans. Figure 11A An illustrative flowchart of the process, where the state of the associated devices is ready. Continue Figure 11A In process 1300, at step 1318B, the backend system determines that the output device is in a ready state. In some embodiments, the ready state of the output device can be determined by a response to a status request sent from the second electronic device to the backend system. This response may be presented in the form of metadata. In some embodiments, the ready state of the output device can be determined by a saved status update sent by the output device to the backend system. This saved status update may be sent periodically or only when the output device becomes ready. In some embodiments, a test may have been sent to the output device. This test may have been successfully run, thereby determining that the second electronic device is capable of outputting the requested content.

[0170] At step 1320B, the backend system receives content in response to the first audio data. The content can be anything playable on either the first or second electronic device. In some embodiments, the content can be a song. If the NLU determines that the first audio data signals that "Song" should be played, the backend system can receive audio data from the song category. The song category can be similar to... Figure 3The category server / skill module 262 or therein, and the same description applies. If the backend system is unsure whether it has retrieved the correct song, it can generate text data representing a confirmation message. This text data can be converted into audio data using TTS functionality. Additionally, the backend system can receive answer instructions. These answer instructions can be similar to those in process 1000, and the same description applies. Once the backend system receives the audio data, it can send it to a first electronic device, causing an audio message to be played by one or more speakers on the first electronic device. This message could be, for example, “Do you mean Song?” After the audio is played, the backend system will send an answer instruction to the first electronic device. Once this confirmation message and answer instruction have been sent, the first electronic device can receive responsive audio data. This audio data can be sent to the backend system, where STT functionality will be used to convert the audio data into text data. The NLU will then analyze the text data to determine if the backend system has the correct song. If the NLU determines that the response indicates the backend system does not have the correct song, the backend system can stop the process. In some implementations, this step can be omitted. In some implementations, the content can be stored locally.

[0171] At step 1322B, the backend system sends the content to the output device. Continuing with the Song instance, the audio data received from the category server or skill server is transferred to the output device. The output device in process 1300 can be similar to... Figure 3 The output electronics 300, and the same description applies. Audio data can then be played by one or more speakers of the output device. The one or more speakers of the output device can be similar to... Figure 3 The speaker 314, and the same description applies.

[0172] In some implementations, the backend system can determine which user accounts are authorized to access content. This can be correlated with... Figure 9B and Figure 9C Steps 1134B and 1134C are performed similarly, and the same description applies. The backend system can also generate URLs that allow a second electronic device to stream the received content. This can be correlated with... Figure 9B and Figure 9C Steps 1136B and 1136C are performed similarly, and the same description applies. Furthermore, the generated URL can be sent to a second electronic device, thereby allowing the second electronic device to stream the requested content. This can be done in conjunction with... Figure 9B and Figure 9C Steps 1138B and 1138C are performed similarly, and the same description applies.

[0173] At step 1324B, the backend system receives notification text data indicating that the output device is ready. In some embodiments, the backend system may generate text indicating that the output device is ready. This notification text can be used... Figure 3 The content routing module 270 generates the content, and the same description applies. For example, the text indicating a notification could state: "Your TV is ready."

[0174] In step 1326B, the backend system generates notification audio data representing the notification text data. Once the backend system receives the notification text data, it converts the notification text data into audio data. The notification text data is converted into audio data by performing TTS functionality on the notification text data. The TTS functionality can be similar to... Figure 3 The same description applies to the TTS module 264. For example, confirming audio data could be stated as: "Your TV is ready."

[0175] At step 1328B, the backend system sends notification audio data to the first electronic device. The notification audio data generated from the TTS is sent to the first electronic device. The first electronic device can then play the audio data using one or more microphones on it. The one or more microphones on the first electronic device can be similar to... Figure 3 One or more microphones 208, and the same description applies. For example, a voice-activated electronic device may state: “Your television is ready.”

[0176] Figure 11C It is based on the succession of various implementation plans. Figure 11A An illustrative flowchart of the process, where the status of the associated devices is available. Continue Figure 11A In process 1300, at step 1318C, the backend system determines that the output device is available. In some embodiments, the availability status of the output device can be determined by a response to a status request sent from the output device. This response may be presented in the form of metadata. In some embodiments, the availability status of the output device can be determined by a saved status update sent to the backend system by a second electronic device. This saved status update may be sent periodically or only when the output device becomes available.

[0177] At step 1320C, the backend system generates an instruction to change the state of the output device. In response to determining that the output device is available, the backend system may generate an instruction to change the state of the output device from an available state to a ready state. While an available state may be an indication that the output device is powered on, it may not necessarily allow the output device to play any content. In some implementations, the output device may be in an available state because it is already playing content. If this is the case, the generated instruction may include an instruction to stop playing content.

[0178] At step 1322C, the backend system sends a command to the output device. Once the command has been generated, the backend system can send the generated command to change the state of the output device from an available state to a ready state, thereby allowing content to be transmitted and played by the output device. In some embodiments, once the output device changes its state, an acknowledgment notification can be sent from the output device to the backend system. This notification confirms that the output device is in a ready state and can receive and output content.

[0179] At step 1324C, the backend system receives content in response to the first audio data. Step 1324C can be similar to step 1320B, and the same description applies. In some embodiments, the content can be a movie. If the NLU determines that the content signals that a "movie" should be played, then the backend system can receive video data from the movie category. The movie category can be similar to... Figure 3 The category server / skill module 262 or therein, and the same description applies. In some implementations, this step may be omitted.

[0180] At step 1326C, the backend system sends the content to the output device. Continuing the example, the video data received from the category server or skill server is transferred to the second electronic device. The second electronic device in process 1300 can be similar to... Figure 3 The output electronics 300, and the same description applies. Video data can then be played by one or more speakers and a display screen of the output device. The one or more speakers of the output device can be similar to... Figure 3 The speaker 314, and the same description applies. The display of the output device may be similar to... Figure 3 The same description applies to the display 312.

[0181] In some implementations, the backend system can determine which user accounts are authorized to access content. This can be correlated with... Figure 9B and Figure 9C Steps 1134B and 1134C are performed similarly, and the same description applies. The backend system can also generate URLs that allow a second electronic device to stream the received content. This can be correlated with... Figure 9B and Figure 9C Steps 1136B and 1136C are performed similarly, and the same description applies. Furthermore, the generated URL can be sent to a second electronic device, thereby allowing the second electronic device to stream the requested content. This can be done in conjunction with... Figure 9B and Figure 9C Steps 1138B and 1138C are performed similarly, and the same description applies.

[0182] At step 1328C, the backend system receives notification text data indicating that the output device is ready. This step can be similar to... Figure 11B Step 1324B, and the same description applies. In some embodiments, the backend system may receive a text notification message indicating that a second electronic device is ready. This notification text can use... Figure 3 The content routing module 270 generates the content, and the same description applies. For example, the text indicating a notification could state: "Your TV is ready."

[0183] At step 1330C, the backend system generates notification audio data representing the notification text data. This step can be similar to... Figure 11B Step 1326B, and the same description applies. Once the backend system receives the notification text data, it converts the notification text data into audio data. The notification text data is converted into audio data by performing TTS functionality on the notification text data. The TTS functionality can be similar to... Figure 3 The same description applies to the TTS module 264. For example, confirming audio data could be stated as: "Your TV is ready."

[0184] At step 1332C, the backend system sends notification audio data representing the notification to the first electronic device. This step can be similar to... Figure 11B Step 1328B, and the same description applies. The notification audio data generated via TTS is sent to the first electronic device. The audio data can then be played by the first electronic device through one or more microphones on the first electronic device. The one or more microphones on the first electronic device can be similar to... Figure 3 One or more microphones 208, and the same description applies. For example, the first device could state: “Your television is ready.”

[0185] Figure 11D It is based on the succession of various implementation plans. Figure 11A An illustrative flowchart of the process, where the associated device is in an unavailable state. Continue Figure 11AIn process 1300, at step 1318D, the backend system determines that the output device is unavailable. In some implementations, this can be achieved by not receiving a response to the status request for a predetermined amount of time. For example, if the backend system sends a status request to the output device, it can wait two seconds to receive a response. If no response is received within this two-second window, the backend system can determine that the output device is unavailable. In some implementations, this can be achieved by having a state saved when the output device becomes unavailable. In some implementations, the state can be determined by sending a test to the output device. If the test fails, the backend system can determine that the output device is unavailable.

[0186] At step 1320D, the backend system receives notification text data indicating that the output device is unavailable. In some embodiments, the backend system may receive a text notification message indicating that the output device is unavailable. This notification text can be used... Figure 3 The content routing module 270 generates the content, and the same description applies. For example, the text indicating a notification could state: "Your TV is unavailable."

[0187] In step 1322D, the backend system generates notification audio data representing the notification text data. Once the backend system receives the notification text data, it converts the notification text data into audio data. The notification text data is converted into audio data by performing TTS functionality on the notification text data. The TTS functionality can be similar to... Figure 3 The same description applies to the TTS module 264. For example, confirming audio data could be stated as: "Your TV is unavailable."

[0188] At step 1324D, the backend system sends notification audio data to the first electronic device. The notification audio data generated via TTS is sent to the first electronic device. The first electronic device can then play the audio data through one or more microphones on the first electronic device. The one or more microphones on the first electronic device can be similar to... Figure 3One or more microphones 208, and the same description applies. For example, the first electronic device may play: “Your TV is unavailable.” In some embodiments, the process may stop here. However, in some embodiments, content may be played on the first electronic device. In some embodiments, the backend system may receive text indicating whether the user wants to play content on the first electronic device. In this embodiment, the received text is converted into audio data by performing TTS functionality on the text. The backend system may then generate an answer instruction. As described, the answer instruction may be similar to the answer instruction of process 1100, and the same description applies here. The audio is then sent to the first electronic device. For example, the first electronic device may play: “Would you like to play a song on the voice-activated electronic device?” After sending the audio, the backend system may send the answer instruction to the first electronic device, causing the first electronic device to record and send a response. The backend system may then receive the response. Once the backend system generates text indicating the responsive audio, the NLU then analyzes the text. The NLU will determine whether the response is a positive or negative response. If the response is positive, then content will be played on the first electronic device. A positive response may be, for example, “Yes.” If the response is negative, then the process can stop. A negative response could be, for example, "no".

[0189] At step 1326D, the backend system receives content in response to the first audio data. If the NLU determines that the content requested to be played on the second electronic device can also be played on the first electronic device, then the backend system may receive content in response to the first audio data. Figure 4 , Figure 9A , Figure 9B and Figure 9C The description provides a more detailed account of the process for determining which content can be played on which device, and these descriptions apply here. For example, if an individual requests to play a song on his / her television, the backend system can determine that this content can be played on either a second electronic device (i.e., the television) or a first electronic device. However, if a movie is requested, the backend system can determine that the content cannot be played on the first electronic device (i.e., if the first electronic device does not have a display), and the process can stop. Continuing with the Song example, the backend system can receive Songs from a Song category. This category can be similar to... Figure 3 The category server / skill module 262, and the same description applies. In some implementations, this step may be omitted.

[0190] At step 1328D, the backend system sends content to the first electronic device. Content received from the category server or skill server is transmitted to the first electronic device. The content can then be played by one or more speakers on the first electronic device. The one or more speakers on the first electronic device can be similar to... Figure 3 One or more speakers 210, and the same description applies. In some embodiments, the process may end here, and the playback of content on the first electronic device may be completed.

[0191] In some implementations, the backend system can determine which user accounts are authorized to access content. This can be correlated with... Figure 9B and Figure 9C Steps 1134B and 1134C are performed similarly, and the same description applies. The backend system can also generate a URL that allows the first electronic device to stream the received content. This can be correlated with... Figure 9B and Figure 9C Steps 1136B and 1136C are performed similarly, and the same description applies. Furthermore, the generated URL can be sent to the first electronic device, thereby allowing the first electronic device to stream the requested content. This can be correlated with... Figure 9B and Figure 9C Steps 1138B and 1138C are performed similarly, and the same description applies.

[0192] At step 1330D, the backend system determines that the output device is in a ready state. In some embodiments, the ready state of the output device can be determined by a response from the output device to a status request sent to the backend system. For example, if the output device has just been turned on, a response to a sent status report can be sent to the backend system, indicating that the output device is turned on and ready to receive content. In some embodiments, the ready state of the output device can be determined by a status update sent when the output device is turned on. This can occur each time the output device is turned on and can be stored by the backend system.

[0193] At step 1332D, the backend system receives text data indicating a prompt asking whether content should be moved to the output device. Once the backend system determines that the output device is ready, it can receive prompt text data indicating whether content should be moved to the output device. This notification can be received from... Figure 3 The content routing module 270 receives the message, and the same description applies. For example, the text indicating a notification could state: "Should the content be moved to the TV?"

[0194] In step 1334D, the backend system generates audio data representing the prompt text data. Once the backend system has received the prompt text data, it converts it into audio data. The prompt text data is converted into audio data by performing TTS functionality on the prompt text data. The TTS functionality can be similar to... Figure 3 The same description applies to the TTS module 264. For example, prompt audio data could state: "Should the content be moved to the TV?"

[0195] At step 1336D, the backend system generates an answer instruction. The answer instruction in process 1300 may be similar to the answer instruction in process 1100, and the same description applies. Before sending the prompt, the backend system may generate an answer instruction for the first electronic device. The answer instruction may instruct the first electronic device to record a response to the prompt and send that response to the backend system. In some embodiments, the answer instruction instructs the first electronic device to record without waiting for a wake word.

[0196] At step 1338D, the backend system sends the prompt audio data to the first electronic device. The prompt audio data received from the TTS is transmitted to the first electronic device. The first electronic device can then play the audio data via one or more microphones on its device. The one or more microphones on the first electronic device can be similar to... Figure 3 One or more microphones 208, and the same description applies. For example, a voice-activated electronic device could state: "Should the content be moved to the television?"

[0197] At step 1340D, the backend system sends an answer instruction to the first electronic device. The generated answer instruction can be sent to the first electronic device to instruct it to record the response to the prompt. The recorded response can then be sent back to the backend system. Step 1338D can be similar to... Figure 9A Step 1122 of process 1100, and the same description applies.

[0198] At step 1342D, the back-end system receives second audio data from the first electronic device. The second audio data may represent a response to a cues recorded by the first electronic device. The first electronic device may record the response using one or more of its microphones. The one or more microphones on the first electronic device may be similar to... Figure 3 One or more microphones 208, and the same description applies. For example, audio data could represent a response stating "Yes, content is playing on my TV".

[0199] At step 1344D, the backend system generates text data representing the second audio data. Step 1344D can be similar to step 1006 of process 1000, and the same description applies. The text data can be generated by performing STT functionality on the received second audio data. The STT functionality of process 1300 can be achieved by using... Figure 3 The ASR module 258 shown is used to complete this. Figure 3 ASR module 258 and Figure 3 The disclosure of STT module 266 applies here. Continuing with the example described, once the backend system receives audio data stating "Yes, content is playing on my TV," ASR module 258 can perform STT functionality on the audio data. This will create text data representing the received audio.

[0200] At step 1346D, the backend system determines the intent of the second audio data. Once the backend system generates text representing the second audio data, the NLU then analyzes the text. The NLU determines whether the response is affirmative or negative. If the response is affirmative, the content is played on the second electronic device. An affirmative response could be, for example, "yes". If the response is negative, the content remains on the voice-activated electronic device. A negative response could be, for example, "no". For example, if a person responds to the prompt with "no", the song can continue playing on the first electronic device.

[0201] At step 1348D, the backend system generates a stop command. Step 1346D can be similar to... Figure 10 Step 1224, and the same description applies. A stop command can be used to stop content being played by the first electronic device. A stop command can instruct the first electronic device to stop playing content currently being output on the first electronic device.

[0202] At step 1350D, the backend system sends a stop instruction to the first electronic device. The first electronic device can receive the instruction and stop the content. In some embodiments, the backend system can receive text representing a notification message. The purpose of the notification message may be to inform an individual that content will continue to play on the second electronic device. The notification text will be converted into notification audio by performing TTS functionality on the notification text. Once the backend system has generated the notification audio, it will transmit the notification audio to the first electronic device so that it is played by one or more speakers on the first electronic device. For example, a voice-activated electronic device may state: "Your content will be played on your TV." In some embodiments, the notification audio may be played by one or more speakers on the second electronic device. In this embodiment, instead of sending it to the first electronic device, or in addition to sending it to the second electronic device, the notification audio will be sent to the second electronic device.

[0203] At step 1352D, the backend system receives content in response to the first audio data. The backend system may receive the same content played on the first electronic device. In some embodiments, the backend system may generate text data representing an acknowledgment message. Step 1352D may be similar to step 1326D, and the same description applies. In some embodiments, this step may be omitted.

[0204] At step 1354D, the backend system sends the content to the output device. The output device can then play the content via one or more microphones. In some embodiments, the backend system may receive text representing an acknowledgment message. This text is then converted into audio by performing TTS functionality. The acknowledgment message can be sent to a first electronic device. The first electronic device can then output the acknowledgment message using one or more speakers. For example, a voice-activated electronic device may state, “Play Content.” In some embodiments, the acknowledgment audio can be played by one or more speakers on the output device. In this embodiment, instead of sending it to the first electronic device, the acknowledgment audio will be sent to the output device.

[0205] In some implementations, the backend system can determine which user accounts are authorized to access content. This can be correlated with... Figure 9B and Figure 9C Steps 1134B and 1134C are performed similarly, and the same description applies. The backend system can also generate URLs that allow a second electronic device to stream the received content. This can be correlated with... Figure 9B and Figure 9C Steps 1136B and 1136C are performed similarly, and the same description applies. Furthermore, the generated URL can be sent to a second electronic device, thereby allowing the second electronic device to stream the requested content. This can be done in conjunction with... Figure 9B and Figure 9C Steps 1138B and 1138C are performed similarly, and the same description applies.

[0206] Various embodiments of the present invention can be implemented in software, but can also be implemented in hardware, or a combination of hardware and software. The present invention can also be embodied as computer-readable code on a computer-readable medium. The computer-readable medium can be any data storage device that can subsequently be read by a computer system.

[0207] The above content can also be understood in accordance with the following terms.

[0208] 1. A method comprising:

[0209] Receive first audio data representing a first utterance at an electronic device, the first audio data being received from a voice-activated electronic device;

[0210] Receive a customer identifier associated with the voice-activated electronic device;

[0211] Identify the user account associated with the customer identifier;

[0212] By performing speech-to-text functionality on the first requested audio data, first text data representing the first requested audio data is generated;

[0213] Using the first text data, we determine that the first intent of the first utterance is for the target device to output information;

[0214] It was determined that the output device capable of presenting visual data was also associated with the user account;

[0215] It is determined that visual information responses to the first utterance are available;

[0216] Determine that the target device is the output device, such that the visual information response will be displayed on the display screen of the output device;

[0217] Determine whether to send the first audio response to the voice-activated electronic device;

[0218] Determine that the second audio response should be sent to the output device;

[0219] It is determined that the video response must also be sent to the output device;

[0220] Generate first response text data representing the first audio response;

[0221] By performing text-to-speech functionality on the first response text data, first audio data representing the first response text data is generated;

[0222] The first audio data is sent to the voice-activated electronic device, causing the first audio response to be played by the speaker of the voice-activated electronic device;

[0223] Generate second response text data in response to the second audio response to the first utterance, which includes receiving at least a portion of the second response text data from the application;

[0224] By performing the text-to-speech functionality on the second response text data, second audio data representing the second response text data is generated;

[0225] Generate responsive video data representing the video response to the first utterance;

[0226] Sending the second audio data to the output device, causing the second audio response to be played by the second speaker of the output device; and

[0227] The responsive video data is sent to a television, causing the video response to be played on the display screen of the output device.

[0228] 2. The method as described in Clause 1, further comprising:

[0229] Receive second request audio data representing a second utterance from the voice-activated electronic device;

[0230] By performing speech-to-text functionality on the second requested audio data, second text data representing the second requested audio data is generated;

[0231] The second intention of the second utterance is determined using the second text data in the following manner:

[0232] A first confidence score is received from a first domain, the first confidence score indicating a first likelihood that the second utterance is a request to play the first content on the voice-activated electronic device;

[0233] Receive a second confidence score from the second domain, the second confidence score indicating a second likelihood that the second utterance is a request to play the second content on the output device;

[0234] Determining that the first confidence score is greater than a predetermined threshold indicates that the first functionality of the first domain can serve the second discourse;

[0235] Determining that the second confidence score is also greater than the predetermined threshold indicates that the second functionality of the second domain can also serve the second discourse;

[0236] Determine whether the first or second functionality needs to be selected to respond to the second discourse in order to determine the second intent;

[0237] Generate query text data that expresses the intent to ask whether the second utterance should be responded to by the first domain or the second domain;

[0238] By performing the text-to-speech functionality on the query text data, query audio data representing the query text data is generated;

[0239] Generate an instruction to cause the voice-activated user device to continue sending additional audio data representing local audio after playing the query audio, the additional audio data being captured by the voice-activated electronic device;

[0240] The query audio data is sent to the voice-activated electronic device, causing the intent question to be played by the speaker;

[0241] Send the instruction to the voice-activated user device;

[0242] Receive the additional audio data from the voice-activated user device;

[0243] By performing the speech-to-text functionality on the additional audio data, a third text data representing the additional audio data is generated; and

[0244] The third text data is used to determine that the local audio includes a third utterance of intent response, the intent response having a third intent to play the second content on the output device;

[0245] It is determined that the user account has access to the second content;

[0246] Generate a URL that allows the output device to stream the second content; and

[0247] The URL is sent to the output device, causing the output device to stream the second content.

[0248] 3. The method as described in Clause 1, further comprising:

[0249] Receive second request audio data representing a second utterance from the voice-activated electronic device;

[0250] By performing speech-to-text functionality on the second requested audio data, second text data representing the second requested audio data is generated;

[0251] The second intent of the second text data is determined to be a request to play the first song on the voice-activated electronic device;

[0252] Generate a first URL that allows the voice-activated electronic device to stream the first song;

[0253] The first URL is sent to the voice-activated electronic device, causing the song to be played using the speaker;

[0254] Receive third request audio data representing a third utterance from the voice-activated user equipment;

[0255] By performing speech-to-text functionality on the third requested audio data, third text data representing the third requested audio data is generated;

[0256] The third intent of the third text data is determined to be another request to play the first song on the output device;

[0257] Generate an instruction to stop the voice-activated user device from playing the first song;

[0258] The instruction is sent to the voice activation device to stop playing the first song on the voice-activated user device;

[0259] Generate a second URL that allows the output device to stream the first song;

[0260] Generate song video data for the output device;

[0261] The second URL is sent to the output device, causing the first song to be played by the second speaker;

[0262] The song video data is sent to the output device, so that the song video data is played on the display screen while the first song is played by the second speaker;

[0263] Receive fourth text data indicating a television confirmation message;

[0264] By performing text-to-speech functionality on the fifth text data, fifth audio data representing the fourth text data is generated; and

[0265] The fifth audio data is sent to the voice-activated user device, causing the television confirmation to be played by the speaker.

[0266] 4. A method for routing content, the method comprising:

[0267] Receive first audio data representing a first utterance from a first electronic device;

[0268] Determine that the first user account is associated with the first electronic device;

[0269] Generate first text data representing the first audio data;

[0270] Use the first text data to determine the first intent of the first utterance;

[0271] It was determined that the second electronic device was also associated with the user account;

[0272] Determine that the image response to the first utterance can be sent to the second electronic device;

[0273] Generate second text data representing a first response to the first utterance;

[0274] Generate third text data representing a second response to the first utterance;

[0275] Generate second audio data representing the second text data;

[0276] The second audio data is sent to the first electronic device, causing the first electronic device to output the first response;

[0277] Generate third audio data representing the third text data;

[0278] The third audio data is sent to the second electronic device, causing the second electronic device to output the second response;

[0279] Generate image data representing the image response; and

[0280] The image data is sent to the second electronic device, causing the image response to be output on the second electronic device.

[0281] 5. The method as described in Clause 4, wherein determining the first user account further includes:

[0282] Receive a customer identifier associated with the first electronic device; and

[0283] Determine that the customer identifier is associated with the user account.

[0284] 6. The method as described in Clause 4, further comprising:

[0285] Determining the first intent also includes: determining that the first intent is to play content on the target device; and

[0286] Determining that the second electronic device is also associated with the user account also includes determining that the target device is the second electronic device.

[0287] 7. The method as described in Clause 4, further comprising:

[0288] Receive fourth audio data representing the second utterance from the first electronic device;

[0289] Generate fourth text data representing the fourth audio data;

[0290] The second intent of the fourth text data is determined by the following method:

[0291] Receive the first confidence score that exceeds a predetermined threshold;

[0292] Receive a second confidence score that exceeds the predetermined threshold;

[0293] Generate the fifth text data representing the query message;

[0294] Generate fifth audio data representing the fifth text data;

[0295] Generate a first instruction for the first electronic device;

[0296] The fifth audio data is sent to the first electronic device, so that the first electronic device outputs the fifth audio data;

[0297] Send the first instruction to the first electronic device, causing the first electronic device to send the sixth audio data;

[0298] Receive the sixth audio data representing a response to the query message from the first electronic device;

[0299] Generate sixth text data representing the sixth audio data; and

[0300] The sixth text data is used to determine a third intent in response to the query message;

[0301] Receive the first content in response to the second utterance; and

[0302] The first content is sent to the second electronic device, so that the second electronic device outputs the first content.

[0303] 8. The method as described in Clause 7, further comprising:

[0304] Generate the seventh text data;

[0305] Generate seventh audio data representing the seventh text data; and

[0306] The seventh audio data is sent to the first electronic device, so that the first electronic device outputs the seventh audio data.

[0307] 9. The method as described in Clause 4, wherein determining the intent further comprises:

[0308] Based on the first text data, it is determined that at least two domains are capable of responding to the first utterance.

[0309] 10. The method as described in Clause 9, wherein the at least two fields comprise:

[0310] The first field indicates that the second utterance is a request to play a song with a title on the first electronic device; and

[0311] The second field indicates that the second discourse is a request to play a movie with the title on the second electronic device.

[0312] 11. The method as described in Clause 4, further comprising:

[0313] Receive fourth audio data representing the second utterance from the first electronic device;

[0314] Generate fourth text data representing the fourth audio data;

[0315] Determine the second intent of the fourth text data;

[0316] Receive the first content in response to the second utterance;

[0317] Send the first content to the first electronic device, so that the first electronic device outputs the first content;

[0318] Receive fifth audio data representing the third utterance from the first electronic device;

[0319] Generate fifth text data representing the fifth audio data;

[0320] The third intent of the fifth text data is determined to be to play the second content on the second electronic device;

[0321] It has been determined that the second content and the first content are the same;

[0322] Generate a first instruction for the first electronic device;

[0323] Send the first instruction to the first electronic device so that the first electronic device no longer outputs the first content;

[0324] Receive the second content; and

[0325] The second content is sent to the second electronic device, so that the second electronic device outputs the second content.

[0326] 12. The method as described in Clause 11, further comprising:

[0327] Generate second image data; and

[0328] The second image data is sent to the second electronic device.

[0329] 13. An electronic device comprising:

[0330] A communication circuit receives first audio data representing speech from a first electronic device;

[0331] Memory; and

[0332] At least one processor, said at least one processor being operable to:

[0333] The communication circuit is used to receive first audio data representing a command from a first electronic device;

[0334] Determine that the first electronic device is associated with the first user account;

[0335] Generate first text data representing the first audio data;

[0336] Use the first text data to determine the first intent of the utterance;

[0337] It is determined that the second electronic device is associated with the user account;

[0338] Determine that the image response to the first utterance can be sent to the second electronic device;

[0339] Generate second text data representing a first response to the first utterance;

[0340] Generate third text data representing a second response to the first utterance;

[0341] Generate second audio data representing the second text data;

[0342] This causes the communication circuit to send the second audio data to the first electronic device, so that the first electronic device outputs the first response;

[0343] Generate third audio data representing the third text data;

[0344] This causes the communication circuit to send the third audio data to the second electronic device, so that the second electronic device outputs the second response;

[0345] Generate image data representing the image response; and

[0346] This causes the communication circuit to send the image data to the second electronic device, thereby causing the image response to be output on the second electronic device.

[0347] 14. The electronic device as described in Clause 13, wherein the communication circuitry further receives a client identifier from the first electronic device, and the at least one processor is further operable to:

[0348] Determine that the customer identifier is associated with the user account.

[0349] 15. The electronic device as described in Clause 13, wherein the at least one processor is further operable to:

[0350] Determining the first intent also includes: determining that the first intent is to play content on the target device; and

[0351] Determining that the second electronic device is associated with the user account also includes determining that the target device is the second electronic device.

[0352] 16. The electronic device as described in Clause 13, wherein the communication circuitry further receives fourth audio data representing a second utterance from the first electronic device, and the at least one processor is further operable to:

[0353] Generate fourth text data representing the fourth audio data;

[0354] Receive the first confidence score that exceeds a predetermined threshold;

[0355] Receive a second confidence score that exceeds the predetermined threshold;

[0356] Generate the fifth text data representing the query message;

[0357] Generate fifth audio data representing the fifth text data;

[0358] Generate an answer command for the first electronic device;

[0359] This causes the communication circuit to send the fifth audio data to the first electronic device, so that the first electronic device outputs the query message;

[0360] This causes the communication circuit to send the answer instruction to the first electronic device, causing the first electronic device to send a sixth audio signal;

[0361] It is determined that the communication circuit receives the sixth audio data from the first electronic device;

[0362] Generate sixth text data representing the sixth audio data;

[0363] Determine the second intent of the sixth text data;

[0364] Receive the first content in response to the second utterance; and

[0365] This causes the communication circuit to send the first content to the second electronic device, so that the second electronic device outputs the first content.

[0366] 17. The electronic device as described in Clause 16, wherein the at least one processor is further operable to:

[0367] Receive the seventh text data;

[0368] Generate seventh audio data representing the seventh text data; and

[0369] This causes the communication circuit to send the seventh audio data to the first electronic device, so that the first electronic device outputs the seventh audio data.

[0370] 18. The electronic device as described in Clause 16, wherein determining the intent further includes:

[0371] Based on the first text data, it is determined that at least two domains are capable of responding to the first utterance.

[0372] 19. The electronic device as described in Clause 13, wherein the communication circuitry further receives fourth audio data representing a second utterance from the first electronic device, and the at least one processor is further operable to:

[0373] Receive fourth audio data representing the second utterance from the first electronic device;

[0374] Generate fourth text data representing the fourth audio data;

[0375] Determine the second intent of the fourth text data;

[0376] Receive the first content in response to the second utterance;

[0377] This causes the communication circuit to send the first content to the first electronic device, so that the first electronic device outputs the first content.

[0378] Determine that the communication circuit receives fifth audio data representing a third utterance from the first electronic device;

[0379] Generate fifth text data representing the fifth audio data;

[0380] The third intent of the fifth text data is determined to be to output the first content on the second electronic device;

[0381] Generate a first instruction for the first electronic device;

[0382] This causes the communication circuit to send the first instruction to the first electronic device, so that the first electronic device no longer outputs the first content;

[0383] Receive second content, such that the second content is the same as the first content; and

[0384] This causes the communication circuit to send the second content to the second electronic device, so that the second electronic device outputs the second content.

[0385] 20. The electronic device as described in Clause 19, wherein the at least one processor is further operable to:

[0386] Generate second image data; and

[0387] The second image data is sent to the second electronic device.

[0388] 21. A method comprising:

[0389] Receive first audio data representing a first utterance at an electronic device, the first audio data being sent by a voice-activated electronic device;

[0390] Receive a customer identifier associated with the voice-activated electronic device;

[0391] Identify the user account associated with the customer identifier;

[0392] Determine that an association exists stored on the user account, wherein the association is an association between an input device and an output device;

[0393] It is determined that the input device is the voice-activated electronic device;

[0394] It is determined that the output device is an auxiliary device;

[0395] It is determined that the first statement was received by the input device;

[0396] A first content type is determined such that when the input device requests content of the first type, the content of the first type is sent to the output device;

[0397] By performing speech-to-text functionality on the first requested audio data, first text data representing the first audio data is generated;

[0398] The intent of the first utterance is determined using the first text data to be to play the first content;

[0399] Based on the fact that the first content belongs to the first content type, it is determined that the first content should be sent to the auxiliary device;

[0400] Determine that the user account has access to the first content;

[0401] Send a first status request to the auxiliary device to inquire about the functional status of the auxiliary device, wherein the functional status includes at least one of a ready state, an available state, or an unavailable state.

[0402] It is determined that no first status response to the first status request is received within a predetermined amount of time after the first status request is sent;

[0403] Since no first status response is received within the predetermined time, the first status of the auxiliary device is determined to be the unavailable status.

[0404] Generate second text data representing a notification message, the notification message indicating that the auxiliary device cannot play the first content and the voice-activated electronic device will play the first content;

[0405] By performing text-to-speech functionality on the second text data, second audio data representing the second text data is generated;

[0406] The second audio data is sent to the voice-activated electronic device, causing the notification message to be played by the speaker of the voice-activated electronic device;

[0407] Generate a URL that allows the voice-activated electronic device to stream the first content; and

[0408] The URL is sent to the voice-activated electronic device, causing the first content to be played using the speaker after the notification message.

[0409] 22. The method as described in Clause 21, further comprising:

[0410] In response to the first status request, a status update is received from the auxiliary device, the status update indicating that the auxiliary device is in the ready state;

[0411] Generate third text data, which represents a status update indicating that the auxiliary device is in the ready state;

[0412] By performing text-to-speech functionality on the third text data, third audio data representing the third text data is generated;

[0413] The third audio data is sent to the voice-activated electronic device, causing the status update to be played by the speaker;

[0414] Generate a fourth text data, which represents a prompt asking whether the song should be played on the auxiliary device;

[0415] By performing text-to-speech functionality on the fourth text data, fourth audio data representing the fourth text data is generated;

[0416] Generate an answer command for the voice-activated electronic device to send a fifth audio message after the query audio is played;

[0417] The fourth audio data is sent to the voice activation device, causing the prompt to be played by the speaker;

[0418] The answer instruction is sent to the voice-activated electronic device, so that once the prompt is played, the voice-activated electronic device sends the fifth audio data;

[0419] Receive fifth audio data representing a response to the prompt from the voice activation device;

[0420] By performing speech-to-text functionality on the fifth audio data, fifth text data representing the fifth audio data is generated;

[0421] The intent of the fifth text data is determined to be a response to the prompt, the response indicating that the first content should be played on the auxiliary device;

[0422] Generate a first instruction to stop streaming the first content on the voice-activated electronic device;

[0423] The first instruction for stopping the streaming of the first content is sent to the voice-activated electronic device, causing the voice-activated electronic device to stop playing the song;

[0424] Generate a URL that allows the auxiliary device to stream the first content; and

[0425] The URL is sent to the auxiliary device, causing the first content to be played using the second speaker of the auxiliary device.

[0426] 23. The method as described in Clause 21, further comprising:

[0427] Receive third audio data representing the second utterance from the voice-activated electronic device;

[0428] It is determined that the second statement was received by the input device;

[0429] By performing speech-to-text functionality on the third audio data, third text data representing the third audio data is generated;

[0430] The intent of the third text data is determined to be a request to play the second content;

[0431] Based on the fact that the second content belongs to the first content type, it is determined that the second content should be sent to the auxiliary device;

[0432] Send a second state request for the second state to the auxiliary device;

[0433] Receive a second status response from the auxiliary device, the second status response indicating that the second status of the output device is the ready state;

[0434] Generate a URL that allows the auxiliary device to stream the second content;

[0435] The URL is sent to the auxiliary device, causing the second content to be played using the second speaker;

[0436] Generate a fourth text data representing a confirmation message, the confirmation message indicating that the content is being played on the auxiliary device;

[0437] By performing text-to-speech functionality on the fourth text data, fourth audio data representing the fourth text data is generated; and

[0438] The fourth audio data is sent to the voice-activated electronic device, causing the confirmation message to be played by the speaker.

[0439] 24. The method as described in Clause 21, further comprising:

[0440] Receive second request audio data representing a second utterance from the voice-activated electronic device;

[0441] It is determined that the second statement was received by the input device;

[0442] By performing speech-to-text functionality on the third audio data, third text data representing the third audio data is generated;

[0443] The intent of the third text data is determined to be a request to play the second content;

[0444] Based on the fact that the second content belongs to the first content type, it is determined that the second content should be sent to the auxiliary device;

[0445] Send a second state request for the second state to the auxiliary device;

[0446] Receive a second status response from the auxiliary device, the second status response indicating that the second status of the auxiliary device is the available status;

[0447] Generate a ready instruction for the auxiliary device, such that the instruction causes the auxiliary device to change its state from the available state to the ready state;

[0448] The ready command is sent to the output device, causing the second state of the auxiliary device to change from the available state to the ready state;

[0449] Receive a status update from the auxiliary device indicating that the auxiliary device is in the ready state;

[0450] Generate a URL that allows the auxiliary device to stream the second content;

[0451] The URL is sent to the auxiliary device, causing the second content to be played using the second speaker;

[0452] Generate a fourth text data representing a confirmation message, the confirmation message indicating that the content is being played on the auxiliary device;

[0453] By performing text-to-speech functionality on the fourth text data, fourth audio data representing the fourth text data is generated; and

[0454] The fourth audio data is sent to the voice-activated electronic device, causing the confirmation message to be played by the speaker.

[0455] 25. A method for routing content, the method comprising:

[0456] Receive first audio data representing a first command from a first electronic device;

[0457] Determine that there is an association between the first electronic device and the second electronic device;

[0458] Generate first text data representing the first audio data;

[0459] Determine the first intent of the first text data, the first intent indicating that first content should be output;

[0460] Determine whether to send the first content to the second electronic device;

[0461] The first state of the second electronic device is determined to be unavailable;

[0462] Generate second text data;

[0463] Generate second audio data representing the second text data;

[0464] Sending the second audio data to the first electronic device, causing the first electronic device to output the second audio data; and

[0465] The first content is sent to the first electronic device, so that the first electronic device outputs the first content.

[0466] 26. The method as described in Clause 25, wherein determining the first state further includes:

[0467] Send a first status request to the second electronic device;

[0468] It was determined that no first status response to the first status request was received within a predetermined time period; and

[0469] The first state of the second electronic device is determined to be unavailable.

[0470] 27. The method of claim 25, wherein determining the association between the first electronic device and the second electronic device further comprises:

[0471] Receive the customer identifier associated with the first electronic device;

[0472] Identify the user account associated with the customer identifier;

[0473] It is determined that an association exists stored on the user account, and the association is an association between an input device and an output device;

[0474] It is determined that the input device is the first electronic device;

[0475] It is determined that the output device is the second electronic device; and

[0476] A first content type is determined such that when the input device requests content of the first type, the content of the first type is sent to the output device.

[0477] 28. The method of claim 27, wherein determining to send the first content to the second electronic device further comprises:

[0478] It is determined that the first content belongs to the first content type.

[0479] 29. The method as described in Clause 25, further comprising:

[0480] The second state of the second electronic device is determined to be ready;

[0481] Generate third-party text data;

[0482] Generate third audio data representing the third text data;

[0483] Generate a first instruction for the first electronic device;

[0484] The third audio data is sent to the first electronic device, so that the first electronic device outputs the third audio data;

[0485] The first instruction is sent to the first electronic device, causing the first electronic device to send the third audio data;

[0486] Receive the fourth audio data representing a response to the third audio data from the first electronic device;

[0487] Generate fourth text data representing the fourth audio data;

[0488] Determine the second intent of the fourth text data;

[0489] Generate a second instruction for the first electronic device;

[0490] Send the second instruction to the first electronic device, causing the first electronic device to stop outputting the first content; and

[0491] The first content is sent to the second electronic device, so that the second electronic device outputs the first content.

[0492] 30. The method as described in Clause 25, further comprising:

[0493] Receive third audio data representing a second command from the first electronic device;

[0494] Generate third text data representing the third audio data;

[0495] Determine the second intent of the third text data, the second intent indicating that second content should be output;

[0496] Determine whether to send the second content to the second electronic device;

[0497] The second state of the second electronic device is determined to be available;

[0498] Generate a second instruction for the second electronic device;

[0499] The second instruction is sent to the second electronic device, causing the second electronic device to change from the second state to the third state;

[0500] Determine that the third state of the second electronic device is ready; and

[0501] The second content is sent to the second electronic device, so that the second electronic device outputs the second content.

[0502] 31. The method as described in Clause 25, further comprising:

[0503] Receive third audio data representing a second command from the first electronic device;

[0504] Generate third text data representing the third audio data;

[0505] Determine the second intent of the third text data, the second intent indicating that second content should be output;

[0506] Determine whether to send the second content to the second electronic device;

[0507] The second state of the second electronic device is determined to be ready; and

[0508] The second content is sent to the second electronic device, so that the second electronic device outputs the second content.

[0509] 32. The method as described in Clause 30, wherein determining the second state further comprises:

[0510] Send a second status request to the second electronic device; and

[0511] Receive a second state response indicating that the second electronic device is in the ready state.

[0512] 33. An electronic device comprising:

[0513] Communication circuits;

[0514] Memory; and

[0515] At least one processor, said at least one processor being operable to:

[0516] The communication circuit is used to receive first audio data representing a first command from a first electronic device;

[0517] Determine that there is an association between the first electronic device and the second electronic device;

[0518] Generate first text data representing the first audio data;

[0519] Determine the first intent of the first text data, the first intent indicating that first content should be output;

[0520] Determine whether to send the first content to the second electronic device;

[0521] The first state of the second electronic device is determined to be unavailable;

[0522] Generate second text data;

[0523] Generate second audio data representing the second text data;

[0524] This causes the communication circuit to send the second audio data to the first electronic device, causing the first electronic device to output the second audio data; and

[0525] This causes the communication circuit to send the first content to the first electronic device, so that the first electronic device outputs the first content.

[0526] 34. The electronic device as described in Clause 33, wherein the at least one processor is further operable to:

[0527] This causes the communication circuit to send a first status request to the second electronic device; and

[0528] It was determined that no first status response to the first status request was received within a predetermined time period.

[0529] 35. The electronic device as described in Clause 33, wherein the at least one processor is further operable to:

[0530] This causes the communication circuit to receive a client identifier associated with the first electronic device;

[0531] Identify the user account associated with the customer identifier;

[0532] It is determined that an association exists stored on the user account, and the association is an association between an input device and an output device;

[0533] It is determined that the input device is the first electronic device;

[0534] It is determined that the output device is the second electronic device; and

[0535] A first content type is determined such that when the input device requests content of the first type, the content of the first type is sent to the output device;

[0536] 36. The electronic device as described in Clause 35, wherein the at least one processor is further operable to:

[0537] It is determined that the first content belongs to the first content type.

[0538] 37. The electronic device as described in Clause 33, wherein the at least one processor is further operable to:

[0539] The second state of the second electronic device is determined to be ready;

[0540] Generate third-party text data;

[0541] Generate third audio data representing the third text data;

[0542] Generate a first instruction for the first electronic device;

[0543] This causes the communication circuit to send the third audio data to the first electronic device, so that the first electronic device outputs the third audio data;

[0544] This causes the communication circuit to send the first instruction to the first electronic device, causing the first electronic device to send the fourth audio data;

[0545] The communication circuit is used to receive fourth audio data from the first electronic device;

[0546] Generate fourth text data representing the fourth audio data;

[0547] Determine the second intent of the fourth text data;

[0548] Generate a second instruction for the first electronic device;

[0549] This causes the communication circuit to send the second instruction to the first electronic device, causing the first electronic device to stop outputting the first content; and

[0550] This causes the communication circuit to send the first content to the second electronic device, so that the second electronic device outputs the first content.

[0551] 38. The electronic device as described in Clause 33, wherein the at least one processor is further operable to:

[0552] The communication circuit is used to receive third audio data representing a second command from the first electronic device;

[0553] Generate third text data representing the third audio data;

[0554] Determine the second intent of the third text data, the second intent indicating that second content should be output;

[0555] Determine whether to send the second content to the second electronic device;

[0556] The second state of the second electronic device is determined to be available;

[0557] Generate a second instruction for the second electronic device;

[0558] This causes the communication circuit to send the second instruction to the second electronic device, causing the second electronic device to change from the second state to the third state;

[0559] The third state of the second electronic device is determined to be ready;

[0560] Receive the second content; and

[0561] This causes the communication circuit to send the second content to the second electronic device, so that the second electronic device outputs the second content.

[0562] 39. The electronic device as described in Clause 33, wherein the at least one processor is further operable to:

[0563] The communication circuit is used to receive third audio data representing a second command from the first electronic device;

[0564] Generate third text data representing the third audio data;

[0565] Determine the second intent of the third text data, the second intent indicating that second content should be output;

[0566] Determine whether to send the second content to the second electronic device;

[0567] The second state of the second electronic device is determined to be ready; and

[0568] This causes the communication circuit to send the second content to the second electronic device, so that the second electronic device outputs the second content.

[0569] 40. The electronic device as described in Clause 33, wherein the at least one processor is further operable to:

[0570] This causes the communication circuit to send a first status request to the second electronic device;

[0571] It was determined that no first status response to the first status request was received within a predetermined time period; and

[0572] The first state of the second electronic device is determined to be unavailable.

[0573] The above embodiments of the invention are provided for illustrative purposes and are not intended to be limiting. Although the subject matter has been described in language specific to structural features, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features described. In fact, specific features are disclosed as illustrative forms of implementing the claims.

Claims

1. A method for routing content, comprising: receiving first audio data representing a first command from a first electronic device; determining that an association exists between the first electronic device and a second electronic device; generating first text data representing the first audio data; determining a first intent of the first text data, the first intent indicating that first content is to be output; determining that the first content is to be sent to the second electronic device; determining that a first status of the second electronic device is unavailable; generating second text data; generating second audio data representing the second text data; sending the second audio data to the first electronic device such that the second audio data is output by the first electronic device; and sending the first content to the first electronic device such that the first content is output by the first electronic device.

2. The method for routing content of claim 1, further comprising: sending a first status request to the second electronic device; determining that a first status response to the first status request is not received within a predetermined amount of time; and determining that the first status of the second electronic device is unavailable.

3. The method for routing content of claim 1, further comprising: receiving a customer identifier associated with the first electronic device; determining a user account associated with the customer identifier; determining that an association is stored on the user account, the association being between an input device and an output device; determining that the input device is the first electronic device; determining that the output device is the second electronic device; and determining a first type of content such that when the input device requests the first type of content, the first type of content is sent to the output device.

4. The method for routing content of claim 1, 2, or 3, further comprising: determining that a second status of the second electronic device is ready; generating third text data; generating third audio data representing the third text data; generating first instructions for the first electronic device; sending the third audio data to the first electronic device such that the third audio data is output by the first electronic device; sending the first instructions to the first electronic device such that the first electronic device sends the third audio data; receiving fourth audio data from the first electronic device representing a response to the third audio data; generating fourth text data representing the fourth audio data; determining a second intent of the fourth text data; generating second instructions for the first electronic device; sending the second instructions to the first electronic device such that the first electronic device stops outputting the first content; and sending the first content to the second electronic device such that the first content is output by the second electronic device.

5. The method for routing content of claim 1, 2, or 3, further comprising: receiving third audio data representing a second command from the first electronic device; generating third text data representing the third audio data; ​ ​ ​ ​ determining a second intent of the third text data, the second intent indicating that second content is to be output; determining that the second content is to be sent to the second electronic device; determining that a second status of the second electronic device is available; generating second instructions for the second electronic device; sending the second instructions to the second electronic device such that the second electronic device changes from the second status to a third status; determining that the third status of the second electronic device is ready; and sending the second content to the second electronic device such that the second content is output by the second electronic device.

6. The method for routing content of claim 1, 2, or 3, further comprising: receiving third audio data representing a second command from the first electronic device; generating third text data representing the third audio data; determining a second intent of the third text data, the second intent indicating that second content is to be output; determining that the second content is to be sent to the second electronic device; determining that a second status of the second electronic device is ready; and sending the second content to the second electronic device such that the second content is output by the second electronic device.

7. The method for routing content of claim 4, wherein determining a second status further comprises: sending a second status request to the second electronic device; and receiving a second status response indicating that the second electronic device is in a ready state.

8. An electronic device for routing content, comprising: communication circuitry; a memory; and at least one processor operable to: receive first audio data representing a first command from a first electronic device; determine that an association exists between the first electronic device and a second electronic device; generate first text data representing the first audio data; determine a first intent of the first text data, the first intent indicating that first content is to be output; determine that the first content is to be sent to the second electronic device; determine that a first status of the second electronic device is unavailable; generate second text data; generate second audio data representing the second text data; send the second audio data to the first electronic device such that the second audio data is output by the first electronic device; and send the first content to the first electronic device such that the first content is output by the first electronic device.

9. The electronic device for routing content of claim 8, further comprising: sending a first status request to the second electronic device; determining that a first status response to the first status request is not received within a predetermined amount of time; and determining that the first status of the second electronic device is unavailable.

10. The electronic device for routing content of claim 8, further comprising: receiving a customer identifier associated with the first electronic device; determining a user account associated with the customer identifier; determining that an association exists on the user account, the association being an association between an input device and an output device; determining that the input device is the first electronic device; ​ ​ ​ ​ determining that the output device is the second electronic device; and determining a first type of content such that the first type of content is sent to the output device when the input device requests the first type of content.

11. The electronic device for routing content of claim 8, 9, or 10, further comprising: determining that a second state of the second electronic device is ready; generating third textual data; generating third audio data representing the third textual data; generating first instructions for the first electronic device; sending the third audio data to the first electronic device such that the third audio data is output by the first electronic device; sending the first instructions to the first electronic device such that the first electronic device sends the third audio data; receiving fourth audio data from the first electronic device representing a response to the third audio data; generating fourth textual data representing the fourth audio data; determining a second intent of the fourth textual data; generating second instructions for the first electronic device; sending the second instructions to the first electronic device such that the first electronic device stops outputting the first content; and sending the first content to the second electronic device such that the first content is output by the second electronic device.

12. The electronic device for routing content of claim 8, 9, or 10, further comprising: receiving third audio data representing a second command from the first electronic device; generating third textual data representing the third audio data; determining a second intent of the third textual data, the second intent indicating that a second content is to be output; determining that the second content is to be sent to the second electronic device; determining that a second state of the second electronic device is available; generating second instructions for the second electronic device; sending the second instructions to the second electronic device such that the second electronic device changes from the second state to a third state; determining that the third state of the second electronic device is ready; and sending the second content to the second electronic device such that the second content is output by the second electronic device.

13. The electronic device for routing content of claim 8, 9, or 10, further comprising: receiving third audio data representing a second command from the first electronic device; generating third textual data representing the third audio data; determining a second intent of the third textual data, the second intent indicating that a second content is to be output; determining that the second content is to be sent to the second electronic device; determining that a second state of the second electronic device is ready; and sending the second content to the second electronic device such that the second content is output by the second electronic device.

14. The electronic device for routing content of claim 11, wherein determining a second state further comprises: sending a second state request to the second electronic device; and receiving a second state response indicating that the second electronic device is in a ready state.

Citation Information

Patent Citations

  • Audiovisual multi-room support

    CN102123066A

  • Transferring state information between electronic devices

    CN103782588A