Automated Assistant with Audio Presentation Interaction
By analyzing the entities in the audio presentation and receiving user queries, the automation assistant can effectively handle the interaction between the user and the audio presentation, solving the problem of limited interaction functions in the prior art, and achieving a richer user experience.
Patent Information
- Application Number
- CN202080100658.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-15
- Filing Date
- 2020-12-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-12-14
AI Technical Summary
Existing automation assistants interact with users, especially in audio presentation, with very limited interaction functions, making it difficult to effectively handle user queries related to specific points or entities of audio presentation.
By analyzing the verbal audio content associated with the audio presentation, identifying the entities proposed therein, and receiving a user query during the audio presentation playback, determining whether the query is for the audio presentation, generating an appropriate response.
It realizes the interaction between users and audio presentation, and can generate responses for specific points or entities in audio presentation, expands the interactive function of the automation assistant and improves the user experience.
Smart Images

Figure CN115605840B_ABST
Abstract
Description
Background Art
[0001] Humans can engage in a human-computer dialogue with an interactive software application, which is herein referred to as an "automation assistant" (also known as a "chatbot", "interactive personal assistant", "intelligent personal assistant", "personal voice assistant", "conversation agent", etc.). For example, humans (who may be referred to as "users" when they interact with the automation assistant) can use oral natural language input (i.e., utterances) - in some cases, the oral natural language input can be converted to text and then processed - and / or provide commands and / or requests to the automation assistant by providing text (e.g., typed) natural language input. The automation assistant typically responds to commands or requests by providing a response user interface output, which can include auditory and / or visual user interface output.
[0002] Automation assistants enable users to obtain information, access services, and / or perform various tasks. For example, users can perform searches, get directions, and in some cases, interact with third-party computing services. Users may also be able to perform various operations, such as hailing a ride from a ridesharing app, ordering goods or services (e.g., pizza), controlling smart devices (e.g., light switches), making reservations, etc.
[0003] Automation assistants can talk to users using speech recognition and natural language processing, and some automation assistants also utilize machine learning and other artificial intelligence techniques, for example, to predict user intent. Automation assistants can be good at having a conversation with users in a natural, intuitive way, partly because they understand the conversation context. To utilize the conversation context, automation assistants can save recent inputs from the user, questions from the user, and / or responses / questions provided by the automation assistant. For example, a user might ask, "Where is the closest coffeeshop?", and the automation assistant might answer, "Two blocks east." The user might then ask, "How late is it open?" By retaining at least some form of the conversation context, the automation assistant can determine that the pronoun "it" refers to the "coffee shop" (i.e., coreference resolution).
[0004] Many automated assistants are also used to play back audio content such as music, podcasts, radio stations or streams, audiobooks, etc. Automated assistants running on mobile devices or standalone interactive speakers often include speakers or can be connected to headphones through which users can listen to the audio content. However, traditionally, interaction with such audio content has been mainly limited to controlling playback, such as starting playback, pausing, ending playback, skipping forward or backward, muting or changing the playback volume, or querying the automated assistant for overall information about the audio content, such as obtaining the title of a song or information about the artist who recorded the song. In particular, for audio presentations that contain spoken content, the scope of automated assistant interaction is very limited. SUMMARY OF THE INVENTION
[0005] Techniques are described herein for supporting a user's interaction with an audio presentation of an automated assistant, and in particular, techniques for interacting with the spoken content of such an audio presentation at a specific point within the audio presentation. Analysis of the audio presentation can be performed to identify one or more entities presented by, mentioned in, or otherwise associated with the audio presentation, and discourse classification can be performed to determine whether a received discourse during the playback of the audio presentation is directed at the audio presentation and, in some cases, at a specific entity and / or playback point within the audio presentation, enabling an appropriate response to the discourse to be generated.
[0006] Accordingly, in one aspect consistent with the present invention, a method can include analyzing spoken audio content associated with an audio presentation to identify one or more entities presented in the audio presentation, receiving a user query during the playback of the audio presentation, and determining whether the user query is directed at the audio presentation, and if the user query is determined to be directed at the audio presentation, generating a response to the user query, wherein determining whether the user query is directed at the audio presentation or generating a response to the user query uses the identified one or more entities.
[0007] In some embodiments, analyzing spoken audio content associated with an audio presentation includes performing speech recognition processing on the spoken audio content to generate a transcribed text, and performing natural language processing on the transcribed text to identify the one or more entities. Additionally, in some embodiments, during the playback of the audio presentation by an assistant device, speech recognition processing, natural language processing, and receiving a user query are performed on the assistant device.
[0008] Furthermore, in some embodiments, receiving a user query is performed on the assistant device during the playback of the audio presentation by the assistant device, and at least one of performing speech recognition processing and performing natural language processing is performed before the audio presentation is played back. In some embodiments, at least one of performing speech recognition processing and performing natural language processing is performed by a remote service.
[0009] In addition, some embodiments may also include determining one or more suggestions using the identified one or more entities based on a particular point in the audio presentation. Some embodiments may also include presenting the one or more suggestions on an assistant device during playback of the particular point in the audio presentation by the assistant device. In addition, some embodiments may also include preprocessing responses to one or more potential user queries using the identified one or more entities before receiving a user query. Further, in some embodiments, generating a response to a user query includes generating a response to the user query using a preprocessed response from among the one or more preprocessed responses.
[0010] In some embodiments, determining whether a user query is directed to an audio presentation includes providing the transcribed text from the audio presentation and the user query to a neural network-based classifier that is trained to output an indication of whether a given user query is likely directed to a given audio presentation. Some embodiments may also include buffering audio data from the audio presentation before receiving the user query, and analyzing the spoken audio content associated with the audio presentation includes analyzing the spoken audio content from the buffered audio data after receiving the user query to identify one or more entities presented in the buffered audio data, and determining whether the user query is directed to the audio presentation or generating a response to the user query using the identified one or more entities presented in the buffered audio data. Further, in some embodiments, the audio presentation is a podcast.
[0011] In some embodiments, determining whether a user query is directed to an audio presentation includes using the identified one or more entities to determine whether the user query is directed to the audio presentation. Further, in some embodiments, generating a response to the user query includes using the identified one or more entities to generate a response to the user query. In some embodiments, determining whether a user query is directed to an audio presentation includes determining whether the user query is directed to a particular point in the audio presentation. Further, in some embodiments, determining whether a user query is directed to an audio presentation includes determining whether the user query is directed to a particular entity in the audio presentation.
[0012] In addition, in some embodiments, receiving a user query is performed on an assistant device, and determining whether the user query is directed to an audio presentation includes determining whether the user query is directed to the audio presentation rather than a general query for the assistant device. In some embodiments, receiving a user query is performed on an assistant device, and determining whether the user query is directed to an audio presentation includes determining that the user query is directed to the audio presentation rather than a general query for the assistant device.
[0013] In addition, in some embodiments, determining whether a user query is directed to an audio presentation further includes determining that the user query is directed to the assistant device rather than non-query utterances. Some embodiments may also include determining whether to pause the audio presentation in response to receiving the user query.
[0014] In addition, in some embodiments, determining whether to pause audio presentation includes determining whether a query can be responded to with a visual response, and the method further includes, in response to determining that the query can be responded to with a visual response, visually presenting the generated response without pausing the audio presentation, and in response to determining that the query cannot be responded to with a visual response, pausing the audio presentation and presenting the generated response when the audio presentation is paused. In addition, in some embodiments, determining whether to pause audio presentation includes determining whether the audio presentation is being played on a pausable device, and the method further includes, in response to determining that the audio presentation is not being played on a pausable device, presenting the generated response without pausing the audio presentation, and in response to determining that the audio presentation is being played on a pausable device, pausing the audio presentation and presenting the generated response when the audio presentation is paused.
[0015] In accordance with another aspect of the present invention, a method may include, during playback of an audio presentation that includes verbal audio content, receiving a user query and determining whether the user query is directed to the audio presentation, and if the user query is determined to be directed to the audio presentation, generating a response to the user query, wherein determining whether the user query is directed to the audio presentation or generating a response to the user query uses one or more entities identified from an analysis of the audio presentation.
[0016] In accordance with another aspect of the present invention, a method may include, during playback of an audio presentation that includes verbal audio content, buffering audio data from the audio presentation and receiving a user query, after receiving the user query, analyzing the verbal audio content of the buffered audio data to identify one or more entities presented in the buffered audio data, and determining whether the user query is directed to the audio presentation, and if the user query is determined to be directed to the audio presentation, generating a response to the user query, wherein determining whether the user query is directed to the audio presentation or generating a response to the user query uses the identified one or more entities.
[0017] In addition, some embodiments may include a system that includes one or more processors and a memory operatively coupled to the one or more processors, where the memory stores instructions that, in response to execution of the instructions by the one or more processors, cause the one or more processors to perform any of the foregoing methods. Some embodiments may also include an automated assistant device that includes an audio input device (e.g., a microphone, a line input, a network or storage interface that receives digital audio data, etc.) and the one or more processors coupled to the audio input device and executing locally stored instructions to cause the one or more processors to perform any of the foregoing methods. Some embodiments also include at least one non-transitory computer-readable medium containing instructions that, in response to execution of the instructions by the one or more processors, cause the one or more processors to perform any of the foregoing methods.
[0018] It should be understood that all combinations of the foregoing concepts and additional concepts described in greater detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a block diagram of an example computing environment in which embodiments disclosed herein may be implemented.
[0020] Figure 2 is a block diagram of an example embodiment of an example machine learning stack in which embodiments disclosed herein may be implemented.
[0021] Figure 3 is a flowchart showing an example sequence of operations for capturing and analyzing audio content from an audio presentation according to various embodiments.
[0022] Figure 4 is a flowchart showing an example sequence of operations for capturing and analyzing audio content from an audio presentation using a remote service according to various embodiments.
[0023] Figure 5 is a flowchart showing an example sequence of operations for capturing and analyzing audio content from an audio presentation using audio buffering according to various embodiments.
[0024] Figure 6 is a flowchart showing an example sequence of operations for presenting suggestions associated with an audio presentation according to various embodiments.
[0025] Figure 7 is a flowchart showing an example sequence of operations for processing a utterance and generating a response to the utterance according to various embodiments.
[0026] Figure 8 Illustrates an example architecture of a computing device. Detailed implementation
[0027] Now turning to Figure 1 , an example environment 100 in which the techniques disclosed herein can be implemented is shown. The example environment 100 includes an assistant device 102 interfacing with one or more remote and / or cloud-based automated assistant components 104, which can be implemented on one or more computing systems (collectively referred to as "cloud" computing systems) communicatively coupled to the assistant device 102 via one or more local area networks and / or wide area networks generally indicated at 106 (e.g., the Internet). The assistant device 102 and the computing devices operating the remote or cloud-based automated assistant components 104 can include one or more memories for storing data and software applications, one or more processors for accessing data and executing applications, and other components facilitating communication over a network. Operations performed by the assistant device 102 and / or the automated assistant components 104 can be distributed across multiple computer systems, e.g., as computer programs running on one or more computers at one or more locations coupled to each other via a network. In various embodiments, for example, some or all of the functionality of the automated assistant can be distributed among multiple computer systems, or even distributed to client computing devices. In some embodiments, for example, the functionality discussed herein can be performed entirely within a client computing device, e.g., such that such functionality is available to a user even when no online connection exists. Thus, in some embodiments, the assistant device can include a client device, while in other embodiments, the assistant device can include one or more computer systems remote from the client device, or even a combination of a client device and one or more remote computer systems, whereby the assistant device is a distributed combination of devices. Accordingly, in various embodiments, the assistant device can be considered to include any electronic device implementing any functionality of the automated assistant.
[0028] The assistant device 102 in the illustrated embodiment is generally a computing device on which an instance of an automated assistant client 108, through its interaction with one or more remote and / or cloud-based automated assistant components 104, can form what appears to the user to be a logical instance of an automated assistant that the user can utilize to engage in a human-machine conversation. For simplicity and brevity, the term "automated assistant" as used herein to refer to serving a particular user will refer to the combination of the automated assistant client 108 executing on the assistant device 102 operated by the user and one or more remote and / or cloud-based automated assistant components 104 (which, in some embodiments, can be shared among multiple automated assistant clients).
[0029] Assistant device 102 may also include instances of various applications 110 which, in some embodiments, may interact with or otherwise be supported by the automated assistant. Among the various applications 110 that may be supported are, for example, audio applications such as podcast applications, audiobook applications, audio streaming applications, and the like. Additionally, from a hardware perspective, assistant device 102 may include, for example, one or more of the following: a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device of the user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker, a smart appliance such as a smart TV, and / or a wearable device of the user that includes a computing device (e.g., the user's watch having a computing device, the user's glasses having a computing device, a virtual or augmented reality computing device). Additional and / or alternative computing devices may be used in other embodiments, and it will be understood that the assistant device in various embodiments may utilize the assistant functionality as its sole function, while in other embodiments, the assistant functionality may be a feature of a computing device that performs a number of other functions.
[0030] As described in more detail herein, the automated assistant participates in a human-machine dialogue session with one or more users via the user interface input and output device assistant device 102. Additionally, various additional components reside in assistant device 102 that are related to supporting such a session, and specifically, support user interaction with the audio presentation on the assistant device.
[0031] For example, the audio playback module 112 may be used to control the playback of various audio presentations, such as one or more audio presentations residing in the audio presentation storage 114 or one or more audio presentations streamed from a remote service. For example, one or more speakers of assistant device 102 may be used, or alternatively, one or more speakers in, for example, headphones, earbuds, a car stereo system, a home stereo system, a television, etc., that communicate with assistant device 102 may be used to present the audio playback to the user. In addition to or instead of audio playback, the audio recording module 116 may be used to capture at least a portion of an audio presentation that is being played back by another device in the same environment as the assistant device, such as a radio playing near the assistant device.
[0032] In this regard, an audio presentation can be considered any presentation of audio content, and in many cases, a presentation of audio content where at least a portion of the audio content is oral audio content that includes human language speech. While in some embodiments the audio content in an audio presentation can include music and / or singing, in many of the embodiments discussed below, the focus is on audio presentations that include non-singing spoken audio content with which a user may wish to interact, such as podcasts, audiobooks, radio shows, talk shows, news shows, sports shows, educational shows, and the like. In some embodiments, an audio presentation can be directed to fictional and / or non-fictional topics, and in some embodiments, in addition to the audio content, can include visual or graphical content, although in many embodiments, an audio presentation can be limited to the audio content.
[0033] In the embodiments discussed below, the oral audio content associated with an audio presentation can be analyzed to identify one or more entities presented in the audio presentation, and such analysis can be used to perform various operations, such as generating suggestions associated with the oral audio content for display or presentation to a user, and / or responding to a user query posed by the user during playback of the audio presentation. In some embodiments, for example, a user query can be received during playback of the audio presentation, and it can be determined whether the user query is directed to the audio presentation such that if the user query is determined to be directed to the audio presentation, an appropriate response to the user query can be generated and presented to the user. As will become more apparent below, for example, the entities identified through the analysis can be used when attempting to determine whether a user query is directed to the audio presentation and / or when generating a response to the user query.
[0034] To support such functionality, the assistant device 102 may include various additional modules or components 118 to 132. The speech recognition module 118 may, for example, be used to generate or transcribe text (and / or other suitable representations or embeddings) from audio data, while the natural language processing module 120 may be used to generate one or more entities. The module 118 may, for example, receive an audio recording of a voice input (e.g., in the form of digital audio data), and convert the digital audio data into one or more text words or phrases (also referred to herein as tokens). In some embodiments, the speech recognition module 118 is also a streaming module such that the voice input is converted into text on a token-by-token basis in real time or near real time, such that tokens can be effectively output from the module 118 simultaneously with the user's speech, and thus output before the user has issued a complete verbal request. The speech recognition module 118 may rely on one or more acoustic and / or language models that together model the relationship between the audio signal and the speech units in the language as well as the sequence of words in the language. In some embodiments, a single model may be used, while in other embodiments, multiple models may be supported, e.g., to support multiple languages, multiple speakers, etc.
[0035] The speech recognition module 118 converts speech into text, while the natural language processing module 120 attempts to discern the semantics or meaning of the text output by the module. For example, the natural language processing module 120 may rely on one or more grammar models to map action text to specific computer-based actions, and identify entity text and / or other text that constrains the performance of these actions. In some embodiments, a single model may be used, while in other embodiments, multiple models may be supported, e.g., to support different computer-based actions or computer-based action domains (i.e., a collection of related actions such as communication-related actions, search-related actions, audio / video-related actions, calendar-related actions, device control-related actions, etc.). As an example, a grammar model (stored on the assistant device 102 and / or a remote computing device) may map a computer-based action to an action term for a voice-based action query, such as the action terms "tell me more about", "directions to", "navigate to", "watch", "call", "email", "contact", etc.
[0036] In addition, although modules 118 and 120 may be used to process voice inputs or queries from a user in some embodiments, in the illustrated embodiment, modules 118 and 120 are also used to process spoken audio content from an audio presentation. Specifically, the audio presentation analysis module 122 may analyze the audio presentation, at least in part, by utilizing modules 118 and 120 to generate various entities associated with the audio presentation. Alternatively, speech recognition and natural language processing may be performed using separate functions, such as those embedded within module 122, from modules 118, 120. In this regard, an entity may refer to virtually any logical or semantic concept incorporated into the spoken audio content, for example, including but not limited to topics, people, places, things, events, opinions, facts, organizations, dates, times, addresses, URLs, email addresses, measurements, etc. associated with the spoken audio content in the audio presentation. Either of modules 118, 120 may also use additional content metadata (e.g., podcast title, description, etc.) to assist in identifying entities and / or disambiguating entities. Entities may also be logically associated with specific points in the audio presentation, for example, an entity for the topic "Battle of Verdun" includes an associated timestamp indicating that the topic was mentioned at 13:45 in a podcast about World War I. By associating entities with specific points in the audio presentation, the entities can be useful for responding to more ambiguous user queries such as "what year did this happen" or "tell me more about this" because, in many cases, knowing what entity is being presented when the user makes a query at a specific point during playback can assist in resolving the ambiguous aspects of the query.
[0037] In the illustrated embodiment, the audio presentation analysis module 122 may be used to provide feedback to a user in at least two ways, although the present invention is not limited thereto. In some embodiments, it may be desirable for an automated assistant, for example, to provide suggestions to a user during audio presentation playback, such as by displaying interesting facts or suggestions for queries the user may want to make at different points in the audio presentation. Thus, module 122 may include a suggestion generator 124 that is capable of generating suggestions based on the entities identified in the audio presentation (e.g., "Tap here to learn more about <celebrity> interviewed in the podcast") <personality>"interviewed in podcast)" or "Tap here to check-out the offer from <service being advertised>"). In some embodiments, it may also be desirable for the automated assistant to respond to specific queries issued by the user, and as such, module 122 may also include a response generator 126 for generating responses to specific queries. As will become more apparent below, either of generators 124, 126 may be used to generate suggestions and / or responses on demand (i.e., during playback and / or in response to a specific query), and in some embodiments, either of the generators may be used to generate pre-processed suggestions and / or responses prior to playback of the audio presentation to reduce processing overhead during playback and / or query processing.
[0038] To support module 122, the entity and action store 128 may store entities identified in the audio presentation and any actions (e.g., suggestions, responses, etc.) that may be triggered in response to user input associated with any stored entity. While the present invention is not limited thereto, in some embodiments, actions are analogous to verbs and entities are analogous to nouns or pronouns such that a query may identify an action to be performed and one or more entities that are the focus of the action, or otherwise associated with the action to be performed and the one or more entities that are the focus of the action. Thus, when executed, given one or more entities involved (directly or indirectly via the surrounding environment) in the query, the user query may cause the manifestation of a computer-based action (e.g., when a query is issued during a discussion of the Battle of Verdun, "what year did this happen" may be mapped to a web search for the start date of the Battle of Verdun).
[0039] It should be understood that in some embodiments, stores 114, 128 may reside locally in the assistant device 102. However, in other embodiments, stores 114, 128 may reside partially or entirely in one or more remote devices.
[0040] As described above, the suggestion generator 124 of the audio presentation analysis module 122 may generate suggestions for presentation to the user of the assistant device 102. In some embodiments, the assistant device 102 may include a display, and as such, it may be desirable to include a visual rendering module 130 to render a visual representation of the suggestions on the integrated display. Additionally, in cases where a visual response to a query is supported, module 130 may also be adapted to generate a textual and / or graphical response to the query.
[0041] Another module used by the automated assistant 108 is the utterance classifier module 132, which is used to classify any speech utterances detected by the automated assistant 108. Module 132 is generally used to detect voice-based queries from utterances spoken within the environment in which the assistant device 102 is located, and to attempt to determine the intent (if any) associated with the utterance. As will become more apparent below, in the context of the present disclosure, module 132 can be used to determine, for example, whether the utterance is a query, whether the utterance is directed to the automated assistant, whether the utterance is directed to an audio presentation, or even whether the utterance is directed to a specific entity and / or point within the audio presentation.
[0042] It will be understood that in other embodiments, some or all of the functions of any of the foregoing modules and components shown as residing in the assistant device 102 may be implemented in a remote automated assistant component. Thus, the present invention is not limited to Figure 1 the specific function assignments shown.
[0043] Although the functions described herein may be implemented in a number of different ways in different embodiments, Figure 2 an example embodiment is next shown that utilizes an end-to-end audio understanding stack 150 that includes four stages 152, 154, 156, 158 adapted to support interaction with an audio presentation of the automated assistant.
[0044] In this embodiment, the first speech recognition stage 152 generates or transcribes text from the audio presentation 160, which is then processed by the second natural language processing stage 154 to annotate the text with appropriate entities or metadata. The third stage 156 includes two different components, a suggestion generation component 162 that generates suggestions from the annotated text, and a query processing component 164 that detects and determines the intent of a query issued by the user (e.g., a query provided as utterance 166). Each of the components 162, 164 can also utilize context information 168, such as previous user queries or preferences, conversation information, etc., which can further be used to determine the user's intent and / or generate useful and informative suggestions for a particular user. The fourth feedback generation stage 158 can incorporate a ranking system that takes the cumulative options from the components 162, 164 and presents the most likely action to be performed to the user as user feedback 170 based on passive listening input or context of the action. In some embodiments, the stages 152, 154 can be implemented using neural networks similar to those used in the assistant stack to process the speech recognition pipeline and text annotation used to process user utterances, and in many cases can run locally on the assistant device or alternatively at least partially on one or more remote devices. Similarly, in some embodiments, the stages 156, 158 can be implemented as an extension of the assistant stack or alternatively use a custom machine learning stack separate from the assistant stack and be implemented locally on the assistant device or partially or fully on one or more remote devices.
[0045] In some embodiments, the suggestion generation component 162 can be configured to generate suggestions that match the content the user has or is currently listening to and can perform actions such as performing a search or integrating with other applications, e.g., by surfacing deep links related to an entity (or other application functionality). Additionally, the query processing component 164 can determine the intent of a query or utterance issued by the user. Further, as will become more apparent below, the query processing component 164 may also be able to interrupt or pause the playback of the audio presentation in response to a particular query and can include a machine learning model that is capable of classifying whether a query is an auxiliary command related to the audio presentation or generally unrelated and, in some instances, classifying whether the query is related to a particular point in the audio presentation or a particular entity referenced in the audio presentation. In some embodiments, the machine learning model can be a multi-layer neural network classifier that is trained to output an indication of whether a given user query is likely to be directed at a given audio presentation and uses as input embedding layers both, e.g., the transcribed audio content as well as the user query, and returns one or more available actions if the user query is determined to be related to the audio content.
[0046] In some embodiments, the feedback generation phase may combine the outputs from both components 162, 164 and rank what to present to the user and what not to present. For example, it may be the case that a user query is determined to be relevant to an audio presentation, but no actions are returned, yet one or more suggestions may still be desired to be presented to the user.
[0047] Turning now to Figures 3 to 5 , as described above, in some embodiments, in some implementations, the spoken audio content from an audio presentation may be analyzed to identify the various entities referenced by or otherwise associated with the audio presentation. For example, Figure 3 FIG. shows an example operation sequence 200 that may be performed on an audio presentation, e.g., by Figure 1 the audio presentation analysis module 122. In some embodiments, sequence 200 may be performed in real-time fashion - e.g., during audio presentation playback. However, in other embodiments, sequence 200 may be performed prior to audio presentation playback, e.g., as part of a preprocessing operation performed on the audio presentation. For example, in some embodiments, it may be desirable to preprocess multiple audio presentations in batch and store the entities, timestamps, metadata, preprocessing suggestions, and / or preprocessing responses for later retrieval during audio presentation playback. Doing so may reduce the processing overhead associated with supporting user interaction with the audio presentation via an assistant device during playback. In this regard, it may be desirable for such batch processing to be performed by a remote or cloud-based service rather than by an individual user device.
[0048] Thus, as shown in block 202, to analyze an audio presentation, the audio presentation may be presented, retrieved, or captured. In this regard, being presented generally refers to playing back the audio presentation on an assistant device, which may include a local stream or a stream from a remote device or service, and it will be understood that the spoken audio content from the audio presentation is generally obtained as a result of presenting the audio presentation. In this regard, being retrieved generally refers to retrieving the audio presentation from storage. In some instances, the retrieval may be combined with playback, while in other cases, e.g., when preprocessing audio presentations as part of a batch, the retrieval may be separate from any playback of the audio presentation. In this regard, being captured generally refers to obtaining audio content from the audio presentation when the audio presentation is being presented by a device other than the assistant device. In some embodiments, for example, the assistant device may include a microphone that may be used to capture audio being played back by another device in the same environment as the assistant device, e.g., audio being played by a radio, television, home or commercial audio system, or another user's device.
[0049] Regardless of the source of the spoken audio content in an audio presentation, automated speech recognition can be performed on the audio content to generate a transcription of the audio presentation (block 204), and the transcription can then be used to perform natural language processing to identify one or more entities that are presented or otherwise associated with the audio presentation (block 206). Additionally, in some embodiments, one or more pre-processed responses and / or suggestions can be generated and stored based on the identified entities in a manner similar to the way responses and suggestions are generated as described elsewhere herein (block 208). Additionally, as indicated by the arrow from block 208 to block 202, in some embodiments, sequence 200 can be performed incrementally on the audio presentation, whereby playback and analysis effectively occur in parallel during audio presentation playback.
[0050] As described above, the various functions associated with the analysis of an audio presentation can be distributed among multiple computing devices. For example, Figure 4 An operation sequence 210 is shown that relies on an assistant device (left column) communicating with a remote or cloud-based service (right column) to perform real-time streaming and analysis of an audio presentation. Specifically, in block 212, the assistant device can present the audio presentation to the user, e.g., by playing back audio from the audio presentation to the user of the assistant device. During such playback, the audio presentation can also be streamed to the remote service (block 214), and the remote service performs automated speech recognition and natural language processing (blocks 216 and 218) in a manner that is nearly identical to Figure 3 blocks 204 and 206. Additionally, in some embodiments, one or more pre-processed responses and / or suggestions can be generated and stored based on the identified entities in a manner similar to the way Figure 3 block 208. And the entities (and optionally, the responses and / or suggestions) can be returned to the assistant device (block 222). The assistant device then stores the received information (block 224), and the presentation of the audio presentation continues until the presentation ends or is prematurely paused or stopped.
[0051] Now turning to Figure 5 , in some embodiments, it may be desirable to defer or delay the analysis of an audio presentation until a user query requires it, an example of which is shown in sequence 230. Thus, in some embodiments, rather than continuously performing speech recognition and natural language processing on the entire audio presentation during playback, it may be desirable to simply buffer or otherwise store the audio data from the audio presentation until an appropriate user query is received, and then analyze a portion of the audio presentation (e.g., the last N seconds) near the specific point in the audio presentation where the user query is received. Thus, in a representative example, if a user issues a query such as "what year did this happen?", the last approximately 30 seconds of the audio presentation can be analyzed to determine that the current topic being discussed is the Battle of Verdun, and a search can be performed to generate a representative response such as "The Battle of Verdun occurred in 1916."
[0052] Thus, as shown in block 232, the audio presentation can be presented by or captured by the assistant device (if presented by another device), and the last N seconds of the audio presentation are buffered (block 234). In some cases, buffering can include storing the captured audio data, while in other cases, e.g., where the presentation is initiated by the assistant device and the audio presentation is currently stored on the assistant device, buffering can include simply maintaining a reference to the extent of the audio presentation prior to the current playback point, such that the appropriate extent can be retrieved from storage when analysis is required.
[0053] Playback continues in this manner until a user query is received (block 236), and once a user query is received, in a manner similar to that discussed above Figures 3 to 4 automated speech recognition and natural language processing are performed on the buffered audio or extent of the audio presentation (blocks 238 and 240). Then, the user query is processed in a manner similar to that described below Figure 7 but primarily based on the analysis of the buffered portion of the audio presentation.
[0054] Now turning to Figure 6 , in some embodiments, it may be desirable to actively display suggestions on an assistant device during the presentation of an audio presentation, an example of which is shown by the operation sequence 250. Specifically, during the presentation or capture of an audio presentation (block 252), it can be determined whether any stored suggestions are associated with the current playback point in the audio presentation. If not, the playback of the audio presentation continues and control returns to block 252. However, if any suggestions are associated with the current playback point, control goes to block 256 to temporarily display one or more of the suggestions on the assistant device, for example, using cards, notifications, or snippets displayed on the assistant device. In some embodiments, the suggestions can also be associated with actions such that user interaction with the suggestions can trigger specific actions. Thus, for example, if a particular celebrity is being discussed in a podcast, a suitable suggestion could state "Tap here to learn more about <celebrity> interviewed in the podcast" <personality>“interviewed in Podcast)”, whereby the selection of the suggestion can open a browser tab with additional information about the celebrity.
[0055] In different embodiments, suggestions can be generated in a plurality of ways. As described above, for example, suggestions can be generated during the analysis of an audio presentation, either in real time during the playback of the audio presentation or pre-generated, e.g., as a result of a preprocessing of the audio presentation. Thus, the suggestions stored may have been stored only recently while the audio presentation is being analyzed. Additionally, in some embodiments, suggestions can be associated with a particular entity such that whenever the particular entity is referenced at a particular point during the audio presentation and is identified as a result of the analysis of the audio presentation, any suggestions associated with the particular entity can be retrieved, optionally ranked, and presented to the user. Suggestions can also or alternatively be associated with a particular point in the audio presentation such that the retrieval of the suggestions can be based on the current playback point of the audio presentation rather than on what entity is currently being presented in the audio presentation.
[0056] Figure 7 Next, an example operation sequence 260 suitable for processing utterances or user queries is shown, which can be performed, for example, at least in part by Figure 1 the utterance classifier module 132 and, in some embodiments, using a neural network-based classifier trained to output an indication of whether a given user query is likely to be directed at a given audio presentation. The sequence begins with receiving an utterance during the presentation of an audio presentation, e.g., an utterance captured by a microphone of an assistant device (block 262). The utterance can then be processed (block 264) to determine the utterance intent. In some instances, the utterance intent can be based at least in part on one or more identified entities from the audio presentation currently being presented and / or the current playback point in the audio presentation. In some embodiments, the classification of the utterance can be multi-layered and can attempt to determine one or more of the following: (1) whether the utterance is directed at the automated assistant (e.g., as opposed to a non-query utterance directed at another person in the environment, not directed at anyone in particular, or background noise), (2) whether the utterance is directed at the audio presentation (e.g., as opposed to a general query directed at the automated assistant), (3) whether the utterance is directed at a particular point in the audio presentation, or (4) whether the utterance is directed at a particular entity associated with the audio presentation. The response of the assistant device to the utterance can thus vary based on the classification. Additionally, in some embodiments, acoustic echo cancellation can be utilized such that the audio presentation itself is not processed as an utterance.
[0057] In an example implementation, the classification of the utterance is based on whether the utterance is directed to the assistant (block 266); if so, whether the utterance is more specifically directed to the currently presented audio presentation (block 268); and if so, whether the utterance is even more specifically directed to a specific point in the audio presentation (block 270) or to a specific entity in the audio presentation (block 272).
[0058] If the utterance is determined to not be directed to the assistant, the utterance can be ignored, or alternatively, a response such as "I didn't understand that, could you please repeat" can be generated (block 274). If the utterance is determined to be directed to the assistant but not specifically to the audio presentation (e.g., "what is the weather tomorrow"), a response can be generated in a conventional manner (block 276). Similarly, if the utterance is determined to be directed to the audio presentation but not to any specific point or entity in the audio presentation (e.g., "pause podcast" or "what is this podcast called?"), an appropriate response can be generated (block 276).
[0059] However, if the utterance classification determines that the utterance is directed to a specific point and / or entity in the audio presentation, control can transfer to block 278 to optionally determine whether any preprocessed responses are available (e.g., as discussed above in connection with Figures 3 to 4 ). If response preprocessing is not used, block 278 can be omitted. If no preprocessed response is available, a response can be generated (block 280). One or more identified entities and / or the current playback point in the audio presentation can optionally be used to determine an appropriate response when preparing the response.
[0060] Next, if a preprocessed response is available, or a new response has been generated, the response can be presented to the user. In the illustrated embodiment, depending on the type of the response and the context in which the response has been generated, the response can be presented in a number of different ways. Specifically, in the illustrated embodiment, it is determined whether the response is a visual response (block 282), which means that the response can be presented to the user by visual means (e.g., via a display of the assistant device). If so, the response can be presented without pausing the playback of the audio presentation (block 284). An example of such a response can be a notification, card, or snippet displayed on the assistant device in response to a query "What year did this happen" during a discussion of the Battle of Verdun, stating "The Battle of Verdun took place in 1916". However, in other embodiments, no visual response may be supported (e.g., for non-display assistant devices), and thus block 282 can be omitted.
[0061] If the response is not a visual response, it can be determined whether the playback is pausable (block 286). For example, if the playback is on a radio or on a device other than the assistant device and / or is thus uncontrollable, or if the audio presentation is a live stream that is desired not to be paused, the playback may not be pausable. In such a case, control can transfer to block 284 to present the response without pausing the playback. However, if the playback is pausable, control can transfer to block 288 to temporarily pause the playback and present the response, and generally resume the playback of the audio presentation when the response presentation is complete.
[0062] Figure 8 is a block diagram of an example computing device 300 that is adapted to implement all or part of the functionality described herein. Computing device 300 generally includes at least one processor 302 that communicates with a number of peripheral devices via a bus subsystem 304. These peripheral devices can include a storage subsystem 306, including, for example, a memory subsystem 308 and a file storage subsystem 310, user interface input devices 312, user interface output devices 314, and a network interface subsystem 316. The input and output devices allow a user to interact with computing device 300. The network interface subsystem 316 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.
[0063] The user interface input device 312 can include a keyboard, a pointing device such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen incorporated into a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into the computing device 300 or onto a communication network.
[0064] The user interface output device 314 can include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visual image. The display subsystem can also provide a non-visual display such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways of outputting information from the computing device 300 to a user or to another machine or computing device.
[0065] The storage subsystem 306 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 306 can include logic that executes Figure 5 and / or Figure 7 selected aspects of the various sequences shown.
[0066] These software modules are typically executed by the processor 302 alone or in combination with other processors. The memory subsystem 308 used in the storage subsystem 306 can include multiple memories, including a main random access memory (RAM) 318 for storing instructions and data during program execution and a read-only memory (ROM) 420 for storing fixed instructions. The file storage subsystem 310 can provide persistent storage for program and data files and can include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of certain embodiments can be stored in the storage subsystem 306 by the file storage subsystem 310 or in other machines accessible by the processor 302.
[0067] The bus subsystem 304 provides a mechanism for enabling the various components and subsystems of the computing device 300 to communicate with each other as expected. Although the bus subsystem 304 is schematically shown as a single bus, alternative embodiments of the bus subsystem can use multiple buses.
[0068] The computing device 300 can be of various types, including mobile devices, smartphones, tablets, notebook computers, desktop computers, wearable computers, programmable electronic devices, set-top boxes, dedicated assistant devices, workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, Figure 8 the description of the computing device 300 depicted in Figure 8 is only intended as a specific example for the purpose of illustrating some embodiments. Many other configurations of the computing device 300 may have more or fewer components than the computing device 300 depicted in
[0069] In cases where the systems described herein collect personal information about a user or can make use of personal information, the user may be provided with an opportunity to control whether the program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographic location), or to control whether and / or how content that may be more relevant to the user is received from a content server. Additionally, before storing or using certain data, it may be processed in one or more ways so as to remove personal identity information. For example, a user's identity may be processed so that no personal identity information can be determined for the user, or in cases where geographic location information is obtained, the user's geographic location may be generalized (such as to a city, zip code, or state level) so that the user's specific geographic location cannot be determined. Thus, the user can have control over how information about the user is collected and / or used.
[0070] Although several embodiments have been described and illustrated herein, various other means and / or structures can be utilized for performing the functions and / or obtaining the results and / or one or more of the advantages described herein, and each such variation and / or modification is considered to be within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications of the teachings. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation many equivalents to the specific embodiments described herein. Accordingly, it is to be understood that the foregoing embodiments are presented by way of example only, and that within the scope of the appended claims and their equivalents, embodiments may be practiced in a manner different from that specifically described and claimed. Embodiments of the present disclosure are directed to each and every separate feature, system, article, material, kit, and / or method described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the scope of the present disclosure.< / personality> < / personality>
Claims
1. A computer-implemented method, comprising: Analyzing oral audio content associated with an audio presentation to identify one or more entities presented in the audio presentation; Receiving a first user query and a second user query in an assistant device capable of generating an audio response and a visual response during playback of the audio presentation; In response to receiving the first user query: Determining that the first user query is directed to the assistant device; In response to determining that the first user query is directed to the assistant device, determining that the first user query is directed to the audio presentation by determining that the first user query references at least one of the identified one or more entities; And In response to determining that the first user query is directed to the audio presentation, generating a first response to the first user query, wherein generating the first response to the first user query uses at least one of the identified one or more entities; In response to receiving the second user query: Determining that the second user query is directed to the assistant device; In response to determining that the second user query is directed to the assistant device, determining that the second user query is not directed to the audio presentation by determining that the second user query does not reference at least one of the identified one or more entities; And In response to determining that the second user query is not directed to the audio presentation, generating a second response to the second user query that is independent of the audio presentation; Determining whether the first user query can be responded to with a visual response; In response to determining that the first user query can be responded to with a visual response, visually presenting the first response and not pausing the audio presentation; In response to determining that the first user query cannot be responded to with a visual response, determining whether the audio presentation is being played on a pausable device, wherein determining whether the audio presentation is being played on a pausable device includes determining whether the audio presentation is being played on a device other than the assistant device that cannot be controlled by the assistant device; In response to determining that the audio presentation is being played on a pausable device, pausing the audio presentation and presenting the first response when the audio presentation is paused; And In response to determining that the audio presentation is not being played on a pausable device, presenting the first response without pausing the audio presentation.
2. The method according to claim 1, wherein Analyzing the oral audio content associated with the audio presentation includes: Performing speech recognition processing on the oral audio content to generate a transcribed text; and Performing natural language processing on the transcribed text to identify the one or more entities.
3. The method according to claim 2, wherein Receiving the first user query and the second user query is performed on the assistant device during playback of the audio presentation by the assistant device, and wherein at least one of performing the speech recognition processing and performing the natural language processing is performed before playback of the audio presentation.
4. The method according to claim 1, further comprising determining one or more suggestions using the identified one or more entities based on a specific point in the audio presentation.
5. The method according to claim 4, further comprising presenting the one or more suggestions on the assistant device during playback of the specific point in the audio presentation on the assistant device.
6. The method according to claim 1, further comprising, before receiving the first user query and the second user query, using the one or more identified entities to preprocess responses to one or more potential user queries.
7. The method according to claim 6, wherein, Generating the response to the first user query includes using a preprocessed response from among the one or more preprocessed responses to generate the response to the first user query.
8. The method according to claim 1, wherein Determining that the first user query is directed to the audio presentation includes providing the transcribed text from the audio presentation and the first user query to a neural network-based classifier that is trained to output an indication of whether a given user query is likely to be directed to a given audio presentation.
9. The method according to claim 1, further comprising buffering audio data from the audio presentation before receiving the first user query and the second user query, wherein analyzing the spoken audio content associated with the audio presentation includes analyzing the spoken audio content from the buffered audio data after receiving the first user query and the second user query to identify one or more entities presented in the buffered audio data, and wherein determining that the first user query is directed to the audio presentation or generating the first response to the first user query uses the one or more identified entities presented in the buffered audio data.
10. The method according to claim 1, wherein, Determining that the first user query is directed to the audio presentation includes determining that the first user query is directed to a specific point in the audio presentation.
11. The method according to claim 1, wherein, Determining that the first user query is directed to the audio presentation includes determining that the first user query is directed to a specific entity in the audio presentation.
12. The method according to claim 1, wherein, Determining that the second user query is not directed to the audio presentation includes determining that the second user query is a general query directed to the assistant device.
13. A system for supporting user interaction with an audio presentation, the system including one or more processors and a memory operably coupled to the one or more processors, wherein the memory stores instructions that, when executed by the one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 12.
14. An assistant device, comprising: an audio input device; and one or more processors coupled to the audio input device and executing locally stored instructions to cause the one or more processors to perform the method according to any one of claims 1 to 12.
15. A non-transitory computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Playback apparatus and method of controlling the same
US20100232759A1
Detection of creative works on broadcast media
US20130080159A1
Methods, systems, and media for processing queries relating to presented media content
US20160306797A1
Anchored speech detection and speech recognition
US20170270919A1
Enhancing digital media with supplemental contextually relevant content
US20180069914A1