Automated assistant with audio presentation interaction

KR103005373B1Active Publication Date: 2026-08-14GOOGLE LLC
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
KR1020237001956
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-15
Filing Date
2020-12-14
Publication Date
2026-08-14
Estimated Expiration
2040-12-14

Smart Images

  • Figure 112023006218365-PCT00001_ABST
    Figure 112023006218365-PCT00001_ABST
Patent Text Reader

Abstract

User interaction may be supported by audio presentations by an automated assistant, particularly by the voice content of such audio presentations presented at specific points within the audio presentation. Analysis of the audio presentation may be performed to identify one or more entities addressed, mentioned, or otherwise associated with the audio presentation, and utterance classification is performed to determine whether a utterance received during the playback of the audio presentation is for the audio presentation, and in some cases, for a specific entity and / or point of playback within the audio presentation, thereby enabling the generation of an appropriate response to the utterance.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] A person may engage in a person-computer conversation with an interactive software application referred herein as an "automated assistant" (also referred to as a "chatbot," "interactive personal assistant," "intelligent personal assistant," "personal voice assistant," and "conversational agents"). For example, a person (who may be referred to as a "user" when interacting with an automated assistant) may provide commands and / or requests to the automated assistant by providing natural language input (i.e., utterances) that, in some cases, may be converted into text and then processed, and / or text (e.g., typed) natural language input. The automated assistant responds to the commands or requests by providing response user interface outputs, which may generally include auditory and / or visual user interface outputs.

[0002] Through automated assistants, users can obtain information, access services, and perform various tasks. For example, users can launch searches, find directions, and, in some cases, interact with third-party computing services. Users can also perform various actions in ride-sharing applications, such as hailing a car, ordering goods or services (e.g., pizza), controlling smart devices (e.g., light switches), and making reservations.

[0003] Automated assistants can converse with users using speech recognition and natural language processing, and some also use machine learning and other artificial intelligence technologies, for example, to predict user intent. Because automated assistants understand conversational context in part, they can be adept at conducting conversations with users in a natural and intuitive manner. To leverage conversational context, automated assistants can preserve recent input from the user, questions from the user, and / or responses / questions provided by the automated assistant. For example, if a user asks, "Where is the closest coffee shop?", the automated assistant might answer, "Two blocks east." The user could then ask, "How late is it open?" By preserving at least some form of conversational context, the automated assistant can determine whether the pronoun "it" refers to the "coffee shop" (i.e., a common-reference resolution).

[0004] Many automated assistants are used to play audio content such as music, podcasts, radio stations or streams, and audiobooks. Automated assistants running on mobile devices or standalone interactive speakers often include the speaker or can be connected to headphones that allow the user to listen to the audio content. However, generally, interaction with this audio content has been limited primarily to playback controls, such as starting, pausing, stopping, skipping forward or backward, muting, or changing the playback volume, or querying the automated assistant for information about the entire audio content, such as obtaining information about the song title or the artist who recorded the song. In particular, the scope of automated assistant interaction is significantly limited when presenting audio that includes voice content.

[0005] Techniques for supporting user interaction with audio presentations by an automated assistant are described herein, and may be supported in particular by voice content of such audio presentations presented at specific points within the audio presentations. Analysis of the audio presentations may be performed to identify one or more entities addressed, mentioned, or otherwise associated with the audio presentations, and speech classification is performed to determine whether a speech received during the playback of the audio presentations is for the audio presentations, and in some cases, for specific entities and / or points of playback within the audio presentations, so as to generate an appropriate response to the speeches.

[0006] Accordingly, according to one aspect of the present invention, a method comprises the steps of: analyzing voice audio content associated with an audio presentation to identify one or more entities addressed in the audio presentation; receiving a user query during playback of the audio presentation; determining whether the user query is for the audio presentation; and, if the user query is determined to be for the audio presentation, generating a response to the user query, wherein determining whether the user query is for the audio presentation or generating a response to the user query uses the identified one or more entities.

[0007] In some embodiments, the step of analyzing the voice audio content associated with the audio presentation includes the step of performing speech recognition processing on the voice audio content to generate transcribed text, and the step of performing natural language processing on the transcribed text to identify one or more entities. Additionally, in some embodiments, the step of performing the speech recognition processing, the step of performing the natural language processing, and the step of receiving the user query are performed on the assistant device during the playback of the audio presentation by the assistant device.

[0008] Additionally, in some embodiments, the step of receiving the user query is performed on the assistant device during the playback of the audio presentation by the assistant device, and at least one of the step of executing the speech recognition processing and the step of executing the natural language processing is performed before the playback of the audio presentation. In some embodiments, at least one of the step of executing the speech recognition processing and the step of executing the natural language processing is performed by a remote service.

[0009] Additionally, some embodiments further include the step of determining one or more suggestions using one or more identified entities based on a specific point of the audio presentation. Additionally, some embodiments further include the step of presenting one or more suggestions on the assistant device during playback of a specific point of the audio presentation by the assistant device. Additionally, some embodiments further include the step of pre-processing a response to one or more potential user queries using one or more identified entities before receiving a user query. Additionally, in some embodiments, the step of generating a response to the user query includes the step of using one pre-processed response from one or more pre-processed responses to generate a response to the user query.

[0010] In some embodiments, the step of determining whether the user query is for an audio presentation includes providing the text transcribed from the audio presentation and the user query to a neural network-based classifier trained to output an indication of whether the given user query is likely for the given audio presentation. Additionally, some embodiments further include the step of buffering audio data from the audio presentation before receiving the user query, and the step of analyzing voice audio content associated with the audio presentation includes the step of analyzing voice audio content from the buffered audio data after receiving the user query to identify one or more entities addressed in the buffered audio data, and the step of determining whether the user query is for an audio presentation or generating a response to the user query uses one or more identified entities addressed in the buffered audio data. Furthermore, in some embodiments, the audio presentation is a podcast.

[0011] In some embodiments, the step of determining whether the user query is for the audio presentation includes the step of determining whether the user query is for the audio presentation using the identified one or more entities. Additionally, in some embodiments, the step of generating a response to the user query includes the step of generating a response to the user query using the identified one or more entities. In some embodiments, the step of determining whether the user query is for the audio presentation includes the step of determining whether the user query is for a specific point of the audio presentation. Additionally, in some embodiments, the step of determining whether the user query is for the audio presentation includes the step of determining whether the user query is for a specific entity of the audio presentation.

[0012] Additionally, in some embodiments, the step of receiving a user query is performed at an assistant device, and the step of determining whether the user query is for the audio presentation includes the step of determining whether the user query is for the audio presentation rather than a general query to the assistant device. In some embodiments, the step of receiving a user query is performed at an assistant device, and the step of determining whether the user query is for the audio presentation includes the step of determining whether the user query is for the audio presentation rather than a general query to the assistant device.

[0013] Additionally, in some embodiments, the step of determining whether the user query is for audio presentation further includes the step of determining that the user query is for the assistant device rather than a non-query utterance. Additionally, some embodiments further include the step of determining whether to pause the audio presentation in response to receiving the user query.

[0014] Additionally, in some embodiments, the step of determining whether to pause the audio presentation includes the step of determining whether the query can be answered with a visual response, and the method includes the step of presenting the generated response visually without pausing the audio presentation in response to the determination that the query can be answered with a visual response, and the step of pausing the audio presentation and presenting the generated response while the audio presentation is paused in response to the determination that the query cannot be answered with a visual response. Additionally, in some embodiments, the step of determining whether to pause the audio presentation includes the step of determining whether the audio presentation is being played on a pauseable device, and the method includes: the step of presenting the generated response without pausing the audio presentation in response to the determination that the audio presentation is not being played on a pauseable device, and the step of pausing the audio presentation and presenting the generated response while the audio presentation is paused in response to the determination that the audio presentation is being played on a pauseable device.

[0015] According to another aspect of the present invention, the method comprises receiving a user query during playback of an audio presentation including voice audio content, determining whether the user query is for the audio presentation, and, if the user query is determined to be for the audio presentation, generating a response to the user query, wherein determining whether the user query is for the audio presentation or generating a response to the user query uses one or more entities identified from an analysis of the audio presentation.

[0016] According to another aspect of the present invention, a method comprises: buffering audio data from an audio presentation and receiving a user query during playback of an audio presentation including voice audio content; analyzing the voice audio content from the buffered audio data to identify one or more entities addressed in the buffered audio data after receiving the user query; determining whether the user query is for the audio presentation; and, if the user query is determined to be for the audio presentation, generating a response to the user query, wherein determining whether the user query is for the audio presentation or generating a response to the user query uses the identified one or more entities.

[0017] Additionally, some embodiments include a system comprising one or more processors and memory operably connected to said one or more processors, said memory comprising instructions, said instructions causing said one or more processors to perform the method described above in response to the execution of said instructions by said one or more processors. Also, some embodiments may include an audio input device (e.g., a microphone, a wired input, a network or storage interface receiving digital audio data) and an automated assistant device comprising one or more processors coupled to the audio input device and executing locally stored instructions to cause said one or more processors to perform any of the method described above. Some embodiments include at least one non-transient computer-readable medium comprising instructions, said instructions causing said one or more processors to perform any of the method described above in response to the execution of said instructions by said one or more processors.

[0018] All combinations of the above concepts and additional concepts described in great detail in this specification shall be considered as part of the invention disclosed herein. For example, all combinations of the claimed invention appearing at the end of this specification shall be considered as part of the invention disclosed herein. Brief explanation of the drawing

[0019] FIG. 1 is a block diagram of an exemplary computing environment in which the embodiments disclosed in this specification can be implemented. FIG. 2 is a block diagram of an exemplary embodiment of an exemplary machine learning stack in which the embodiments disclosed in this specification can be implemented. FIG. 3 is a flowchart illustrating an exemplary sequence of operations for capturing and analyzing audio content from audio presentation according to various embodiments. FIG. 4 is a flowchart illustrating an exemplary sequence of operations for capturing and analyzing audio content from audio presentation using a remote service according to various embodiments. FIG. 5 is a flowchart illustrating an exemplary sequence of operations for capturing and analyzing audio content from audio presentation using audio buffering according to various embodiments. FIG. 6 is a flowchart illustrating an exemplary sequence of operations for presenting a proposal associated with audio presentation according to various embodiments. FIG. 7 is a flowchart illustrating an exemplary sequence of operations for processing a statement and generating a response thereto according to various implementations. Figure 8 illustrates an exemplary architecture of a computing device. Specific details for implementing the invention

[0020] Now, returning to FIG. 1, an exemplary environment (100) in which the techniques disclosed herein may be implemented is illustrated. The exemplary environment (100) includes an assistant device (102) that interfaces with one or more remote and / or cloud-based automated assistant components (104), which may be optionally implemented in one or more computing systems (collectively referred to as “cloud” computing systems) that are communicably connected to the client device (102) via one or more local and / or wide-area networks (e.g., the Internet) generally shown in (106). The computing device(s) operating the assistant device (102) and the remote or cloud-based automated assistant components (104) may include one or more memories for storing data and software applications, one or more processors for accessing data and executing applications, and other components that support communication over a network. The actions performed by the assistant device (102) and / or the automated assistant component(s) (104) may be distributed across multiple computer systems, for example, as computer programs running on one or more computers at one or more computer locations connected to each other via a network. In various embodiments, for example, some or all of the functions of the automated assistant may be distributed among multiple computer systems or even to client computing devices. In some embodiments, for example, the functions discussed herein may be performed entirely within a client computing device so that the user can use these functions even when there is no online connection.As such, in some embodiments, the assistant device may include a client device, and in other embodiments, the assistant device may include one or more computer systems remote from the client device or even a combination of the client device and one or more remote computer systems, and accordingly, the assistant device is a distributed combination of devices. Accordingly, the assistant device may be considered to include any electronic device that implements any function of the automated assistant in various embodiments.

[0021] In the illustrated embodiment, the assistant device (102) is a computing device capable of forming an instance of an automated assistant client (108) that, by means of interactions with one or more remote and / or cloud-based automated assistant components (104), appears from the user's perspective as a logical instance of an automated assistant that enables the user to engage in human-to-computer conversation. For brevity and simplicity, the term "automated assistant" used herein to refer to "serving" a specific user refers to a combination of an automated assistant client (108) running on the assistant device (102) operated by the user and one or more remote and / or cloud-based automated assistant components (104) (which, in some embodiments, may be shared among multiple automated assistant clients).

[0022] The assistant device (102) may also include instances of various applications (110) that may interact with or be supported by an automated assistant in some embodiments. Various applications (110) that may be supported include, for example, audio applications such as podcast applications, audiobook applications, audio streaming applications, etc. Additionally, from a hardware perspective, the assistant device (102) may be, for example, a desktop computer device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device in the user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker, a smart device such as a smart television, and / or a user's wearable device (e.g., a user's watch equipped with a computing device, a user's glasses equipped with a computing device, a virtual or augmented reality computing device). It will be recognized that additional and / or alternative computing devices may be used in other embodiments, that in various embodiments an assistant device may utilize the assistant function as its sole function, and that in other embodiments the assistant function may be a configuration of a computing device performing various functions.

[0023] As described in more detail in this specification, the automated assistant engages in a human-to-computer conversation session with one or more users through the user interface input and output device(s) of the assistant device (102). Furthermore, in relation to supporting such sessions, various additional components reside in the assistant device (102) to support user interaction with audio presentations on the assistant device, in particular.

[0024] For example, the audio playback module (112) may be used to control the playback of various audio presentations, for example, one or more audio presentations residing in the audio presentation repository (114), or one or more audio presentations streamed from a remote service. The audio playback may be presented to the user, for example, using one or more speakers of the assistant device (102), or alternatively, using one or more speakers communicating with the assistant device (102), such as headphones, earbuds, car stereo, home stereo, television, etc. In addition to or instead of audio playback, the audio recording module (116) may be used to capture at least some of the audio presentations played by another device in the same environment as the assistant device, for example, radio playing near the assistant device.

[0025] In this regard, an audio presentation may be considered as any presentation of audio content, and in many cases, as a presentation of audio content in which at least part of the audio content is voiced audio content containing human language speech. In some embodiments, the audio content of the audio presentation may include music and / or songs, but in many embodiments discussed below, the focus is on audio presentations containing non-songd voiced audio content that the user may wish to interact with, e.g., podcasts, audiobooks, radio programs, talk shows, news programs, sports programs, educational programs, etc. In some embodiments, the audio presentation may be about fictional and / or non-fiction topics, and in many embodiments, the audio presentation may be limited to audio content only, but in some embodiments, it may include visual or graphic content in addition to audio content.

[0026] In the embodiments discussed below, voice audio content associated with an audio presentation may be analyzed to identify one or more entities addressed in the audio presentation, and such analysis may be used to perform various actions, such as generating a proposal associated with the voice audio content to be displayed or presented to a user and / or responding to a user query raised by a user during the playback of the audio presentation. In some embodiments, for example, a user query may be received during the playback of the audio presentation, and a determination may be made as to whether the user query is related to the audio presentation; if it is determined that the user query is related to the audio presentation, an appropriate response to the user query may be generated and presented to the user. As becomes more apparent below, the entities identified by the analysis may be used, for example, when attempting to determine whether the user query is related to the audio presentation and / or when generating a response to the user query.

[0027] To support these functions, the assistant device (102) may include various additional modules or components (118-132). For example, a speech recognition module (118) may be used to generate or transcribe text (and / or other appropriate expressions or embeddings) from audio data, while a natural language processing module (120) may be used to generate one or more entities. The module (118) may receive, for example, an audio recording of a voice input in the form of digital audio data and convert said digital audio data into one or more text words or phrases (also referred to herein as “tokens”). In some embodiments, the speech recognition module (118) is a streaming module that converts the voice input into text in tokens and in real-time or near real-time so that the tokens can be effectively output from the module (118) simultaneously and thus before the complete spoken request is spoken. The speech recognition module (118) may rely on one or more acoustic and / or language models that model the relationship between audio signals and phonetic units in a language, along with word sequences in the language. In some embodiments, a single model may be used, whereas in other embodiments, multiple models may be supported to support, for example, multiple languages, multiple speakers, etc.

[0028] While the speech recognition module (118) converts speech into text, the natural language processing module (120) attempts to identify the semantics or meaning of the text output by the module. For example, the natural language processing module (120) may rely on one or more grammar models to map action text to a specific computer-based action and to identify entity text and / or other text that restricts the performance of such action. In some embodiments, a single model may be used, whereas in other embodiments, multiple models may be supported to support different computer-based actions or computer-based action domains (i.e., a set of related actions such as communication-related actions, search-related actions, audio / visual-related actions, calendar-related actions, device control-related actions, etc.). For example, a grammar model (stored in the assistant device (102) and / or remote computing device(s)) can map computer-based actions to action terms of voice-based action queries such as “tell me more,” “directions,” “navigation,” “view,” “phone,” “email,” “contacts,” etc.

[0029] Furthermore, while modules (118 and 120) may be used in some embodiments for processing voice input or queries from a user, in the illustrated embodiments, modules (118 and 120) are additionally used to process voice audio content from audio presentations. In particular, the audio presentation analysis module (122) may partially analyze the audio presentation by utilizing modules (118, 120) to generate various entities associated with the audio presentation. Alternatively, speech recognition and natural language processing may be performed using functions distinct from modules (118, 120), for example, embedded within module (122). In this regard, an entity may represent substantially any logical or semantic concept integrated into the voice audio content (e.g., a topic, person, place, thing, event, opinion, fact, organization, date, time, address, URL, email address, measure, etc. associated with the voice audio content of the audio presentation). Any of the modules (118, 120) may use additional content metadata (e.g., podcast title, description, etc.) to help identify and / or clarify the entity. An entity may also be logically associated with a specific point in the audio presentation, for example, an entity for the topic "Battle of Verdun" containing an associated timestamp indicating that it was mentioned at 13:45 in a podcast about World War I. By associating an entity with a specific point in audio presentation, knowing which entity is being addressed when a user issues a query at a specific point during playback can help resolve ambiguous aspects of the query in many cases; therefore, the entity can be useful for responding to more ambiguous user queries, such as "the year this happened" or "tell me more about this."

[0030] In the illustrated embodiment, the audio presentation analysis module (122) may be used to provide feedback to the user in at least two ways, but the invention is not limited thereto. In some embodiments, it may be desirable for an automated assistant to provide suggestions to the user during the playback of the audio presentation, for example, by displaying facts of interest or suggestions regarding a query that the user wishes to publish at different points in the audio presentation. As such, the module (122) may include a suggestion generator (124) capable of generating suggestions based on entities identified in the audio presentation (e.g., "Tap here to learn more about <personality>interviewed in podcast" or "Tap here to check-out the offer from<service being advertised> In addition, in some embodiments, it may be desirable for an automated assistant to respond to a specific query issued by a user, and thus the module (122) may also include a response generator (126) for generating a response to a specific query. As will become more apparent below, one of the generators (124, 126) may be used to generate a suggestion and / or response upon request (i.e., during playback and / or in response to a specific query), and in some embodiments, one of the generators may be used to generate a pre-processed suggestion and / or response before audio presentation playback to reduce processing overhead during playback and / or query processing.

[0031] To support the module (122), the entity and action store (128) may store any action (e.g., suggestion, response, etc.) that can be triggered in response to user input associated with any stored entity as well as entities identified in the audio presentation. Although the invention is not so limited, in some embodiments, since actions are similar to verbs and entities are similar to nouns or pronouns, a query may identify or be associated with one or more entities that are the focus of the action and the action to be performed. Thus, when executed, a user query may cause the performance of a computer-based action associated with one or more entities referenced in the query (directly or indirectly through surrounding context) (e.g., if the query is issued during a discussion about the Battle of Verdun, "the year this happened?" may map to a web search for the start date of the Battle of Verdun).

[0032] It will be understood that in some embodiments, the storage (114, 128) may reside locally on the assistant device (102). However, in other embodiments, the storage (114, 128) may reside partially or entirely on one or more remote devices.

[0033] As mentioned above, the suggestion generator (124) of the audio presentation analysis module (122) can generate suggestions to be presented to a user of the assistant device (102). In some embodiments, the assistant device (102) may include a display, and thus it may be desirable to include a visual rendering module (130) to render visual representations of the suggestions on the integrated display. Additionally, if a visual response to a query is supported, the module (130) may be suitable for generating a text and / or graphic response to the query.

[0034] Another module utilized by the automated assistant (108) is a speech classifier module (132) used to classify any voice speech detected by the automated assistant (108). The module (132) is generally used to detect voice-based queries from speech uttered within the environment where the assistant device (102) is located and to attempt to determine the intent (if any) associated with the speech. As will become more apparent below, in the context of the present disclosure, the module (132) may be used to determine, for example, whether the speech is a query, whether the speech is about the automated assistant, whether the speech is about an audio presentation, or even whether the speech is about a specific entity and / or point of the audio presentation.

[0035] It will be understood that some or all of the functions of any of the aforementioned modules and components depicted as residing in the assistant device (102) may be implemented in a remotely automated assistant component in other embodiments. Accordingly, the present invention is not limited to the specific assignment of functions depicted in FIG. 1.

[0036] Although the function described in this specification may be implemented in a number of different ways in different embodiments, FIG. 2 illustrates one exemplary embodiment utilizing an end-to-end audio understanding stack (150) comprising four steps (152, 154, 156, 158) suitable for supporting automated assistant and audio presentation interaction.

[0037] In this embodiment, the first speech recognition step (152) generates or transcribes text from the audio presentation (160), and this is then processed by the second natural language processing step (154) to annotate the text with appropriate entities or metadata. The third step (156) includes two different components: a suggestion generation component (162) that generates suggestions from the annotated text, and a query processing component (164) that detects and determines the intent of a query issued by a user, provided, for example, as a utterance (166). Each component (162, 164) may also utilize context information (168), for example, previous user queries or biases, conversational information, etc., which can be used to determine the user's intent and / or generate useful and beneficial suggestions for a specific user. The fourth feedback generation step (158) may incorporate a ranking system that provides the user with the action most likely to be performed based on user input or context, which is passively listened to as user feedback (170), and accumulated options from components (162, 164) and surfaces. In some embodiments, steps (152, 154) may be implemented using a speech recognition pipeline and a neural network similar to that used in the assistant stack to process text annotations used to process user utterances, and in many cases, may be executed locally on the assistant device or alternatively, at least partially on one or more remote devices. Similarly, steps (156, 158) may be implemented in some embodiments as an extension of the assistant stack or alternatively using a custom machine learning stack separate from the assistant stack, and may be implemented locally on the assistant device or partially or wholly on one or more remote devices.

[0038] In some embodiments, the suggestion generation component (162) may be configured to generate suggestions that match the content the user has or is currently listening to, and may perform actions such as performing a search or integrating with other applications, for example, by surfacing deep links related to entities (or other application functions). Furthermore, the query processing component (164) may determine the intent of a query or utterance issued by the user. Additionally, as will become more apparent below, the query processing component (164) may also stop or pause playback of an audio presentation in response to a specific query and may include a machine learning model capable of classifying whether the query is related to the audio presentation or a general unrelated assistant command, and in some cases, whether the query is related to a specific point in the audio presentation or a specific entity referenced in the audio presentation. In some implementations, the machine learning model may be a multi-layer neural network classifier trained to output an indication of whether a given user query is likely to be related to a given audio presentation, for example, taking both the user query and the transcribed audio content as input and returning one or more available actions if it is determined that the user query is related to the audio content.

[0039] In some implementations, the feedback generation step may combine outputs from all components (162, 164) and rank what to present to the user and what not to present. For example, this may be a case where a user query is considered to be related to audio presentation but no action is returned, and one or more suggestions may still be desirable to display to the user.

[0040] Now, returning to FIGS. 3 through 5, as illustrated, in some embodiments, voice audio content from an audio presentation may be analyzed in some embodiments to identify various entities associated with or referenced by the audio presentation. FIG. 3 illustrates an exemplary sequence of operations (200) that may be performed on an audio presentation, for example, as may be performed by the audio presentation analysis module (122) of FIG. 1. In some embodiments, the sequence (200) may be performed on a real-time basis, for example, during the playback of the audio presentation. However, in other embodiments, the sequence (200) may be performed prior to the playback of the audio presentation, for example, as part of a pre-processing operation performed on the audio presentation. In some embodiments, it may be desirable to pre-process a number of audio presentations in a batch process, for example, and to store entities, timestamps, metadata, pre-processed proposals, and / or pre-processed responses for later retrieval during the playback of the audio presentation. Doing so can reduce the processing overhead associated with supporting user interaction with audio presentation via an assistant device during playback. In this regard, it may be desirable to perform this batch processing by a remote or cloud-based service rather than on a single user device.

[0041] Accordingly, as illustrated in block (202), to analyze the audio presentation, the audio presentation may be presented, retrieved, or captured. In this regard, it will be understood that "presented" generally refers to the playback of the audio presentation on an assistant device, which may include local streaming or streaming from a remote device or service, and that the voice audio content of the audio presentation is generally obtained as a result of the audio presentation. In this regard, "retrieval" generally refers to the retrieval of the audio presentation from a repository. In some cases, retrieval may be combined with playback, but in other cases, retrieval may be separated from the playback of the audio presentation, for example, when the audio presentation is preprocessed as part of a batch process. In this regard, "capture" generally refers to obtaining audio content from the audio presentation when the audio presentation is presented by a device other than the assistant device. In some embodiments, for example, the assistant device may include a microphone that can be used to capture audio that may be played by another device in the same environment as the assistant device, for example, a radio, television, home or business audio system, or another user's device.

[0042] Regardless of the source of the audio content of the audio presentation, automated speech recognition may be performed on the audio content to generate transcribed text of the audio presentation (Block 204), and the transcribed text is used to perform natural language processing to identify one or more entities associated with or addressed to the audio presentation (Block 206). Additionally, in some embodiments, one or more pre-processed responses and / or suggestions may be generated and stored based on the identified entities in a manner similar to how responses and suggestions are generated as described elsewhere in this specification (Block 208). Furthermore, as indicated by the arrow from block (208) to block (202), in some embodiments, the sequence (200) may be performed progressively on the audio presentation, thereby allowing playback and analysis to occur effectively in parallel during the playback of the audio presentation.

[0043] As mentioned above, various functions associated with the analysis of audio presentation can be distributed across multiple computing devices. For example, FIG. 4 illustrates a series of operations (210) that rely on an assistant device (left column) communicating with a remote or cloud-based service (right column) to perform real-time streaming and analysis of audio presentation. Specifically, in block (212), the assistant device may present the audio presentation to the user, for example, by playing audio from the audio presentation to the user of the assistant device. During this playback, the audio presentation may be streamed to a remote service that performs automated speech recognition and natural language processing (blocks (216 and 218)) in a manner similar to blocks (204 and 206) of FIG. 3 (block 214). Additionally, in some embodiments, one or more pre-processed responses and / or suggestions may be generated and stored based on entities identified in a manner similar to block (208) of FIG. 3 (block 220), and the entities (and optionally, responses and / or suggestions) may be returned to an assistant device (block 222). Then, the assistant device stores the received information (block 224), and the presentation of the audio presentation proceeds until the presentation ends or is paused or stopped too early.

[0044] Now, referring to FIG. 5, in some embodiments it may be desirable to defer or delay the analysis of the audio presentation until it is required by a user query, one example of which is illustrated in sequence (230). Thus, rather than continuously performing speech recognition and natural language processing on the entire audio presentation during playback, in some embodiments it may be desirable to simply buffer or store audio data from the audio presentation until an appropriate user query is received, and then analyze a portion of the audio presentation (e.g., the last N seconds) that is close to a specific point in the audio presentation where the user query is received. Thus, in one representative example, when a user issues a query such as "In what year did this happen?", the last 30 seconds or so of the audio presentation can be analyzed to determine if the topic currently being discussed is the Battle of Verdun, and a search can be performed to generate a representative response such as "The Battle of Verdun took place in 1916."

[0045] Accordingly, as illustrated in block (232), the audio presentation may be presented by the assistant device (if presented by another device) or captured along with the last N seconds of the buffered audio presentation (block 234). In some examples, the buffering may consist of storing the captured audio data, and in other examples, for example, when the presentation is initiated by the assistant device and the audio presentation is currently stored in the assistant device, the buffering may simply include maintaining a reference to the range of the audio presentation prior to the current playback point, so that the appropriate range is retrieved from the storage when analysis is required.

[0046] Playback continues in this manner until a user query is received (block 236), and once a user query is received, automated speech recognition and natural language processing are performed on the buffered audio or the range of audio presentation in a manner similar to that discussed above in relation to FIGS. 3 and 4 (blocks (238 and 240)). Then, the user query is processed in a manner similar to that described below in relation to FIG. 7, but primarily based on the analysis of the buffered portion of the audio presentation (block 242).

[0047] Now, referring to FIG. 6, in some embodiments, it may be desirable to display suggestions in advance on the assistant device during the presentation of the audio presentation, an example of which is illustrated by the sequence of actions (250). Specifically, during the presentation or capture of the audio presentation (block 252), a decision may be made as to whether any stored suggestion is associated with the current playback point of the audio presentation. If not, the playback of the audio presentation continues and control returns to block (252). However, if any suggestion is associated with the current point of playback, control returns to block (256) to temporarily display one or more suggestions on the assistant device, for example, using a card, notification, or chip displayed on the assistant device. In some embodiments, suggestions may also be associated with actions so that user interaction with the suggestion can trigger a specific action. Thus, for example, when a specific individual is being discussed in a podcast, an appropriate suggestion would be "Tap here to learn more about <personality>If you enter "interviewed in Podcast." and select the proposal, a browser tab containing additional information about the individual may open.

[0048] Proposals may be generated in various ways in various implementations. As mentioned above, proposals may be generated, for example, in real-time during the playback of the audio presentation, or, for example, in advance during the analysis of the audio presentation as a result of the pre-processing of the audio presentation. As such, the stored proposals may have been saved only recently while the audio presentation is being analyzed. Furthermore, in some implementations, proposals are associated with specific entities, so that whenever a specific entity is referenced at a specific point during the audio presentation and identified as a result of the analysis of the audio presentation, proposals associated with that specific entity are retrieved, optionally ranked, and displayed to the user. Proposals may also be associated with a specific point in the audio presentation, or alternatively, so that the retrieval of proposals can be based on the current playback point of the audio presentation rather than the entity currently being addressed in the audio presentation.

[0049] Next, FIG. 7 illustrates a sequence of exemplary operations (260) suitable for processing a speech or user query, which may be performed, for example, by the speech classifier module (132) of FIG. 1, and in some embodiments, by a neural network-based classifier trained to output an indication of whether a given user query is likely to be about a given audio presentation. The sequence begins by receiving a speech during the presentation of an audio presentation, for example, as captured by a microphone of an assistant device (block 262). Then, the speech may be processed to determine the intent of the speech (block 264). In some cases, the intent of the speech may be based at least partially on one or more identified entities from the audio presentation currently being presented and / or the current playback point of the audio presentation. In some embodiments, the classification of speech may be multilayered and may attempt to determine one or more of (1) whether the speech is directed toward an automated assistant (e.g., non-query speech directed toward another person in the environment, particularly as opposed to no one or background noise), (2) whether the speech is about an audio presentation (e.g., as opposed to a general query toward an automated assistant), (3) whether the speech is about a specific point in the audio presentation, or (4) whether the speech is about a specific entity associated with the audio presentation. Thus, the response of the assistant device to the speech may vary depending on the classification. Additionally, in some embodiments, acoustic echo cancellation is utilized so that the audio presentation itself is not processed as speech.

[0050] In one exemplary embodiment, the classification of a remark is based on whether the remark is about an assistant (block 266); if so, whether the remark is more specifically about an audio presentation currently being presented (block 268); and if so, whether the remark is much more specifically about a specific point in the audio presentation (block 270) or a specific entity in the audio presentation (block 272).

[0051] If it is determined that a remark is not directed at the assistant, the remark may be ignored, or alternatively, a response such as “I didn’t understand that, would you like to repeat it?” may be generated (Block 274). If it is determined that a remark is directed at the assistant but is not specifically about an audio presentation (e.g., “What is the weather tomorrow?”), a response may be generated in the usual manner (Block 276). Similarly, if it is determined that a remark is about an audio presentation but is not about a specific point or entity of the audio presentation (e.g., “Pause podcast” or “What is this podcast called?”), an appropriate response may be generated (Block 276).

[0052] However, if the speech classification determines that the speech relates to a specific point and / or entity of the audio presentation, control is passed to block (278) to optionally determine whether any pre-processed response is available (e.g., as discussed above in relation to FIG. 3-4). If no response pre-processing is used, block (278) may be omitted. If no pre-processed response is available, a response may be generated (block 280). When preparing the response, one or more identified entities and / or current playback points of the audio presentation may be optionally used to determine an appropriate response.

[0053] Next, if a pre-processed response is available or a new response is generated, the response may be presented to the user. In the illustrated embodiment, the response may be presented in various ways depending on the type of response and the context in which the response was generated. In particular, in the illustrated embodiment, a determination is made as to whether the response is a visual response (Block 282), which means that the response may be presented to the user through visual means (e.g., through the display of the assistant device) (Block 282). If so, the response may be provided without pausing the playback of the audio presentation (Block 284). An example of such a response may be a notification, card, or chip in which a message stating "The Battle of Verdun took place in 1916" is displayed on the assistant device in response to the query "When and in what year did this happen?" during a discussion about the entirety of Verdun. However, in other embodiments, no visual response may be supported (e.g. for a non-display assistant device), and thus Block (282) may be omitted.

[0054] If the response is not a visual response, a decision may be made as to whether playback can be paused (block 286). For example, playback cannot be paused if playback is being played on a radio or on a device other than an assistant device and / or is uncontrollable, or if it is a live stream where it is desirable not to pause the audio presentation. In such cases, control may pass to block (284) to present the response without pausing playback. However, if playback can be paused, control may pass to block (288) to temporarily pause playback and present the response, and generally to continue playback of the audio presentation upon completion of the response presentation.

[0055] FIG. 8 is a block diagram of an exemplary computing device (300) suitable for implementing all or part of the functions described herein. The computing device (300) generally includes at least one processor (302) and communicates with a number of peripheral devices through a bus subsystem (304). These peripheral devices may include, for example, a storage subsystem (306) including a memory subsystem (308) and a file storage subsystem (310), a user interface input device (312), a user interface output device (314), and a network interface subsystem (316). The input and output devices enable user interaction with the computing device (300). The network interface subsystem (316) provides an interface to an external network and is connected to corresponding interface devices of other computing devices.

[0056] User interface input devices (312) include pointing devices such as keyboards, mice, trackballs, touchpads or graphic tablets, scanners, touchscreens integrated into displays, audio input devices such as voice recognition systems and microphones, and / or other types of input devices. Generally, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into a computing device (300) or a communication network.

[0057] The user interface output device (314) may include a non-visual display such as a display subsystem, a printer, a fax machine, or an audio output device. The display subsystem may include a flat panel device such as a CRT or LCD, a projection device, or some other mechanism for generating visual images. Additionally, the display subsystem may provide a non-visual display such as an audio output device. Generally, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device (300) to a user or to another machine or computing device.

[0058] The storage subsystem (306) stores programming and data structures for providing the functions of some or all of the modules described herein. For example, the storage subsystem (306) may include logic for performing selected embodiments of various sequences illustrated in FIG. 5, 7 and / or 10.

[0059] These software modules are generally executed by the processor (302) alone or in combination with other processors. The memory (308) used in the storage subsystem (306) may include multiple memories, including a main RAM (318) for storing instructions and data during program execution and a ROM (420) for storing fixed instructions. The file storage subsystem (310) may provide permanent storage for program and data files and may include a hard disk drive, a floppy disk drive, a CD-ROM drive, an optical drive, or a removable media cartridge together with an associated removable media. Modules implementing the functions of specific embodiments may be stored by the file storage subsystem (310) in the storage subsystem (306) or in another machine accessible by the processor(s) (302).

[0060] The bus subsystem (304) provides a mechanism for enabling various components and subsystems of the computing device (300) to communicate with each other as intended. Although the bus subsystem (304) is schematically depicted as a single bus, alternative embodiments of the bus subsystem may use multiple buses.

[0061] The computing device (300) may be of various types, including mobile devices, smartphones, tablets, laptop computers, desktop computers, wearable computers, programmable electronic devices, set-top boxes, dedicated assistant devices, workstations, servers, computing clusters, blade servers, server farms, or any other data processing systems or computing devices. Due to the constantly changing nature of computers and networks, the computing device (300) illustrated in FIG. 8 is intended only as a specific example for the purpose of illustrating some embodiments. Many other configurations of the computing device (300) may have more or fewer components than the computing device (300) illustrated in FIG. 8.

[0062] In cases where the systems discussed herein collect or use personal information regarding users, users may be provided with the opportunity to control whether programs or configurations will collect information regarding user information (e.g., user's social networks, social actions or activities, occupation, user preferences, or user's current geographic location), and to control whether and / or how to receive content from a content server more relevant to the user. Additionally, certain data is treated in one or more different ways before being stored or used so that personally identifiable information is removed. For example, the user's identity is treated so that personally identifiable information regarding the user cannot be determined, or the user's geographic location is generalized from where the location information was obtained (at the city, zip code, or state level) so that the user's specific geographic location cannot be determined. Thus, the user may have control over how information regarding the user is collected and used.

[0063] Although some embodiments have been described and illustrated herein, various other means and / or structures may be utilized to perform the function and / or obtain the result and / or one or more advantages described herein, and such variations and / or modifications are considered to be within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and actual parameters, dimensions, materials, and / or configurations will depend on the specific application or applications in which the teaching is used. A person skilled in the art will be able to recognize or identify many equivalents to the specific embodiments described herein using only routine experimentation. Accordingly, the foregoing embodiments are merely examples, and it should be understood that within the scope of the appended claims and their equivalents, implementations may be carried out differently from those specifically described and claimed. Embodiments of this disclosure relate to each individual configuration, system, article, material, kit, and / or method described herein. Additionally, if the configurations, systems, articles, materials, kits, and / or methods are not mutually inconsistent, any combination of two or more such configurations, systems, articles, materials, kits, and / or methods is included within the scope of the present invention.< / personality> < / personality>

Claims

Claim 1 A method implemented by a computer, comprising: buffering audio data containing speech natural language content from an audio presentation for a certain period of time before receiving a user query; receiving a user query during playback of the audio presentation, wherein the user query is an automated natural language assistant query; performing a semantic analysis on the speech natural language content contained in the audio presentation from the buffered audio data after receiving the user query to identify one or more semantic entities addressed in the audio presentation; determining that the user query contextually refers to the audio presentation based on the one or more identified semantic entities; and generating a response to the user query based on the identified semantic entities in response to the determination that the user query contextually refers to the audio presentation. Claim 2 The method of claim 1, wherein the step of performing semantic analysis on the speech natural language content comprises: the step of performing speech recognition processing on the speech natural language content to generate transcribed text; and the step of performing natural language processing on the transcribed text to identify one or more entities. Claim 3 A method according to claim 2, wherein the steps of executing the speech recognition processing, executing the natural language processing, and receiving the user query are performed on the assistant device during playback of audio presentation by the assistant device. Claim 4 A method according to claim 2, wherein the step of receiving the user query is performed on the assistant device during the playback of the audio presentation by the assistant device, and at least one of the step of executing the speech recognition processing and the step of executing the natural language processing is performed before the playback of the audio presentation. Claim 5 A method according to claim 2, wherein at least one of the step of executing speech recognition processing and the step of executing natural language processing is performed by a remote service. Claim 6 A method according to claim 1, further comprising the step of determining one or more proposals using one or more identified entities based on a specific point of the audio presentation. Claim 7 A method according to claim 6, further comprising the step of presenting one or more suggestions on an assistant device during playback of a specific point of audio presentation by the assistant device. Claim 8 A method according to claim 1, further comprising the step of pre-processing a response to one or more potential user queries before receiving a user query using one or more identified entities. Claim 9 The method of claim 8, wherein the step of generating a response to the user query comprises the step of using one preprocessed response from one or more preprocessed responses to generate a response to the user query. Claim 10 A method according to claim 1, wherein the step of determining whether the user query is for an audio presentation comprises providing text transcribed from the audio presentation and the user query to a neural network-based classifier trained to output an indication of whether the given user query is likely for the given audio presentation. Claim 11 The method of claim 1, wherein the audio presentation is a podcast. Claim 12 A method according to claim 1, wherein the step of determining whether the user query is for the audio presentation includes the step of determining whether the user query is for the audio presentation using one or more identified entities. Claim 13 The method of claim 1, wherein the step of generating a response to the user query comprises the step of generating a response to the user query using one or more identified entities. Claim 14 A method according to claim 1, wherein the step of determining whether the user query is for the audio presentation includes the step of determining whether the user query is for a specific point of the audio presentation. Claim 15 A method according to claim 1, wherein the step of determining whether the user query is for an audio presentation includes the step of determining whether the user query is for a specific entity of the audio presentation. Claim 16 A method according to claim 1, wherein the step of receiving the user query is performed at an assistant device, and the step of determining whether the user query is for the audio presentation includes the step of determining whether the user query is for the audio presentation rather than a general query to the assistant device. Claim 17 In claim 16, the step of determining whether the user query is for audio presentation further comprises the step of determining that the user query is for the assistant device and not a non-query utterance. Claim 18 A method according to claim 1, further comprising the step of determining whether to pause the audio presentation in response to receiving the user query. Claim 19 The method of claim 18, wherein the step of determining whether to pause the audio presentation includes the step of determining whether the query can be answered with a visual response, and the method comprises: the step of visually presenting the generated response without pausing the audio presentation in response to the determination that the query can be answered with a visual response; and the step of pausing the audio presentation and presenting the response generated while the audio presentation is paused in response to the determination that the query cannot be answered with a visual response. Claim 20 The method of claim 18, wherein the step of determining whether to pause the audio presentation includes the step of determining whether the audio presentation is being played on a pauseable device, and the method comprises: the step of presenting the generated response without pausing the audio presentation in response to the determination that the audio presentation is not being played on a pauseable device; and the step of pausing the audio presentation and presenting the generated response while the audio presentation is paused in response to the determination that the audio presentation is being played on a pauseable device. Claim 21 A system comprising one or more processors and a memory operably connected to the one or more processors, wherein the memory comprises instructions, and the instructions enable the one or more processors to perform the method of any one of claims 1 to 20 in response to the execution of the instructions by the one or more processors. Claim 22 An assistant device comprising: an audio input device; and one or more processors connected to the audio input device and executing locally stored instructions, wherein the instructions cause the one or more processors to perform the method of any one of claims 1 to 20. Claim 23 At least one non-transient computer-readable storage medium comprising instructions, wherein the instructions, in response to the execution of instructions by one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 20. Claim 24 delete Claim 25 delete Claim 26 delete Claim 27 delete

Citation Information

Patent Citations

  • Disambiguating input based on context

    KR1020130101505A

  • Method and Apparatus for Voice Recognition

    KR1020180069660A

  • Electronic apparatus, method for determining user utterance intention of thereof, and non-transitory computer readable recording medium

    KR1020180071931A

  • Detection of creative works on broadcast media

    US20130080159A1

  • Method and user device for providing context awareness service using speech recognition

    US20140163976A1