Selective generation and / or selective rendering of continuous content to complete a spoken utterance
A suggestion module addresses latency and resource inefficiencies in information retrieval by generating and rendering ongoing content based on partial spoken utterances, using disfluency detection and machine learning, to provide timely and efficient information without explicit user invocation.
Patent Information
- Application Number
- JP2025171883
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2026-01-27
AI Technical Summary
Existing information retrieval methods, such as search engines and automated assistants, suffer from significant latency and resource utilization when users seek specific information, particularly in spoken interactions, leading to inefficient and computationally wasteful processes.
Implementing a suggestion module that generates and renders ongoing content in response to partial spoken utterances, using techniques like disfluency detection, entropy criteria, and machine learning models to provide continuous content with reduced latency and selective rendering, thereby avoiding explicit user invocation and minimizing resource usage.
The solution reduces latency and conserves computational resources by providing timely and accurate additional information without explicit user input, enhancing user interaction efficiency and minimizing unnecessary processing.
Smart Images

Figure 2026012740000001_ABST
Abstract
Description
[Background technology]
[0001] Search engines, automated assistants, and other applications have facilitated the flow of information from various electronic sources to humans. However, when specific information is sought by a user at a particular time, there may be significant latency in retrieving the specific information from one of these applications and / or there may be significant utilization of client device resources in retrieving the specific information.
[0002] As an example, assume that a user is helping a friend install a smart thermostat and speaks the utterance "HVAC fan wires" and then pauses the utterance when the user realizes that they are unsure about the general color of the wires that control the fan. In such an example, the user could, for example, utilize a search engine or automated assistant to analyze the general color of the wires that control the fan. However, there may be significant latency between the time the user realizes that the color is unsure and the time the color is actually analyzed. Furthermore, analyzing the color may utilize significant client device resources.
[0003] For example, analyzing the color of a common fan wire via a search engine may require accessing a client device (e.g., a smartphone), launching a browser, navigating a search engine page via the browser, determining how to formulate an appropriate query to submit via the search engine page, inputting the query (e.g., speaking or typing), submitting the query, reviewing the results, and / or clicking through the underlying resources to analyze the color. In addition to taking a significant amount of time (resulting in analysis latency), significant smartphone resources are utilized. For example, client device resources are utilized in launching the browser, keeping the smartphone display active during the process, navigating to the search engine page, determining how to formulate an appropriate query, and reviewing the results and / or underlying resources. For example, the client device display may be maintained active throughout those processes, and further, client device processor resources may be required in launching the browser, rendering graphical content throughout those processes, etc.
[0004] Also, for example, analyzing common fan wire colors via an automated assistant may require explicitly invoking the automated assistant (e.g., by speaking a wake phrase), determining how to formulate an appropriate query to submit via the automated assistant, speaking the query, and reviewing the results from the assistant. In particular, due to the syntactic requirements of the automated assistant, the formulated query may need to be longer than a user utterance of "HVAC fan wire is" to obtain accurate (or any) results (e.g., "HVAC fan wire is" may lead to an error by the assistant). For example, it may need to start with "what" and / or include the term "color" (e.g., "What color is HVAC fan wire?"). This may lead to processing longer queries and / or multiple queries when the user attempts to formulate a query using proper syntax. Also, notably, the rendered results may include more than just the accurate answer (green). For example, the rendered results may restate portions of the query (e.g., "HVAC fan wire is generally green") and / or include additional information (e.g., "HVAC fan wire is green, and it connects to terminal G of your thermostat"). Additionally, the rendered results may be rendered in their entirety without an option to stop the rendering, or with only a high latency option to stop the rendering (e.g., tapping a rendered "cancel" software button, speaking a wake phrase and "stop," etc.), which may lead to the rendering of additional content in addition to that required to satisfy the user's information request. Summary of the Invention [Means for solving the problem]
[0005] Implementations described herein relate to generating and rendering ongoing content (e.g., natural language content) that can be used by a user in completing a partial spoken utterance (e.g., a partial spoken utterance of a human user). As an example, assume a user wears earphones and, while conversing with another person, speaks the partial spoken utterance, "The HVAC fan wire is." The earphones' microphones can be used to capture the spoken utterance, and the detected audio data can be processed according to the techniques disclosed herein to determine the ongoing content, which is the natural language content "green." Furthermore, the word "green" can be rendered as synthesized speech output (optionally at a faster speed and / or reduced volume) via the earphones' speakers to provide the user with the natural language content needed to complete the partial spoken utterance. In these and other ways, the need to utilize significant computational resources in analyzing "green" to complete the partial spoken utterance is avoided. For example, there is no need to explicitly invoke an automated assistant, access a search engine interface, formulate and submit an appropriate query thereto, and / or utilize client device resources in rendering results responsive to the appropriate query. Further, an audible rendering of the word "green" may be generated and / or rendered with minimal latency, and still further, "green" may only be rendered selectively (e.g., when disfluency is detected after a partial verbal utterance), thereby mitigating computationally wasteful and unnecessary rendering of "green."
[0006] In various implementations, the generation and / or rendering of the continuous content may be performed automatically (at least when certain conditions are met) in response to partial verbal utterances and without relying on any user input that explicitly invokes the generation and / or rendering. By at least avoiding the need to wait for certain explicit input before generating and / or rendering the continuous content in these or other ways, the continuous content may be rendered with reduced latency. For example, there may be no need to speak a call wake phrase (e.g., "OK Assistant") or physically interact with call software or hardware buttons.
[0007] Furthermore, in many implementations, at least the rendering of the continuous content is selectively performed. For example, automatic rendering of the continuous content may be selectively performed when certain conditions are met. These conditions may include, or may be based on, whether a disfluency is detected following a partial verbal utterance, the duration of the detected disfluency, which modalities or modalities are available for rendering the continuous content, the characteristics of the continuous content, and / or other factors. In these and other ways, automatic rendering may be performed only when certain conditions are met. This may prevent computationally wasteful rendering of content in unfavorable situations while still allowing automatic rendering of the content (without first requiring explicit input) in favorable alternative situations. Thus, implementations seek to balance the efficiency achieved by reducing the latency of rendering continuous content with the inefficiency caused by wastefully rendering unfavorable content.
[0008] As an example of selectively rendering continuous content, a decision to automatically render continuous content may be conditioned on detecting a disfluency, or a disfluency of at least a threshold duration. In some implementations, detecting the disfluency and / or the duration of the disfluency may be based on processing audio data of a partial spoken utterance and / or audio data following (e.g., immediately after) the partial utterance. For example, the audio data may be processed to detect the absence of human speech or other interruptions in human speech (e.g., "umm" or "uh") that indicate disfluency. In some implementations, detecting the disfluency and / or the duration of the disfluency may be based in addition to or instead of on readings from motion sensors of the auxiliary device. For example, readings from an inertial measurement unit (IMU), accelerometer, and / or other motion sensors of the auxiliary device may indicate when a user is speaking and when the user is not speaking.
[0009] As another example, the determination to automatically render the continuous content may be contingent on an entropy criterion, a confidence criterion, and / or other criteria for the continuous content to meet a threshold. For example, the entropy criterion for the continuous content text may be based on an inverse document frequency (IDF) measure for the text (e.g., an IDF measure for the text as a whole and / or for the text as individual terms in the text). As a particular case, the continuous content text for "cat" may have a low IDG measure because it occurs frequently within the resources of a corpus (e.g., a subset of Internet resources), while the continuous content text for "rhododendron" may have a higher IDF measure because it occurs less frequently within the documents of the corpus. In some implementations, the IDF measure for a term may be personal to the user in that it is based at least in part on its occurrence within the user's resources, such as personal documents and / or past speech recognition hypotheses generated (with prior permission from the user) based on the user's past verbal utterances. As another example, the confidence metric may be based on one or more measures that indicate whether the continuous content is actually accurate to complete the partial verbal utterance. For example, the confidence metric may be based on how closely the continuous content aligns with a query that is automatically generated and utilized when reviewing the continuous content. Also, for example, the confidence metric may be based on a measure generated using a language model that indicates how likely it is that the complete utterance (the partial verbal utterance accompanying the continuous content) will be utilized and / or possibly occur in a given language.
[0010] As a further example, determining whether to automatically render continuous content may be based on which one or more modalities are available for rendering the continuous content. For example, when the auxiliary device has a display that is not in use (or that may be temporarily disabled), there may be a higher probability that the continuous content will be automatically rendered via the display compared to when the auxiliary device has only speakers for rendering the continuous content (e.g., lack of a display and / or a display that is currently unavailable). In these and other ways, continuous content may, in some implementations, be more likely to be rendered exclusively visually when a display is available, as this may be less confusing and more natural to a user as opposed to audio rendering of the continuous content.
[0011] As yet another example, determining whether to automatically render continuous content may be based on the factors described above, but not necessarily on any one of the factors individually satisfying a condition. For example, the determination may be based on whether a combined score based on multiple factors satisfies a threshold or other condition. In some implementations, the combined score may be generated by combining (optionally weighted) disfluency criteria (e.g., 1 for disfluency, 0 for no disfluency), disfluency duration criteria (e.g., scaled from 0 to 1 based on duration), entropy criteria, confidence criteria, display criteria (e.g., 1 for available, 0 for unavailable), and / or other criteria (which may optionally be normalized prior to combination).
[0012] In some other implementations, the combined score may be generated using a machine learning model by processing a disfluency criterion, a disfluency duration criterion, an entropy criterion, a confidence criterion, and / or other criteria (optionally normalized). For example, the machine learning model may be a classifier that produces an output between 0 and 1 (e.g., 1 indicates the best score), and the combined score may be the output and may be satisfied if the output exceeds 0.7 or other threshold. The machine learning model may be trained using implicit or explicit user feedback signals and may optionally be personalized to the user. For example, the machine learning model may be personalized to the user (exclusively or strictly) through training based on feedback signals from the user. Additionally or alternatively, the machine learning model may be personalized by adjusting the threshold (against which the output from the model is compared) for the user (e.g., adjusting upward if feedback signals from the user indicate excessive triggering of rendering). Through personalization of the machine learning model, false suppression of rendering and / or erroneous rendering may be further mitigated.
[0013] As one example of generating training examples based on implicit feedback signals, a positive training example may be generated based on the continuous content text being rendered, the user speaking the continuous content text (or closely matching it), and the user speaking the continuous content text after rendering is complete (or at least 90% or other threshold is completed). For example, the positive training example may include features of the input to the machine learning model for that situation and a display output of “1.” As another example, a negative training example may be generated based on the continuous content text being rendered and the user (a) speaking alternative (e.g., not closely matching) continuous content text and / or (b) speaking the continuous content text but doing so early in rendering (e.g., when rendering is less than 40% complete or other threshold indicating no rendering is required). In both of these examples, the actual spoken continuous content text may be determined based on automatic speech recognition of the continuous content text.
[0014] In various implementations, monitoring of a user's voice activity may be performed during the rendering of the continuous content (e.g., based on processing audio data and / or based on sensor readings from a motion sensor of an auxiliary device). In some of these implementations, stopping the rendering of the continuous content may occur in response to detecting user voice activity. For example, assume that the continuous content is a "Utility Model Patent" and the rendering is an audible rendering of the continuous content (e.g., using speech synthesis). If user voice activity is detected simply after "Utility Model" is rendered (and before "Patent" is rendered), the rendering may be stopped, thereby preventing "Patent" from being rendered. This may save computational resources because the term "Patent" does not need to be rendered. Furthermore, this may be more natural to the user and may resonate with the user because the user may have already begun speaking the entire "Utility Model Patent," triggered by the partial rendering. Optionally, the user's verbal utterance leading to the stop may be processed using automatic speech recognition to verify that the rendering should continue to be stopped. For example, continuing with the previous example, if speech recognition indicates that the spoken utterance is "wait" or "just a moment" or similar, the pause may be aborted and the remainder of the continuing content may be rendered.
[0015] In some implementations, certain continuous content may be intentionally rendered in multiple separate chunks (i.e., segments). For example, an address may be provided with a first chunk that is a street address, a second chunk that is a town and state, and a third chunk that is a zip code. In these implementations, the first chunk may be rendered first without then automatically rendering the second chunk. Furthermore, rendering of the second chunk may depend on the initial detection of voice activity during or after the rendering of the first chunk and the detection of a cessation of voice activity. This ensures that the user can complete speaking the first chunk before the second chunk is rendered, which may be more natural for the user. Optionally, rendering of the second chunk may further depend on automatic speech recognition verifying that the content of the first chunk was present in the spoken utterance that caused the detected voice activity. Furthermore, rendering of the third chunk may depend on the detection of voice activity during or after the rendering of the second chunk and the detection of a cessation of voice activity. Optionally, rendering of the third chunk may further depend on automatic speech recognition verifying that the content of the second chunk was present in the spoken utterance that gave rise to the detected voice activity.
[0016] Various techniques may be utilized in generating the continuous content. In some implementations, the generation of the continuous content may be selectively performed. For example, it may be performed in response to the detection of any spoken utterance, or the detection of spoken speech and verification that it is from a user of the auxiliary device. The verification that the spoken utterance is from a user of the auxiliary device may be based on text-independent speaker verification. The verification that the spoken utterance is from a user may additionally or alternatively be based on the automatic speech recognition being speaker-dependent and the automatic speech recognition generating at least one predicted hypothesis for the spoken utterance when the speaker-dependent speaker is the user. In other words, the spoken utterance may be verified to be from a user when speaker-dependent automatic speech recognition that is specific to the user is performed and the automatic speech recognition generates at least one speech recognition hypothesis. The speaker-dependent automatic speech recognition may generate one or more predicted hypotheses only when the audio data is from the user. For example, using speaker-dependent automatic speech recognition, the user's voice embeddings are processed along with the audio data to generate predicted hypotheses for only those portions of the audio data that match the voice embeddings. As another example, the generation of continuous content may be further responsive to detecting that the auxiliary device is actively worn by the user and / or detecting that the auxiliary device is in a certain environmental situation (e.g., a certain location, a certain type of location (e.g., a personal location), a certain time of day, a certain day of the week).
[0017] As an example of generating continuous content, a machine learning model may be used to process speech recognition hypotheses for a partial verbal utterance to generate a subsequent content output indicating one or more features of further content predicted to follow the partial verbal utterance. For example, the machine learning model may be a language model trained to predict likely next words for a partial description (e.g., text) based on processing features of the partial description. The machine learning model may be, for example, a Transformer model or a memory network (e.g., a long-short-term memory (LSTM) model). In some implementations, the machine learning model may be a language model (LM).
[0018] In some implementations, the subsequent content output indicates the next word, and those predicted next words may be used as the continuing content. In some other implementations, the predicted next word may be used in identifying the continuing content, but the continuing content will be different from the predicted next word. For example, the entity type of the predicted next word may be determined (e.g., using an entity detector), and that entity type and the speech recognition hypothesis terms may be used in performing the search. Further, the continuing content may be determined from resources responsive to the search. Performing a search using the predicted entity type and one or more predicted terms may include determining that a database input resource (e.g., from a public or personal knowledge graph) is indexed with both the predicted entity type and one or more hypothesis terms, and determining the continuing content from the database input. Performing a search using the predicted entity type and one or more predicted terms may alternatively include identifying a resource using one or more hypothesis terms and extracting additional text from a superset of the text of the resource based on the additional text being of the predicted entity type.
[0019] As a particular example, assume the partial spoken utterance "What is the capital of Kentucky?" A language model can be used to process the speech recognition hypothesis of "What is the capital of Kentucky?" to generate the predicted continuous content of "Louisville." Notably, "Louisville" is inaccurate (Louisville is not the capital of Kentucky), but such inaccuracy may arise from a machine learning model. Nevertheless, a search may be performed using terms from an entity type (e.g., "city") determined based on the hypothesis and the predicted continuous content of "Louisville." For example, a search of a knowledge graph may be performed to identify nodes of type "city" that are also indexed by "state capital" and "Kentucky" (e.g., have edges targeting the corresponding nodes). The identified nodes include an alias for "Frankfort," which can be used as the continuous content to be rendered. As another example, a search may be performed to identify resources (e.g., web pages) responsive to "What is the capital of Kentucky?" Additionally, a snippet of the text "Frankfort" may be extracted from the resource and utilized as continued content in response to it being of the "city" entity type. In these and other ways, language models may be utilized to enable more efficient analysis of accurate continued content, thereby mitigating the occurrence of computationally wasteful rendering of inaccurate content.
[0020] Additional and / or alternative techniques may be utilized in generating the continuous content. For example, an embedding-based lookup may be performed. For example, embeddings may be generated based on recognition hypotheses, and the embeddings are compared to pre-computed embeddings, each for a different fact. Furthermore, one of the facts may be selected as the continuous content in response to a comparison indicating that the distance between the embedding and the recognition embedding satisfies a threshold. For example, it may be selected based on the distance being the shortest distance between the recognition embedding and all pre-computed embeddings, and based on the distance meeting an absolute threshold.
[0021] In some implementations, one or more aspects described herein may be implemented locally at the auxiliary device or another client device in communication with the auxiliary device (e.g., a smartphone paired with the auxiliary device via Bluetooth). For example, generating continuous content may be implemented locally. In some implementations in which generating continuous content is implemented locally, resources for which a search is performed in generating the continuous content may be proactively downloaded at the auxiliary device or client device before they need to be utilized. In some versions of those implementations, these resources may be proactively downloaded based on current environmental conditions (e.g., resources related to the current location) or current conversational conditions (e.g., resources related to a topic inferred from previous utterances in the conversation). Such proactive downloading of resources and / or local generation of continuous content may further reduce latency in rendering the continuous content.
[0022] The above description is provided as an overview of some implementations of the present disclosure, further descriptions of which and other implementations are described in more detail below.
[0023] Other implementations may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform a method such as one or more of the methods described above and / or elsewhere herein. Still other implementations may include one or more computer systems including one or more processors operable to execute stored instructions to perform a method such as one or more of the methods described above and / or elsewhere herein.
[0024] It should be appreciated that all combinations of the above concepts and additional concepts described in more detail herein are contemplated as being part of the subject matter disclosed herein, for example, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. [Brief explanation of the drawings]
[0025] [Figure 1A] FIG. 10 is a diagram of a suggestion module that allows a user to receive suggested information in a timely manner in certain circumstances without an explicit request from the user. [Figure 1B] FIG. 10 is a diagram of a suggestion module that allows a user to receive suggested information in a timely manner in certain circumstances without an explicit request from the user. [Figure 2A] FIG. 1 is a diagram of a user wearing a pair of earphones and being assisted by a suggestion module that may provide suggestions to the user in time without the user necessarily explicitly requesting the suggestions. [Figure 2B]FIG. 1 is a diagram of a user wearing a pair of earphones and being assisted by a suggestion module that may provide suggestions to the user in time without the user necessarily explicitly requesting the suggestions. [Figure 3] FIG. 1 illustrates a system that enables an auxiliary computing device to render content to a user in a timely manner when it is predicted that the user will desire personal assistance and / or assistance when assisting another person. [Figure 4] FIG. 1 illustrates a method for providing supplemental information in a timely manner via a wearable device in certain situations where a user may or may not have explicitly requested such information from the wearable device. [Figure 5] FIG. 1 is a block diagram of an exemplary computer system. DETAILED DESCRIPTION OF THE INVENTION
[0026] Implementations described herein relate to providing a suggestion module for rendering data in time through an interface of a wearable device and / or other computing device, without the suggestion module necessarily being directly invoked by a user. The suggestion module may operate in one or more different devices, including, but not limited to, a pair of earphones, glasses, a watch, a cellular phone, and / or any other computerized device. In some implementations, the suggestion module may operate in a primary computing device (e.g., a cellular phone) and interface with an auxiliary computing device (e.g., earphones and / or glasses). The suggestion module may, in certain circumstances, optionally operate to provide suggested content to a user without being first invoked by the user. In this way, users may receive some helpful information via their wearable devices through a process that involves less interaction with those devices and / or mitigates interference that may sometimes occur when a user attempts to identify helpful content.
[0027] As an example, a user checking in for a flight may be quizzed by an airline agent about details about the flight. For example, the airline agent may ask the user, "What is your flight number?" When the user confirms the airline agent's query, the user may be wearing earphones and / or glasses in communication with a computerized, or possibly computing, device. A suggestion module may operate on the wearable device and / or another computerized device supporting the wearable device, thereby allowing the content of the airline agent's query to be processed by the suggestion module (without prior permission from the user). For example, audio data corresponding to verbal utterances from the airline agent may be processed to determine that the airline agent has requested the user provide a particular instance of data (e.g., a string of characters corresponding to the flight number). When the suggestion module determines that this instance of data is requested, the suggestion module may identify data that may satisfy the airline agent's query.
[0028] In some implementations, with prior permission from the user, the suggestion module may access one or more different sources of data to determine suitable and / or accurate answers to questions from the airline agent. For example, the suggestion module (with prior permission from the user) may access location data, application data, and / or other data to identify a corpus of data that may be useful in the current situation. For example, the location data may indicate that the user is at an airport, and travel data stored in the memory of the user's cellular device may be associated with that airport. The travel data may be processed by the suggestion module when the user arrives at the airport, thereby allowing the suggestion module to buffer certain output that may be useful to the user while the user is at the airport. In some implementations, the suggestion module may buffer audio data that may respond to inquiries the user may encounter at the airport, even if the suggestion module does not ultimately render the audio data. Alternatively or additionally, the suggestion module and / or other applications may process data that indicates whether the user is in a situation in which audio data should or should not be rendered for the user's benefit.
[0029] For example, a verbal utterance provided by an airport agent may be processed by a computing device that provides access to a suggestion module. When the suggestion module determines that the data available to the suggestion module (e.g., travel data and / or buffered audio data) is suitable for responding to the verbal utterance from the airline agent, the suggestion module and / or other applications may determine whether there is any indication that the suggestion module should or should not cause the audio data to be rendered. For example, a heuristic process and / or a machine learning model may be utilized to determine whether the suggestion module should or should not render the buffered audio data. This determination may be based on a variety of different data, including, but not limited to, the amount of time since the user last spoke, a noise or disfluency made by the user (e.g., "hmm"), a gesture made by the user (e.g., the user looking up, nodding, etc.), inertial measurements, the amount of time since the airline agent last spoke, the rate at which the user speaks, the rate at which the airline agent speaks, an explicit input directly requesting suggestions from the suggestion module, another explicit input instructing the airline agent to suppress suggestions from the suggestion module, and / or some other detectable characteristic of the situation in which the airline agent directs the user to make a verbal utterance.
[0030] In some cases, the user may respond to the airport agent by providing another verbal utterance, such as “Let me check…” followed by a short pause (e.g., at least half a second). This short pause may be characterized by data (e.g., input data from a microphone) and processed according to one or more heuristic processes and / or trained machine learning models to determine whether to render a suggestion to the user. For example, when the short pause exhibited by the user lasts for a threshold amount of time and / or the user makes a particular linguistic gesture (e.g., a short humming sound), the suggestion module may cause audio data to be rendered to the user. For example, after the airline agent provides the verbal utterance “What's your flight number?” and the user provides another verbal utterance “Let me check,” followed by a short pause, the suggestion module may render an audible output such as “K2134.” The audible output may be rendered through a pair of earphones worn by the user during the interaction with the airline agent, thereby allowing the user to receive the “flight number” from the suggestion module without having to initialize the display interface of the cellular phone. This may preserve the user time to respond to important inquiries from third parties and may also help conserve the computing resources of any devices that the user may typically rely on in such situations.
[0031] In some implementations, further training of one or more models utilized by the suggestion module may be performed based on whether the user is determined to be assisted by the suggestion and / or whether the user is determined to desire the suggestion. For example, when the user recites a suggestion (e.g., "K-2-1-3-4...") to an airline agent, one or more models may be updated based on the user's utilization of a suggestion from the suggestion module. Alternatively or additionally, in response to the user's utilization of a suggestion from the suggestion module, a score or other ranking for one or more sources of data that formed the basis of the suggestion may be modified. The determination of whether the user utilized a suggestion may be based on the ARS and determining whether a match or correlation exists between the suggestion and a verbal utterance from the user. In some cases, when the user does not utilize a suggestion from the suggestion module, a score or other ranking for one or more sources of data that formed the basis of the suggestion may be modified accordingly. In this manner, as the suggestion module continues to operate to provide subsequent suggestions, reliance on sources of data that are not determined to be helpful to the user may be reduced.
[0032] In some implementations, one or more models and / or processes utilized to determine whether to render suggestions in certain situations may be updated to reflect how a user has historically utilized suggestions from the suggestion module. For example, a user may respond more affirmatively to suggestions rendered in a public space when close acquaintances are present compared to other suggestions rendered in a personal space (e.g., home or work). As a result, the suggestion module may adapt to render suggestions less frequently in personal spaces and more frequently in public spaces. In some implementations, the properties of how suggestions are rendered may be adapted over time according to how the user responds to suggestions from the suggestion module and / or how the user interacts, possibly directly and / or indirectly, with the device. For example, a user's speaking rate and / or the user's accent may be detected by a device facilitating the functionality of the suggestion module and utilized to render suggestions at a similar rate and / or with a similar accent with prior permission from the user.
[0033] For example, a user who speaks at a faster rate may have suggestions rendered at a similar rate (e.g., the same rate plus or minus a threshold tolerance). In some implementations, properties of the rendered suggestions may be dynamically adapted according to the user's context and / or the content of the suggestion. As an example, the suggestion module may determine that the user speaks at a different rate than a third party to which the suggestions may be directed. Based on this determination, the suggestion module may cause suggestions to be rendered at a rate appropriate for the third party (e.g., an airline agent), which may be interpreted as suggestions on how the user should recite the suggested content back to the third party. For example, even though the user may speak at a higher rate than the third party, the user may recite the suggestions at a lower rate because the suggestion module causes the suggestions to be rendered at the lower rate.
[0034] 1A and 1B illustrate diagrams 100 and 120 of suggestion modules of computing device 104 and / or computing device 112 that enable user 102 to receive suggested information in a timely manner without a direct and explicit request from user 102 in certain situations. In some implementations, a direct explicit request from a user to an application may include the user identifying the application, speaking one or more predefined invocation terms to invoke the application (e.g., “Hey, Assistant…”), tapping a touch input to invoke the application (e.g., tapping a “home” button on a cell phone), or performing other direct explicit requests. In some implementations, computing device 104 may be a primary device, and computing device 112 may be an auxiliary computing device in communication with computing device 104. Computing device 112 may process contextual data related to situations in which user 102 may currently be and / or may have previously been, with prior permission from user 102. For example, the computing device 112 may include one or more interfaces for capturing data characterizing the situation of the user 102. In some implementations, the computing device 112 may be a pair of computerized glasses that includes a display interface, a camera, a speaker, and / or a microphone. The microphone and / or camera may capture audio and / or image data based on the environment of the user 102 in certain pre-approved situations.
[0035] For example, the microphone may capture audio data characterizing the audio of the user 102 and / or a nearby person openly speaking a query to one or more other persons. The query may be, for example, "Do you guys know if the trains are running today?", which may be a request for the user 102 to provide an answer to the query. The audio data characterizing the query may undergo processing 108 at the computing device 112 and / or the computing device 104 to identify information responsive to the query. For example, the audio data may undergo voice processing to identify the natural language content of the query and select one or more sources of information from which to select a subset of response data. When the response information is identified in operation 110, a suggestion module of the computing device 112 and / or the computing device 104 may determine whether one or more characteristics of the situation satisfy one or more conditions for rendering suggested information to the user 102.
[0036] For example, the computing device 112 may determine that the user 102 has performed a gesture 122 that may satisfy one or more conditions for rendering the suggestion. In some implementations, the detected gesture may be an audible sound that may or may not correspond to a word or phrase that suggests the user 102 is willing to receive the suggestion. Alternatively or additionally, the detected gesture 122 may be a physical movement of the user 102 (e.g., a non-verbal gesture) that is determined to suggest the user 102 is willing to receive the suggestion. For example, the user 102 may make a sound such as "hmm..." that indicates that the user 102 is considering responding to a query from another person. This sound or disfluency may indicate to the suggestion module that the user 102 is attempting to think of an appropriate response for the other person.
[0037] The suggestion module may determine 124 that one or more conditions for providing the suggested information are met and, in response, render an output 126 that embeds the suggested information. For example, the computing device 112 may render an audible output to the ear of the user 102, which may be audible only to the user 102 because the user 102 is wearing the computing device 112. The audible output may be, for example, "According to the city website, there are no trains running today due to construction." When the user 102 receives this suggested information, the user may then provide a response 128 to another person, such as "There are no trains running today due to construction." In some implementations, the suggestion module may determine whether the user 102 utilized the suggested information and may modify one or more suggestion processes accordingly. For example, one or more trained machine learning modules may be trained based on the user 102 utilizing the suggested information in the scenarios of FIGS. 1A and 1B. Thereafter, when the user 102 is in a similar situation, the suggestion module may generate suggested information based on previous instances in which the user 102 utilized suggested information in similar situations.
[0038] 2A and 2B show diagrams 200 and 220 of a user 202 wearing a pair of earphones and being assisted by a suggestion module that may provide suggestions to the user 202 in a timely manner without the user necessarily explicitly requesting the suggestions. For example, the user 202 may be at a geographic location corresponding to a rental car facility. Prior to arriving at the rental car facility, the user 202 may contact a rental car application to reserve the vehicle 206 with a specific key code for boarding the vehicle 206. In some implementations, the computing device 214 and / or the auxiliary computing device 204 may store and / or provide access to information related to reservations made by the user 202. The suggestion module of the computing device may process the reservation and / or other contextual data, with prior permission from the user 202, to assist the user while the user 202 is at the rental car facility.
[0039] For example, user 202 may perform a verbal gesture 208 or possibly make a sound (e.g., a disfluency) indicating that the user may be trying to remember something. The verbal gesture 208 may be, "My truck entry code is... hmm..." and may be detected by the auxiliary computing device 204 via its microphone. Alternatively or additionally, contextual data, which may include audio data, may be captured by the auxiliary computing device 204 and the primary computing device 214 to identify characteristics of the user's 202 situation. The contextual data, via process 210, may determine that the user 202 is at a rental facility location and that the user 202 is exhibiting a pause or other period of silence. Based on this determination, the suggestion module may perform operation 212 to identify information that may assist the user 202 in the current situation.
[0040] For example, the suggestion module may access reservation data associated with the user 202's location with prior permission from the user 202. Alternatively or additionally, the suggestion module may identify one or more types of data that one or more other users may have accessed while at the user 202's location and / or similar locations. For example, the context data may be processed using one or more trained machine learning models that have been trained according to historical interaction data. The historical interaction data may characterize one or more prior interactions between the user (and / or one or more other users) and one or more computing devices. These prior interactions may include prior instances when the user 202 and / or one or more other users accessed reservation data and / or key code data while at a rental car facility. Thus, when the context data is processed using the one or more trained machine learning models, a determination may be made that the user 202 is likely to benefit from being rendered key code data to the user.
[0041] When the suggestion module determines that the key code data is suitable for suggesting to the user 202, the suggestion module may cause the auxiliary computing device 204 (e.g., earphones, a smartwatch, and / or other wearable device) to audibly render the key code as subsequent content (e.g., “8339,” or “The code for your track is 8339”). Alternatively or additionally, the suggestion module may determine whether one or more conditions are met before causing the auxiliary computing device 204 to render the suggestion output 222. For example, the computing device 214 and / or the auxiliary computing device 204 may determine whether one or more characteristics of the user's 202 situation satisfy one or more conditions for rendering the suggested information. In some implementations, the one or more conditions may include a duration since the user 202 provided a disfluency (e.g., “Hmm…”) without providing another verbal utterance. When the duration since the disfluency meets a threshold duration, the suggestion module may cause the auxiliary computing device 204 to render the suggestion output 222.
[0042] 3 illustrates a system 300 that enables an auxiliary computing device to timely render content to a user when a user is predicted to desire personal assistance and / or assistance when assisting another person. The automated assistant 304 may operate as part of an assistant application provided on one or more computing devices, such as the computing device 302 and / or a server device. The user can interact with the automated assistant 304 through an assistant interface 320, which may be a microphone, a camera, a touchscreen display, a user interface, and / or any other device capable of providing an interface between a user and an application. For example, the user can initialize the automated assistant 304 by providing linguistic, textual, and / or graphical input to the assistant interface 320 to cause the automated assistant 304 to initiate one or more actions (e.g., provide data, control a peripheral device, access an agent, generate input and / or output, etc.). Alternatively, the automated assistant 304 may be initialized based on processing of context data 336 using one or more trained machine learning models. The context data 336 may characterize one or more features of an environment to which the automated assistant 304 has access and / or one or more characteristics of a user who is predicted to intend to interact with the automated assistant 304. The computing device 302 may include a display device, which may be a display panel including a touch interface for receiving touch inputs and / or gestures to enable a user to control the applications 334 of the computing device 302 via the touch interface.In some implementations, the computing device 302 lacks a display device and can thereby provide audible user interface output without providing graphical user interface output. Additionally, the computing device 302 can provide a user interface, such as a microphone, for receiving verbal natural language input from a user. In some implementations, the computing device 302 can include a touch interface and may lack a camera, but may optionally include one or more other sensors.
[0043] The computing device 302 and / or other third-party client devices may communicate with the server device over a network, such as the Internet. Additionally, the computing device 302 and any other computing devices may communicate with each other over a local area network (LAN), such as a Wi-Fi network. The computing device 302 may offload computational tasks to the server device to conserve computational resources at the computing device 302. For example, the server device may host the automated assistant 304, and / or the computing device 302 may send inputs received at one or more assistant interfaces 320 to the server device. However, in some implementations, the automated assistant 304 may be hosted at the computing device 302, and various processes that may be associated with automated assistant operation may be performed at the computing device 302.
[0044] In various implementations, all or less than all aspects of the automated assistant 304 may be implemented on the computing device 302. In some of these implementations, aspects of the automated assistant 304 are implemented via the computing device 302 and may interface with a server device that may implement other aspects of the automated assistant 304. The server device may optionally serve multiple users and their associated auxiliary applications via multiple threads. In implementations in which all or less than all aspects of the automated assistant 304 are implemented via the computing device 302, the automated assistant 304 may be an application that is separate from (e.g., installed “on top of”) the operating system of the computing device 302, or alternatively, may be implemented directly by (e.g., considered an application of, but integral with, the operating system).
[0045] In some implementations, the automated assistant 304 can include an input processing engine 306, which can employ multiple different modules to process input and / or output for the computing device 302 and / or the server device. For example, the input processing engine 306 can include a speech processing engine 308, which can process audio data received at the assistant interface 320 to identify text embodied in the audio data. The audio data can be transmitted from the computing device 302 to a server device, for example, to conserve computational resources at the computing device 302. Additionally or alternatively, the audio data can be processed exclusively at the computing device 302.
[0046] The process for converting audio data to text can include speech recognition algorithms that can employ neural networks and / or statistical models to identify groups of audio data that correspond to words or phrases. The text converted from the audio data can be parsed by a data analysis engine 310 and made available to the automated assistant 304 as text data that can be used to generate and / or identify command phrases, intents, actions, slot values, and / or any other content specified by the user. In some implementations, output data provided by the data analysis engine 310 can be provided to a parameter engine 312 to determine whether the user has provided input that corresponds to a particular intent, action, and / or routine that can be performed by the automated assistant 304 and / or an application or agent that can be accessed via the automated assistant 304. For example, assistant data 338 can be stored at the server device and / or computing device 302 and can include data defining one or more actions that can be performed by the automated assistant 304, as well as parameters necessary to perform those actions. The parameter engine 312 can generate one or more parameters for the intent, action, and / or slot value and provide the one or more parameters to the output generation engine 314. The output generation engine 314 can use the one or more parameters to communicate with the assistant interface 320 to provide output to the user and / or to communicate with one or more applications 334 to provide output to the one or more applications 334.
[0047] In some implementations, the automated assistant 304 may be an application that may be installed “on top of” the operating system of the computing device 302 and / or may itself form part of (or the entirety of) the operating system of the computing device 302. The automated assistant application includes and / or has access to on-device speech recognition, on-device natural language understanding, and on-device fulfillment. For example, on-device speech recognition may be performed using an on-device speech recognition module that processes audio data (detected by a microphone) using an end-to-end speech recognition machine learning model stored locally on the computing device 302. The on-device speech recognition generates recognized text for verbal utterances (if any) present in the audio data. Also, for example, on-device natural language understanding (NLU) may be performed using an on-device NLU module that processes the recognized text generated using on-device speech recognition and, optionally, contextual data, to generate NLU data.
[0048] The NLU data can include an intent corresponding to the verbal utterance and, optionally, parameters for that intent (e.g., slot values). On-device fulfillment can be implemented using an on-device fulfillment module that utilizes the NLU data (from the on-device NLU) and, optionally, other local data, to determine actions to take to analyze the intent of the verbal utterance (and, optionally, parameters for that intent). This can include determining local and / or remote responses (e.g., answers) to the verbal utterance, interactions with locally installed applications to perform based on the verbal utterance, commands to send to Internet-of-Things (IoT) devices (directly or via corresponding remote systems) based on the verbal utterance, and / or other analytical actions to perform based on the verbal utterance. The on-device fulfillment can then initiate local and / or remote implementation / execution of the actions determined to analyze the verbal utterance.
[0049] In various implementations, remote speech processing, remote NLU, and / or remote fulfillment may be utilized at least selectively. For example, recognized text may be at least selectively sent to a remote automated assistant component for remote NLU and / or remote fulfillment. For example, recognized text may be sent for remote fulfillment optionally in parallel with on-device fulfillment or in response to failure of on-device NLU and / or on-device fulfillment. However, on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device execution may be prioritized at least due to the reduced latency they offer when analyzing spoken speech (by not requiring a client-server round trip to analyze the spoken speech). Furthermore, on-device functionality may be the only functionality available in situations where network connectivity is absent or limited.
[0050] In some implementations, the computing device 302 may include one or more applications 334 that may be provided by a third-party entity different from the entity that provided the computing device 302 and / or the automated assistant 304. The application state engine of the automated assistant 304 and / or the computing device 302 may access the application data 330 to determine one or more actions that may be performed by the one or more applications 334, as well as the state of each application of the one or more applications 334 and / or the state of each device associated with the computing device 302. The device state engine of the automated assistant 304 and / or the computing device 302 may access the device data 332 to determine one or more actions that may be performed by the computing device 302 and / or one or more devices associated with the computing device 302. Additionally, the application data 330 and / or any other data (e.g., device data 332) may be accessed by the automated assistant 304 to generate context data 336, which may characterize the circumstances under which a particular application 334 and / or device is running and / or the circumstances under which a particular user is accessing the computing device 302, the application 334, and / or any other device or module.
[0051] While one or more applications 334 are executing on the computing device 302, the device data 332 may characterize the current operational state of each application 334 executing on the computing device 302. Additionally, the application data 330 may characterize one or more characteristics of the executing applications 334, such as the content of one or more graphical user interfaces being rendered toward the one or more applications 334. Alternatively or additionally, the application data 330 may characterize action schemas, which may be updated by the respective applications and / or by the automated assistant 304 based on the current operational status of the respective applications. Alternatively or additionally, the one or more action schemas for one or more applications 334 may remain static but may be accessed by the application state engine to determine suitable actions to initiate via the automated assistant 304.
[0052] The computing device 302 can further include an assistant invocation engine, which can use one or more trained machine learning models to process the application data 330, the device data 332, the context data 336, and / or any other data accessible to the computing device 302. The assistant invocation engine can process this data to determine whether to wait for the user to explicitly speak an invocation phrase to invoke the automated assistant 304, or the data can be considered to indicate the user's intent to invoke the automated assistant instead of requiring the user to explicitly speak an invocation phrase. For example, the one or more trained machine learning models can be trained using training data examples based on scenarios in which a user is in an environment with multiple devices and / or applications exhibiting various operating states. The training data examples can be generated to capture training data characterizing situations in which a user invokes an automated assistant and other situations in which the user does not invoke an automated assistant. When one or more trained machine learning models are trained according to these examples of training data, the assistant invocation engine can cause the automated assistant 304 to detect or limit detection of spoken invocation phrases from the user based on situational and / or environmental characteristics.
[0053] In some implementations, the computing device 302 may be a wearable device, such as a pair of earphones, a watch, and / or a pair of glasses, and / or any other wearable computing device. The computing device 302 may include a language model engine 322 that facilitates access to a transformer neural network model, a recurrent neural network, and / or other language model (LM). For example, the computing device 302 may capture audible sounds via one or more interfaces and process audio data corresponding to the audible sounds using the language model engine 322. In some implementations, the language model engine 322 may employ a transformer neural network model to determine whether a verbal utterance is complete or incomplete and / or the type of entity that can be utilized to complete the verbal utterance. For example, when a user of the computing device 302 and / or another person recites a verbal utterance such as "At 7:30 PM, I should be...", the language model engine 322 may be utilized to determine that an "event" entity type may be utilized to complete the unfinished verbal utterance. The verbal utterance may then be processed to identify suitable content (e.g., a location such as "Cardinal Stadium") corresponding to the entity type to complete the verbal utterance. In some implementations, the content may be identified using a public knowledge graph generated based on prior interactions between one or more other persons and one or more other applications. Alternatively or additionally, the content may be identified using a private knowledge graph generated based on prior interactions between the user and one or more applications.
[0054] In some implementations, the computing device 302 may include a verbal / disfluency gesture engine 324. The verbal / disfluency gesture engine 324 may be utilized to process input data from one or more sensors to determine whether the user has indicated a willingness to receive suggested information and / or content. For example, disfluencies embodied in audio data (e.g., “Hmm…”) may be identified by the verbal / disfluency gesture engine 324 and utilized as factors for determining whether to provide the suggested information in a timely manner at a given point in time. Other factors may include the user's context and / or prior interactions in which the user subsequently utilized information that may not have been explicitly requested by the user. Alternatively or additionally, the verbal / disfluency gesture engine 324 may process input from other sensors, such as motion sensors capable of characterizing the user's movements and / or one or more of the user's extremities. Certain movements that may affect an accelerometer or inertial sensor may be detected and considered as factors when determining whether the user would benefit from unsolicited information. For example, a user making their way through an airport may exhibit abrupt pauses in movement as they rummage through their bag searching for their identification. In response, the computing device 302 may determine that the user is searching for a particular type of information and cause the output interface to render that information to the user in a timely manner (e.g., via a set of earphones worn by the user). In some implementations, the verbal / disfluency gesture engine 324 may process data from multiple sources to determine characteristics of verbal and / or non-verbal gestures. For example, the prosody of speech and / or rate of speech may be detected by the verbal / disfluency gesture engine 324 to determine whether the user would benefit from unsolicited rendering of suggested information.In this way, when a user who normally speaks quickly is currently speaking slowly (e.g., due to a lack of memory), the computing device 302 may determine that the user may need assistance in recalling certain information. Alternatively or additionally, when a user who generally speaks in a lower voice is currently speaking in a higher voice, the computing device 302 may determine that the user may need assistance in recalling certain information.
[0055] In some implementations, the computing device 302 may include a voice filtering engine 316 that can be utilized when processing audio data. The voice filtering engine 316 may be used to remove voices from one or more persons from a portion of audio data to isolate specific voices from a particular person. For example, a person reciting a question in a crowded room may, in response to detecting the person's voice, cause the voice filtering engine 316 to filter the voice from the crowd, thereby allowing the person's question to be accurately processed. The language model engine 322 may then process the question to generate suggested information that the user can recite to answer the person's question, at least if the user expresses a willingness to receive such information. In some implementations, the computing device 302 may include a condition processing engine 318 to determine whether one or more conditions are met for rendering suggested information to the user. In some implementations, the one or more conditions may vary for certain situations and / or certain users and / or may be determined using a heuristic process and / or one or more trained machine learning models. For example, a user who declines to receive unsolicited suggestions in a particular location can cause the condition processing engine 318 to generate location-based conditions for rendering the suggested information. In some implementations, the output of the verbal / disfluency gesture engine 324 can be processed to determine whether one or more conditions are met before rendering the suggested information to the user.
[0056] In some implementations, the computing device 302 may include a situation processing engine 326 that may process context data 336 to identify appropriate content to render to the user and / or determine whether the user is willing to receive content without an explicit request. For example, the user's location may be an indication of whether the user is willing to receive preemptive information via their wearable device. For example, a user may prefer to receive unsolicited information outside a radius of their home. Thus, determining the user's geographic location, with prior permission from the user, may assist in determining whether to render certain unsolicited information to the user. In some implementations, the computing device 302 may include an output generation engine 314 that may determine characteristics of the output to be rendered. For example, the output generation engine 314 may cause the suggested information to be rendered via an audio output interface at a rate based on the user's speaking rate. In some implementations, the rendering of the suggested information may be momentarily paused and / or permanently stopped based on changes in the user's situation and / or how the user responds to the rendering of the suggested information. For example, when a user recites aloud information suggested through speech, rendering of the suggested information may be temporarily paused until the user completely recites the portion of the suggested information that has been rendered so far. Alternatively or additionally, when the user performs verbal and / or non-verbal gestures indicating that the user is not interested in the suggested information being rendered, the output generation engine 314 may stop rendering the suggested information.
[0057] FIG. 4 illustrates a method 400 for providing supplemental information via a wearable device in a timely manner in certain situations where a user may or may not have explicitly requested such information from the wearable device. Method 400 may be implemented by one or more computing devices, applications, and / or any other devices or modules that may be associated with an automated assistant. Method 400 may include operation 402 of determining whether the user is wearing the device in a situation in which supplemental information is available. The supplemental information may be associated with the user's situation and / or retrieved via a primary computing device in communication with a device worn by the user (i.e., an auxiliary computing device). For example, a user may be visiting a storage facility that the user rents, and the storage facility may have a keypad for entering a passcode that may provide access to enter the storage facility. When the user approaches the keypad, the user may be wearing a pair of computerized earphones and / or eyeglasses and / or a smartwatch. In this situation, the primary computing device may determine, based on data available to the primary computing device, that the user is wearing the computerized device during a situation in which supplemental information (e.g., a passcode) is available. The passcode may be available in memory of the primary computing device, the auxiliary computing device, and / or other computing devices available via a network connection.
[0058] When it is determined that the user is wearing the computerized device while in this situation, method 400 may proceed from operation 402 to operation 404, where output data based on the response information is generated. For example, the response information may be based on one or more different sources of data, such as one or more applications and / or devices accessible via the primary computing device and / or the auxiliary computing device. In some implementations, one or more sources of data may be selected to provide information based on how each source of data ranks for the user's situation. For example, a particular application may be prioritized over other applications as a source of information based on the user's location, the content of audio captured by the auxiliary computing device, the content of images captured by the auxiliary computing device, and / or any other data that may be available to the computing device. For example, a security application associated with a keypad may be prioritized over other applications, and thus the security application may be selected as a source of information in a situation in which the user is approaching the keypad.
[0059] From operation 404, method 400 may proceed to operation 406, which determines whether one or more characteristics of the situation satisfy one or more conditions for unsolicited rendering of information to the user. The one or more conditions for rendering information may be selected for a particular user, a particular situation, and / or a particular input to the auxiliary computing device and / or the primary computing device. For example, the one or more conditions may be determined to be satisfied when the user exhibits a certain type of pause in their speech and / or movement. In some implementations, the one or more conditions may be satisfied for a pause in movement (e.g., a pause that creates a detectable amount of inertia) and / or speech that lasts for a threshold duration, and / or other input data (e.g., ambient noise, another person's voice, detected distance from an object and / or location, etc.). These situation characteristics may be considered different from intentional gestures, such as the user's verbal input and / or hand movements, but may still cause method 400 to perform operation 408.
[0060] In some implementations, one or more conditions for rendering information may be established over time as a result of training one or more machine learning models. For example, after certain information has been rendered, a user may provide an affirmation indicating whether the certain information was useful to the user and / or whether it was provided on time. The affirmative affirmation may be, for example, a gesture interpreted by the primary computing device and / or the auxiliary computing device as an indication that the provision of the information was useful to the user. In some implementations, the gesture may include a recitation of at least a portion of the information provided via the auxiliary computing device and / or a physical movement of the user (e.g., a nod) in response to receiving the information via the auxiliary computing device. Such user feedback may then be utilized to further train one or more machine learning models to provide information on time via the auxiliary computing device.
[0061] When the characteristics of the situation satisfy one or more conditions for unsolicited rendering of information to the user, method 400 may proceed from operation 406 to operation 408. Otherwise, method 400 may proceed from operation 406 to operation 410. Operation 408 may include causing the auxiliary computing device to render output for the user based on the output data. For example, the auxiliary computing device may include one or more interfaces for rendering the situation, and the output may be rendered via the one or more interfaces. When the auxiliary computing device is a pair of earphones, the output may be audible content (e.g., a passcode for a keypad) rendered via one or more speakers of the pair of earphones. Alternatively or additionally, when the user is wearing a pair of computerized eyeglasses, the output may be visually and / or audibly rendered via one or more interfaces of the eyeglasses.
[0062] When the situation characteristics do not satisfy one or more conditions, method 400 may proceed to operation 410, which may include determining whether the user has performed an intentional gesture indicating an intent to receive information. For example, an intentional gesture may be any intentional and direct input from the user to the auxiliary computing device and / or the primary computing device. Following the foregoing example, when the auxiliary computing device is ready to render output to the user, the user may perform a gesture, such as saying the phrase "what's this?" and / or pointing to the keypad, such that the camera of the auxiliary computing device captures the gesture. When it is determined that the user has provided an intentional gesture, method 400 may proceed to operation 408. Otherwise, method 400 may return to operation 402.
[0063] In some implementations, method 400 may include the optional operation 412 of training one or more machine learning models according to interactions between the user and the primary computing device and / or the auxiliary computing device. For example, a suggestion module of the primary computing device and / or the auxiliary computing device may provide information to the user in a particular situation without being directly requested. When the user responds affirmatively or positively to providing information, the one or more models may be trained to prioritize a particular type of information in that situation. However, when the user does not respond affirmatively or positively to providing information, the one or more models may be trained not to provide a particular type of information in that situation.
[0064] 5 is a block diagram 500 of an exemplary computer system 510. The computer system 510 generally includes at least one processor 514 that communicates with several peripheral devices via a bus subsystem 512. These peripheral devices may include, for example, a storage subsystem 524 including memory 525 and a file storage subsystem 526, a user interface output device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices enable user interaction with the computer system 510. The network interface subsystem 516 provides an interface to external networks and is coupled to corresponding interface devices in other computer systems.
[0065] The user interface input devices 522 may include a keyboard, a pointing device such as a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touchscreen integrated into a display, a voice recognition system, an audio input device such as a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computer system 510 or onto a communications network.
[0066] The user interface output devices 520 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visual image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computer system 510 to a user or to another machine or computer system.
[0067] Storage subsystem 524 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 524 may include logic for performing selected aspects of method 400 and / or for implementing one or more of system 300, computing device 104, auxiliary computing device 112, computing device 214, auxiliary computing device 204, and / or any other applications, devices, apparatuses, and / or modules discussed herein.
[0068] These software modules are generally executed by the processor 514 alone or in combination with other processors. The memory 525 used in the storage subsystem 524 can include several memories, including a main random access memory (RAM) 530 for storing instructions and data during program execution, and a read-only memory (ROM) 532 in which fixed instructions are stored. The file storage subsystem 526 can provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of certain implementations can be stored in the storage subsystem 524 by the file storage subsystem 526 or in other machines accessible by the processor 514.
[0069] Bus subsystem 512 provides a mechanism for allowing the various components and subsystems of computer system 510 to communicate with each other as intended. Although bus subsystem 512 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0070] Computer system 510 can be of different types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computer system 510 shown in Figure 5 is only a specific example for purposes of illustrating some implementations. Many other configurations of computer system 510 are possible, having more or fewer components than the computer system shown in Figure 5.
[0071] In situations where the systems described herein may collect or utilize personal information about users (or sometimes referred to herein as “participants”), users may be provided with an opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographic location) or to control whether and / or how to receive content from content servers that may be more relevant to the user. Also, certain data may be treated in one or more ways before being stored or used so that personally identifiable information is removed. For example, a user's identity may be treated so that personally identifiable information cannot be determined about the user, or so that if geographic location information is obtained, the user's geographic location may be generalized (such as to the city level, ZIP code level, or state level) so that the user's specific geographic location cannot be determined. Thus, users can control how information is collected and / or used about them.
[0072] While several implementations have been described and illustrated herein, various other means and / or structures for performing the functions and / or obtaining the results and / or one or more of the advantages described herein may be utilized, and each such variation and / or modification is considered to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on the specific application or applications for which the teachings are used. Those skilled in the art will recognize, or will be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. Accordingly, the foregoing implementations are presented by way of example only, and it should be understood that, within the scope of the appended claims and their equivalents, implementations may be practiced other than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
[0073] In some implementations, a method implemented by one or more processors is described as including operations such as performing automatic speech recognition on streaming audio data captured via one or more microphones of a wearable device worn by a user to generate recognized text hypotheses for a partial spoken utterance. The partial spoken utterance is captured in the streaming audio data and provided by the user as part of a conversation with at least one additional person. The method further includes processing the recognized text hypotheses using one or more machine learning models to generate subsequent content output indicating one or more characteristics of further content predicted to follow the partial spoken utterance. The method further includes determining additional text predicted to follow the partial spoken utterance based on the subsequent content output. The method further includes determining whether to render the additional text via at least one interface of the auxiliary device. The determining whether to render the additional text is independent of any explicit user interface input directly requesting rendering and is based on whether a disfluency is detected following the subsequent content output and / or the partial spoken utterance. The method further includes, in response to determining to render the additional text, causing the additional text to be rendered via at least one interface of the auxiliary device.
[0074] These and other implementations of the techniques disclosed herein may include one or more of the following features.
[0075] In some implementations, determining whether to render the additional text is based on whether a disfluency is detected following the partial verbal utterance. In some of those implementations, determining whether to render the additional text is further based on a measured duration of the disfluency. In some versions of those implementations, determining whether to render the additional text includes determining to render the additional text based on the measured duration satisfying a threshold duration. In some versions, determining the additional text occurs prior to the threshold duration being satisfied. In some implementations that determine whether to render the additional text based on disfluency, the method further includes monitoring for disfluency. Monitoring for disfluency may include processing a stream of audio data following the partial verbal utterance and / or processing motion sensor data from one or more motion sensors of the auxiliary device.
[0076] In some implementations, causing the additional text to be rendered includes causing the additional text to be rendered as speech synthesis output via an audio interface of the auxiliary device. In some of these implementations, the method further includes detecting user voice activity during the rendering of the speech synthesis output, and stopping the rendering of the speech synthesis output in response to detecting the user voice activity.
[0077] In some implementations, causing the additional text to be rendered includes causing the additional text to be rendered in multiple separate chunks based on one or more properties of the additional text. In some of these implementations, causing the additional text to be rendered in multiple separate chunks includes causing only a first chunk of the multiple separate chunks to be rendered, and following the rendering of only the first chunk, detecting an onset and then a cessation of voice activity, and in response to detecting the cessation of voice activity, causing only a second chunk of the multiple separate chunks to be rendered.
[0078] In some implementations, determining whether to render the additional text is based on the subsequent content output. In some of these implementations, determining whether to render the additional text based on the subsequent content output includes determining an entropy criterion for the additional text determined based on the subsequent content output, and determining whether to render the additional text according to the entropy criterion. In some versions of these implementations, the entropy criterion is based on an inverse document frequency (IDF) measure of one or more terms of the additional text. In some versions of these implementations, the IDF measure is a personalized user criterion for one or more terms of the additional text based on their frequency of occurrence in speech recognition hypotheses for the user.
[0079] In some implementations, determining whether to render additional text is based on whether a disfluency is detected and also based on subsequent content output.
[0080] In some implementations, determining additional text predicted to follow the partial verbal utterance based on the subsequent content output includes determining a predicted entity type of the further content predicted to follow the partial verbal utterance based on the subsequent content output, conducting a search using the predicted entity type and using one or more hypothesized terms of the recognized text hypotheses, and determining additional text from resources responsive to the search. In some versions of these implementations, the resources responsive to the search are particular database entries, and conducting a search using the predicted entity type and the one or more predicted terms includes determining that the database entries are indexed with both the predicted entity type and the one or more hypothesized terms. In some of these versions, conducting a search using the predicted entity type and the one or more predicted terms includes identifying resources using the one or more hypothesized terms and extracting the additional text from a superset of text of the resources based on the additional text being of the predicted entity type. In some additional or alternative versions of these implementations, the step of conducting the search occurs at a client device in communication with the auxiliary device, and the step of conducting the search includes searching for locally cached resources at the client device. In some of these versions, the method further includes determining a topic of the conversation based on processing a prior verbal utterance preceding the partial verbal utterance, and downloading, at the client device, prior to the partial verbal utterance, at least a portion of resources having a defined relationship to the topic.In some additional or alternative versions of those implementations, the subsequent content output indicates the initial additional text that is different from the additional text, and determining a predicted entity type of the further content based on the subsequent content output includes identifying an entity type of the additional text. In some additional or alternative versions of those implementations, performing the search includes searching one or more personal resources that are personal to the user. In some versions of those additional or alternative versions, searching the one or more personal resources that are personal to the user includes verifying that the partial spoken utterance is from the user, and searching the one or more personal resources conditioned on verifying that the partial spoken utterance is from the user.
[0081] In some implementations, determining additional text predicted to follow the partial spoken utterance is conditioned on verifying that the partial spoken utterance is from the user. In some of these implementations, verifying that the partial spoken utterance is from the user includes performing text-independent speaker verification and / or speaker-dependent automatic speech recognition that facilitates generating predicted hypotheses only when the audio data is from the user.
[0082] In some implementations, a method implemented by one or more processors is described as including operations such as processing, by a computing device, audio data of verbal utterances captured by an auxiliary computing device during an interaction between a user and one or more persons. The user is wearing the auxiliary computing device during the interaction, and the auxiliary computing device communicates the audio data to the computing device via a communication channel. The method further includes determining, by the computing device, based on processing the audio data, that the verbal utterance is incomplete and that information accessible via the computing device supplements the content of the verbal utterance. The method further includes determining, by the computing device, whether one or more characteristics of a situation in which the interaction is occurring satisfy one or more conditions for rendering information to the user via the auxiliary computing device. The method further includes, when it is determined that the one or more characteristics of the situation satisfy one or more conditions for rendering information to the user via the auxiliary computing device, causing the computing device to render an output on an interface of the auxiliary computing device that characterizes information about the user.
[0083] These and other implementations of the techniques disclosed herein may include one or more of the following features.
[0084] In some implementations, determining that one or more characteristics of the situation satisfy one or more conditions for rendering the information includes determining, based on processing the audio data, that a disfluency is embodied in the verbal utterance. In some implementations, determining that one or more characteristics of the situation satisfy one or more conditions for rendering the information includes determining that the disfluency continues for at least a threshold duration of time.
[0085] In some implementations, determining that one or more characteristics of the situation satisfy one or more conditions for rendering the information includes determining that the user has performed a linguistic gesture indicating that the user can receive suggested content via a suggestion module accessible to the user via the computing device.
[0086] In some implementations, determining that one or more characteristics of the situation satisfy one or more conditions for rendering the information includes determining that the user has performed a non-verbal gesture indicating that the user is available to receive the suggested content via a suggestion module accessible to the user via the computing device.
[0087] In some implementations, the auxiliary computing device is a pair of earphones, and the audio data is captured via an audio input interface of at least one earphone of the pair of earphones.
[0088] In some implementations, the auxiliary computing device is a pair of computerized glasses, and the audio data is captured via an audio input interface of the computerized glasses, and the output is rendered via a display interface of the computerized glasses.
[0089] In some implementations, determining that information accessible via the computing device is responsive to a portion of content characterized by the audio data includes determining, based on processing the audio data using one or more trained machine learning models, that the information corresponds to an entity type of information previously utilized by one or more other users when responding to a query of a type corresponding to that portion of content.
[0090] In some implementations, the method further includes, when it is determined that one or more characteristics of the situation do not satisfy one or more conditions for rendering information to the user via the auxiliary computing device, bypassing, by the computing device, rendering of output characterizing information about the user on an interface of the auxiliary computing device.
[0091] In some implementations, processing the audio data of the verbal utterance captured by the auxiliary computing device includes processing the audio data using a Transformer Neural Network model that has been trained using the natural language content data.
[0092] In some implementations, a method implemented by one or more processors is described as including operations such as determining that one or more characteristics of a user's situation satisfy one or more conditions for rendering information to the user via an auxiliary computing device without any direct request from the user. The auxiliary computing device is a wearable device worn by the user in the situation, and the one or more characteristics of the situation indicate that the user or another person is exhibiting a memory lapse. The method further includes processing input data characterizing one or more inputs captured by the auxiliary computing device via one or more input interfaces of the auxiliary computing device based on the one or more characteristics of the situation satisfying the one or more conditions. The method further includes determining, based on the processing of the input data, that particular information accessible via the auxiliary computing device corresponds to a portion of content characterized by the input data. The method further includes causing an output interface of the auxiliary computing device to render output characterizing the particular information to the user.
[0093] These and other implementations of the techniques disclosed herein may include one or more of the following features.
[0094] In some implementations, the auxiliary computing device includes computerized glasses, and the output is rendered through a graphical user interface (GUI) of the computerized glasses. In some implementations, the one or more inputs include an image of the situation captured via a camera of the auxiliary computing device.
[0095] In some implementations, the auxiliary computing device includes a pair of earphones, and the output is rendered through an audio interface of one or more earphones of the pair of earphones.
[0096] In some implementations, the one or more characteristics of the situation include an audible disfluency captured by an audio interface of the auxiliary computing device.
[0097] In some implementations, a method implemented by one or more processors is described as including operations such as causing an output interface of the auxiliary computing device to render, by the auxiliary computing device, output characterizing information that may be utilized by a user in the user's current situation. The method further includes determining whether the user has provided a verbal utterance embodying a portion of the information embodied in the output from the auxiliary computing device while the auxiliary computing device is rendering the output. The method further includes causing the output interface of the auxiliary computing device to pause rendering of the output characterizing the information when it is determined that the user has provided a verbal utterance embodying at least that portion of the information, and causing the output interface of the auxiliary computing device to resume rendering of additional output embodying another portion of the information when the user is no longer providing the verbal utterance. The method further includes causing the output interface of the auxiliary computing device to stop rendering of the output characterizing the information when it is determined that the user has provided another verbal utterance that does not embody at least that portion of the information.
[0098] These and other implementations of the techniques disclosed herein may include one or more of the following features.
[0099] In some implementations, the method further includes processing audio data corresponding to audible sounds in the user's current situation prior to causing the output interface to render the output, where processing the audio data includes filtering out certain audible sounds that do not correspond to the user's voice.
[0100] In some implementations, the method further includes generating embedding data based on the user's current situation, and causing the output interface to render the output is performed when the embedding data corresponds to an embedding that is a threshold distance from the particular embedding in the latent space. In some of these implementations, the particular embedding in the latent space is based on a personal knowledge graph associated with prior interactions between the user and one or more applications. In some other versions of these implementations, the particular embedding in the latent space is based on a public knowledge graph associated with prior interactions between one or more other users and one or more applications.
[0101] In some implementations, the method further includes processing content data using a Transform Neural Network model that has been trained using natural language content prior to causing the output interface to render the output, the content data characterizing audible sounds in the user's current situation, and the output being based on processing the content data. [Explanation of symbols]
[0102] 102 users 104 Computing Devices 108 Processing 110 operation 112 Computing devices, auxiliary computing devices 122 Gestures 124 Verdict 126 Output 128 responses 202 users 204 Assistive Computing Devices 206 vehicles 208 Linguistic Gestures 210 Process 214 Computing Devices, Primary Computing Devices 222 Suggested Output 300 System 302 Computing Devices 304 Automated Assistant 306 Input Processing Engine 308 Audio Processing Engine 310 Data Analysis Engine 312 Parameter Engine 314 Output Generation Engine 316 Voice Filtering Engine 318 Condition Processing Engine 320 Assistant Interface 322 Language Model Engine 324 Verbal / Disfluency Gesture Engine 326 Situation Processing Engine 330 Application Data 332 Device Data 334 Applications 336 Context Data 338 Assistant Data 400 ways 510 Computer Systems 512 Bus Subsystem 514 processor 516 Network Interface Subsystem 520 User Interface Output Device 522 User Interface Input Devices 524 Memory Subsystem 525 memory 526 File Storage Subsystem 530 Main Random Access Memory (RAM) 532 Read-Only Memory (ROM)
Claims
1. 1. A method implemented by one or more processors, comprising: performing automatic speech recognition on streaming audio data captured via one or more microphones of a wearable device worn by a user to generate recognized text hypotheses for partial spoken utterances, the partial spoken utterances being captured in the streaming audio data and provided by the user as part of a conversation with at least one other person; processing the recognized text hypotheses using one or more machine learning models to generate subsequent content output indicative of one or more characteristics of further content predicted to follow the partial spoken utterance; determining additional text that is predicted to follow the partial spoken utterance based on the subsequent content output; determining whether to render the additional text via at least one interface of an auxiliary device; the determining whether to render the additional text is independent of any explicit user interface input that directly requests rendering; said subsequent content output; or Whether disfluencies are detected following the partial oral utterance determining the identity of the subject based on one or both of the following: In response to determining to render the additional text, causing the additional text to be rendered via the at least one interface of the auxiliary device; and A method comprising:
2. The method of claim 1 , wherein the determining whether to render the additional text is based on whether the disfluency is detected following the partial spoken utterance.
3. The method of claim 2 , wherein the determining whether to render the additional text is further based on a measured duration of the disfluency.
4. 4. The method of claim 3, wherein the determining whether to render the additional text comprises determining to render the additional text based on the measured duration satisfying a threshold duration, and wherein determining the additional text occurs prior to the threshold duration being satisfied.
5. further comprising monitoring the disfluency, wherein the monitoring the disfluency comprises: processing the stream of audio data following the partial spoken utterance; or processing motion sensor data from one or more motion sensors of the auxiliary device; 5. The method of claim 1, comprising one or both of:
6. causing the rendering of the additional text includes causing the additional text to be rendered as speech synthesis output via an audio interface of the auxiliary device; detecting voice activity of the user during rendering of the speech synthesis output; in response to detecting the voice activity of the user, stopping the rendering of the speech synthesis output; 6. The method of claim 1, further comprising:
7. the step of causing the additional text to be rendered comprises: causing the additional text to be rendered in a plurality of separate chunks based on one or more properties of the additional text; 6. The method according to any one of claims 1 to 5.
8. causing the rendering of the additional text in a plurality of separate chunks, causing only a first chunk of the plurality of individual chunks to be rendered; detecting an onset and then cessation of voice activity following said rendering of only said first chunk; in response to detecting the cessation of the voice activity, causing rendering of only a second chunk of the plurality of individual chunks; 8. The method of claim 7, comprising:
9. The step of determining whether to render the additional text is based on the subsequent content output, and the step of determining whether to render the additional text based on the subsequent content output comprises: determining an entropy metric for the additional text determined based on the subsequent content output; and determining whether to render the additional text depending on the entropy criterion.
10. The method of claim 9 , wherein the entropy criterion is based on an inverse document frequency (IDF) measure of one or more terms of the additional text.
11. 11. The method of claim 10, wherein the IDF measure is a personalized user measure based on a frequency of occurrence in speech recognition hypotheses for the user for one or more terms of the additional text.
12. 12. The method of claim 1, wherein the determining whether to render the additional text is based on whether the disfluency is detected and also based on the subsequent content output.
13. determining the additional text predicted to follow the partial spoken utterance based on the subsequent content output; determining a predicted entity type of the further content predicted to follow the partial verbal utterance based on the subsequent content output; conducting a search using the predicted entity type and using one or more hypothesis terms of the recognized text hypotheses; and determining the additional text from resources responsive to the search.
14. the resource responsive to the search is a particular database entry; performing the search using the predicted entity type and the one or more hypothesis terms; determining that the database entry is indexed with both the predicted entity type and the one or more hypothesis terms; The method of claim 13.
15. conducting the search using the predicted entity type and using the one or more hypothesis terms; identifying the resource using the one or more hypothesis terms; and extracting the additional text from a superset of text of the resource based on the additional text being of the predicted entity type.
16. 16. The method of claim 13, wherein the step of performing the search occurs at a client device in communication with the auxiliary device, and the step of performing the search includes searching for resources cached locally at the client device.
17. determining a topic of the conversation based on processing of a previous verbal utterance preceding the partial verbal utterance; downloading, at the client device, at least a portion of the resource having a defined relationship to the topic prior to the partial spoken utterance; 17. The method of claim 16, further comprising:
18. the subsequent content output indicates initial additional text that is different from the additional text; determining the predicted entity type of the further content based on the subsequent content output includes identifying an entity type of the additional text; 18. The method of any one of claims 13 to 17.
19. 18. The method of claim 13, wherein said conducting said search comprises searching one or more personal resources that are personal to said user.
20. said searching said one or more personal resources personal to said user, verifying that the partial spoken utterance is from the user. Including, 20. The method of claim 19, wherein the searching the one or more personal resources is conditioned on verifying that the partial spoken utterance is from the user.
21. 20. The method of claim 1, wherein determining the additional text predicted to follow the partial spoken utterance is conditioned on verifying that the partial spoken utterance is from the user.
22. the step of verifying that the partial spoken utterance is from the user comprises: performing text-independent speaker verification; and / or 20. The method of claim 18 or 19, comprising performing speaker-dependent automatic speech recognition that facilitates the generation of predicted hypotheses only when the audio data is from a user.
23. 1. A method implemented by one or more processors, comprising: processing, by the computing device, audio data of verbal utterances captured by the auxiliary computing device during an interaction between the user and the one or more persons; the user is wearing the auxiliary computing device during the interaction, and the auxiliary computing device communicates the audio data to the computing device over a communication channel; processing the determining, by the computing device, based on the processing of the audio data, that the verbal utterance is incomplete and that information accessible via the computing device supplements the content of the verbal utterance; determining, by the computing device, whether one or more characteristics of the situation in which the interaction is occurring satisfy one or more conditions for rendering the information to the user via the auxiliary computing device; when it is determined that the one or more characteristics of the situation satisfy the one or more conditions for rendering the information to the user via the auxiliary computing device; causing, by the computing device, an interface of the auxiliary computing device to render output characterizing the information about the user; A method comprising:
24. determining that the one or more characteristics of the situation satisfy the one or more conditions for rendering the information, determining, based on said processing of said audio data, that disfluencies are embodied in said oral speech; 24. The method of claim 23, comprising:
25. determining that the one or more characteristics of the situation satisfy the one or more conditions for rendering the information, determining that the disfluency has continued for at least a threshold duration of time.
25. The method of claim 24, comprising:
26. determining that the one or more characteristics of the situation satisfy the one or more conditions for rendering the information, determining that the user has performed a linguistic gesture indicating that the user is available to receive suggested content via a suggestion module accessible via the computing device; 24. The method of claim 23, comprising:
27. determining that the one or more characteristics of the situation satisfy the one or more conditions for rendering the information, determining that the user has performed a non-verbal gesture indicating that the user is available to receive suggested content via a suggestion module accessible via the computing device; 24. The method of claim 23, comprising:
28. 24. The method of claim 23, wherein the auxiliary computing device is a pair of earphones, and the audio data is captured through an audio input interface of at least one earphone of the pair of earphones.
29. the auxiliary computing device is computerized glasses, and the audio data is captured via an audio input interface of the computerized glasses; The output is rendered through a display interface of the computerized glasses.
24. The method of claim 23.
30. determining that the information accessible via the computing device is responsive to the portion of content characterized by the audio data, determining, based on processing the audio data using one or more trained machine learning models, that the information corresponds to an entity type of information previously utilized by one or more other users when responding to a query of a type corresponding to the portion of content.
24. The method of claim 23, comprising:
31. when it is determined that the one or more characteristics of the situation do not satisfy the one or more conditions for rendering the information to the user via the auxiliary computing device; by the computing device bypassing the rendering of the output characterizing the information about the user to the interface of the auxiliary computing device.
24. The method of claim 23, further comprising:
32. processing the audio data of the oral utterance captured by the auxiliary computing device, processing the audio data using a transform neural network model that has been trained using natural language content data; 24. The method of claim 23, comprising:
33. 1. A method implemented by one or more processors, comprising: determining that one or more characteristics of a user's situation satisfy one or more conditions for rendering information to the user via an auxiliary computing device without any direct request from the user; the auxiliary computing device is a wearable device worn by the user in the situation, and the one or more characteristics of the situation indicate that the user or another person is exhibiting a memory lapse; determining the processing input data characterizing one or more inputs captured by the auxiliary computing device via one or more input interfaces of the auxiliary computing device based on the one or more characteristics of the situation satisfying the one or more conditions; determining, based on said processing of said input data, that particular information accessible via said auxiliary computing device corresponds to a portion of content characterized by said input data; causing an output interface of the auxiliary computing device to render an output characterizing the particular information to the user; A method comprising:
34. 34. The method of claim 33, wherein the auxiliary computing device comprises computerized glasses, and the output is rendered via a graphical user interface (GUI) of the computerized glasses.
35. 35. The method of claim 34, wherein the one or more inputs include an image of the situation captured via a camera of the auxiliary computing device.
36. 34. The method of claim 33, wherein the auxiliary computing device includes a pair of earphones, and the output is rendered through an audio interface of one or more earphones of the pair of earphones.
37. 34. The method of claim 33, wherein the one or more characteristics of the situation include an audible disfluency captured by an audio interface of the auxiliary computing device.
38. 1. A method implemented by one or more processors, comprising: causing an output interface of the auxiliary computing device to render, by the auxiliary computing device, output characterizing information available to the user in the user's current situation; determining whether the user provided a verbal utterance that embodied a portion of the information embodied in the output from the auxiliary computing device while the auxiliary computing device was rendering the output; when it is determined that the user has provided the verbal utterance embodying at least the portion of the information; causing the output interface of the auxiliary computing device to pause rendering of the output characterizing the information; causing the output interface of the auxiliary computing device to resume rendering additional output embodying another portion of the information when the user is no longer providing the verbal utterance; when it is determined that the user has provided another verbal utterance that does not embody at least the portion of the information; causing the output interface of the auxiliary computing device to stop rendering the output characterizing the information.
39. processing audio data corresponding to audible sounds in the user's current situation prior to causing the output interface to render the output; further comprising The step of processing the audio data includes filtering out certain audible sounds that do not correspond to the user's voice.
39. The method of claim 38.
40. generating embedded data based on the current situation of the user; further comprising causing the output interface to render the output is performed when the embedded data corresponds to an embedding that is a threshold distance from a particular embedding in latent space.
39. The method of claim 38.
41. 41. The method of claim 40, wherein the particular embedding in a latent space is based on a personal knowledge graph associated with previous interactions between the user and one or more applications.
42. 41. The method of claim 40, wherein the particular embedding in the latent space is based on a public knowledge graph associated with previous interactions between one or more other users and one or more other applications.
43. processing content data using a Transform Neural Network model that has been trained using natural language content prior to causing the output interface to render the output; further comprising the content data characterizing an audible sound in the current situation of the user; the output is based on processing the content data 39. The method of claim 38.
44. 44. A computer program comprising instructions that, when executed by one or more processors of a computing system, cause the computing system to perform the method of any one of claims 1 to 43.
45. 44. A computing system configured to perform the method of any one of claims 1 to 43.
46. 46. The computing system of claim 45, including an auxiliary device and / or a client device separate from but in communication with the auxiliary device.
47. 47. The computing system of claim 46, wherein the client device is a smartphone.