Proactive caching of assistant action content on client devices to enable on-device analysis of spoken or typed utterances

Proactive caching on client devices for automated assistants allows local processing of voice inputs, addressing connectivity issues and reducing latency and bandwidth, enhancing response times and resource efficiency.

JP7747834B2Active Publication Date: 2025-10-01GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024125733
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-05-06
Filing Date
2024-08-01
Publication Date
2025-10-01
Estimated Expiration
2039-05-31

AI Technical Summary

Technical Problem

Existing client-server architectures for automated assistants require continuous online connectivity, consume significant bandwidth, and exhibit latency due to remote processing of voice inputs, which can be problematic in situations with unreliable internet connections.

Method used

Implementing proactive caching of assistant actions on client devices for local processing of voice inputs, using on-device speech recognition and natural language understanding to identify and respond to cached entries without server communication.

Benefits of technology

Reduces latency and bandwidth consumption by enabling quick responses to voice inputs even in offline conditions, optimizing resource usage and improving network performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007747834000001
    Figure 0007747834000001
  • Figure 0007747834000002
    Figure 0007747834000002
  • Figure 0007747834000003
    Figure 0007747834000003
Patent Text Reader

Abstract

To provide a proactive caching of an assistant action content in a client device for enabling an on-device analysis of a speech in oral or to be typed.SOLUTION: By passing through a local proactive caching in a client device of a proactive assistant cache entry and passing through an on-device use of the proactive assistant cache entry, a time required for acquiring a response from an automatic assistant is reduced. When a different proactive cache entry is provided to a difference client device, and which proactive cache entry is provided to which client device, a remote system selects a subset of a cache entry for providing a given client device from a super set of a candidate proactive cache entry.SELECTED DRAWING: Figure 5A
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Humans can engage in human-to-computer interactions with interactive software applications referred to herein as “automated assistants” (also referred to as “digital agents,” “interactive personal assistants,” “intelligent personal assistants,” “assistant applications,” “conversational agents,” etc.). For example, humans (sometimes referred to as “users” when they interact with an automated assistant) can provide commands and / or requests to the automated assistant using oral natural language input (i.e., oral utterances), which in some cases can be converted to text and then processed, and / or by providing textual (e.g., typed) natural language input (i.e., typed utterances). The automated assistant responds to the request by providing a responsive user interface output, which may include audible and / or visual user interface output.

[0002] As described above, many automated assistants are configured to interact via verbal utterances. A user may submit queries and / or commands to an automated assistant interface on a client device via verbal utterances that verbally indicate what information the user is interested in being provided and / or what actions the user is interested in being performed. Typically, the verbal utterances are detected by a microphone on the client device and captured as audio data. The audio data is transmitted to a remote system for further processing. The remote system processes the audio data to determine an appropriate response and transmits the response to the client device.

[0003] Components of the remote system may devote significant computing resources to processing audio data, allowing more complex speech recognition and semantic analysis functionality to be implemented than could otherwise be implemented locally within the client device. However, the client-server approach necessarily requires that the client be online (i.e., communicating with the remote system) when processing voice input. In various situations, continuous online connectivity may not be guaranteed at all times and in all locations; therefore, the client-server voice-based user interface may be disabled within the client device whenever the device is “offline” and therefore not connected to an online service. Furthermore, because the client-server approach requires the transmission of high-bandwidth audio data from the client to components of the remote system, the client-server approach may consume significant bandwidth. Bandwidth consumption is amplified in the typical situation where the remote system is processing requests from a large number of client devices. Furthermore, the client-server approach may exhibit significant latency in provisioning responses to the user, which may prolong the voice-based user-client interaction and utilize client device resources for extended durations. Summary of the Invention [Means for solving the problem]

[0004] Implementations disclosed herein help reduce the time required to obtain a response from an automated assistant through local proactive caching of proactive assistant cache entries at the client device and through on-device utilization of proactive assistant cache entries.

[0005] In various implementations, an automated assistant application running on a client device may use on-device voice processing to process locally detected audio data and generate recognized text corresponding to the spoken utterance (e.g., user interface input) captured by the audio data. The automated assistant application may further utilize the recognized text and / or NLU data generated by on-device natural language understanding (NLU) to identify a locally stored proactive cache entry that corresponds to the spoken utterance. Identification may also be performed based on other user interface input, such as a typed phrase or utterance. A locally stored proactive cache entry may be identified by determining that the recognized text and / or NLU data (from the typed or spoken utterance) matches the assistant request parameters of the proactive cache entry. The automated assistant application may then appropriately respond to the spoken utterance using the action content of the identified proactive cache entry in response to determining a match. The automated assistant application's response may include, for example, executing a deep link included within the action content, rendering text, graphics, and / or audio included within the content (or audio data converted from the response using an on-device speech synthesizer), and / or sending a command to the peripheral device (e.g., via WiFi and / or Bluetooth) to control the peripheral device. Additionally, the automated assistant application's response may optionally be provided without any live communication with the server, or at least without having to wait for a response from the server, thereby further reducing the time over which a response can be provided.

[0006] As described in more detail herein, one or more (e.g., all) proactive cache entries stored in a client device's local proactive cache are "proactive" in the sense that the entries are not stored in the proactive cache in response to a recent request at the client device based on user input. Rather, proactive cache entries may be proactively prefetched from a remote system, such as a server, and then stored in the proactive cache even though user input that conforms to the proactive cache's assistant request parameters has never been received.

[0007] Different proactive cache entries may be provided to different client devices, and various implementations disclosed herein relate to techniques utilized by a remote system when determining which proactive cache entries to provide to which client devices. In some of these implementations, when determining which proactive cache entries to provide to a given client device (proactively or on demand), the remote system selects a subset of cache entries to provide to the given client device from a superset of candidate proactive cache entries. Various considerations may be utilized in selecting the subset, including considerations that take into account attributes of the given client device and / or one or more attributes of the proactive cache entries. Attributes of / for a client device may include any data related to the operation of the client device, such as, for example, data related to the operating system version level, the model of the client device, which applications are installed on the client device (and which versions), the search history of a user of the client device, the location history of the client device, the power mode or battery level of the client device, etc.

[0008] As an example, when selecting one or more proactive cache entries of the subset, which applications are installed on a given client device (and, optionally, which of them are implicitly or explicitly flagged as preferred) may be considered and compared with the action content (and / or metadata) of the candidate proactive cache entries. For example, multiple candidate proactive cache entries may be generated, each with the same Assistant request parameters but different action content. Each different action content may be tailored to a specific application, e.g., a corresponding deep link that is locally executable by the Assistant client application to cause a corresponding additional application to open in a first state to perform the given action. For example, assume Assistant request parameters of "adjust thermostat schedule," "change thermostat schedule," and / or "{intent=change / adjust; device=thermostat; property=schedule}." For such Assistant request parameters, various different applications may be utilized to implement an appropriate response, such as a first application for a first thermostat manufacturer and a second application for a second thermostat manufacturer. The first and second proactive cache entries may both have the same Assistant request parameters but include different deep links. For example, the first proactive cache entry may have a first deep link to the first application (which, when executed, causes the first application to open in a corresponding schedule change interface), and the second proactive cache entry may have a separate second deep link to the second application (which, when executed, causes the second application to open in a corresponding schedule change interface).

[0009] In such an example, a first proactive cache entry (rather than a second) may be selected for provisioning to a given client device based on determining that the given client device has the first application installed but not the second application. This allows the given client device to respond quickly and efficiently to typed or spoken utterances that conform to the assistant request parameters because the given client device can select only the first deep link of the action content of that entry without having to locally determine which of multiple disparate deep links should be utilized. Furthermore, the action content of the first and second proactive cache entries is reduced in storage space compared to action content that includes multiple disparate deep links. This allows each of them to be individually stored more efficiently in the corresponding local proactive cache and consume less bandwidth when transmitted to the corresponding client device. Thus, fewer computer and network resources may be consumed. Furthermore, the use of deep links in the proactive cache may facilitate the implementation of actions with less user input and fewer processing steps, which may reduce resources consumed by the client device when implementing the actions. The requirement for less user input may also be beneficial for users with reduced dexterity and may improve the utility of the device.

[0010] As another example, when selecting one or more proactive cache entries of the subset, optionally, depending on the magnitude of the event, a proactive cache entry for an entity having one or more determined events may be more likely to be selected. Furthermore, a proactive cache entry for the entity may even be generated in response to a determined occurrence of an event. Some examples of determining an event for an entity are determining an increase in demand for the entity, determining an increase in Internet content for the entity, and / or predicting an increase in demand for the entity.

[0011] For example, if it is determined that the volume of assistant requests (and / or legacy search requests) for "Jane Doe" will spike, along with the spike in Internet content related to "Jane Doe," proactive cache entries for Jane Doe may be more likely (more likely than before the spike) to be generated and / or provided to various client devices for local proactive caching. Providing proactive cache entries for "Jane Doe" to various client devices may optionally be further based on determining that those various client devices have corresponding attributes related to Jane Doe (e.g., past searches for Jane Doe, and / or related entities (e.g., other entities of the same type), geographic locations associated with Jane Doe, and / or other attributes).

[0012] In these and other ways, proactive cache entries for an entity may be provided in anticipation of successive requests for the entity (and the associated bandwidth consumption and server-side processor consumption for processing the requests) to enable the eventual request to be responded to more quickly. Furthermore, provisioning of proactive cache entries for many client devices may occur during periods of relatively low network usage (e.g., overnight when those client devices are idle and charging), allowing for use of network resources during low-usage periods while mitigating network resource usage during periods of higher usage. In other words, a proactive cache entry for Jane Doe may be pre-stored at the client device during low-usage periods and then utilized to locally respond to typed or spoken utterances at the client device during periods of high usage. When this occurs across a large number of client devices, as contemplated herein, this may enable effective redistribution of network resources. Thus, network performance may be improved. Furthermore, various implementations select only a subset of client devices for provisioning proactive cache entries for “Jane Doe” and provide the subset based on the client devices of the subset having attributes related to Jane Doe. By selecting and providing only a subset of proactive cache entries, a reduction in bandwidth usage and network resources may be achieved while still selectively utilizing network resources to provision proactive cache entries to client devices that are likely to utilize such entries, resulting in faster provisioning of responses at and reduced interaction duration with those client devices.

[0013] As another example of determining and utilizing events for an entity, an increase in demand for a Hypothetical Artist (in this example, a virtual musical artist) may be predicted even though an increase in demand has not yet been observed. For example, the increase in demand may be predicted based on determining from one or more additional systems that the Hypothetical Artist is scheduled to release a new song and / or a new album. In response to the prediction of an increase in demand, a proactive cache entry for the Hypothetical Artist may be generated and / or may be more likely (more likely than before the spike) to be provided to various client devices for local proactive caching. For example, a proactive cache entry having assistant request parameters of "play Hypothetical Artist," "listen to a Hypothetical Artist," and / or "{intent=listen to music; artist= Hypothetical Artist}." The proactive cache entry may further include assistant action content, such as, for example, one or more deep links that, when executed, each cause a corresponding application to open with music from the Hypothetical Artist streamed and audibly rendered on the client device. Providing proactive cache entries for "Hypothetical Artist" to various client devices may optionally be further based on determining that those various client devices have corresponding attributes related to the Hypothetical Artist (e.g., past streaming of music from the Hypothetical Artist and / or related musical artists, music files from the Hypothetical Artist stored on the device, search history related to the Hypothetical Artist, etc.).Providing a proactive cache entry to a given client device may be based on determining that the given client device has an application that corresponds to one of the deep links of the proactive cache entry (e.g., in some implementations, to the only deep link).

[0014] In addition to local proactive caches, each stored locally at a corresponding client device, some implementations may further include remote proactive caches, each generated for a subset of client devices. The subset of client devices for a remote proactive cache may be a single client device or a group of client devices grouped based on those client devices having attributes in common with each other. For example, a remote proactive cache may be for only a single client device, or it may be for 1,000 client devices having the same or similar attributes.

[0015] A remote proactive cache for a given client device (either for that given client device only, or for that given client device and other client devices with the same / similar attributes) includes (or is limited to) proactive cache entries that are additional to the proactive cache entries stored in the given client device's local proactive cache. The proactive cache entries of a remote proactive cache, in many implementations, are still a subset of all available candidate cache entries. The proactive cache entries of a remote proactive cache may include proactive cache entries determined to be relevant to a given client device based on attributes of the given client device and / or attributes of the proactive cache entries (e.g., based on a comparison of client device attributes and proactive cache entry attributes, and / or based on proactive cache entry attributes). However, the proactive cache entries of a remote cache include proactive cache entries that are not provided for local storage in the local proactive cache. The decision not to provide them for local storage in the local proactive cache may be based on storage space limitations for the local proactive cache (e.g., providing them would cause the storage limitations for the local proactive cache to be exceeded) and may also be based on determining that the entries are not highly related to those provided for storage in the local proactive cache of a given client device.In other words, the remote proactive cache for a given client device may include proactive cache entries that are determined to be relevant to the given client device but not stored in the given client device's local cache based on storage space limitations and based on the proactive cache entries being determined to be less relevant to those provided for local storage in the given client device's local proactive cache.

[0016] The remote proactive cache may be utilized by the remote automated assistant component when responding to an assistance request from a client device assigned to the remote proactive cache. For example, in the case of spoken speech detected in audio data at a given client device, the given client device may transmit the audio data and / or locally determined recognized text to the remote automated assistant component. The transmission to the remote automated assistant component may optionally be responsive to a determination that no local proactive cache entry is responsive, or may occur in parallel with a determination of whether a local proactive cache entry is responsive. The transmission may be accompanied by an identifier of the given client device and an identifier utilized to identify the remote proactive cache for the given client device. A remote fulfillment module of the automated assistant component may determine whether a remote proactive cache entry is responsive to the assistance request. If responsive, the remote fulfillment module may utilize the response entry to determine assistant action content for responding to the assistance request. The assistant action content may be executed remotely by the remote automated assistant component and / or transmitted to the given client device for local execution.

[0017] Utilization of a remote proactive cache by a remote automated assistant component enables the remote automated assistant component to respond more quickly to automated assistant requests (e.g., typed or spoken utterances directed to the automated assistant requesting performance of a given action). This may be the result of an assistant action being directly mapped to assistant request parameters in a proactive cache entry, which may enable efficient identification of an assistant action from a proactive cache entry without having to generate the assistant action live in response to the assistant request. For example, without an assistant action in a proactive cache entry, the assistant action would have to be generated on the fly, optionally through communication with one or more additional remote systems, which may be computationally burdensome and introduce latency as a result of communicating with the additional remote systems. Thus, although a client-server round trip is required, utilization of a remote proactive cache in analyzing an assistant request still provides reduced latency provisioning of a response, resulting in a reduced duration of the user-assistant interaction. Furthermore, in some implementations, a remote proactive cache specific to a subset of client devices may be stored in one or more servers that are geographically proximate to the subset of client devices, such as a server that receives assistance requests from a given geographic region. This may further reduce latency in analyzing assistance requests, as assistance actions may be determined more quickly from the proactive cache without requiring inter-server communication between multiple geographically distant servers.

[0018] Various implementations disclosed herein are directed to client devices (e.g., smartphones and / or other client devices) that include at least one or more microphones and an automated assistant application. The automated assistant application may be installed “on top of” the client device's operating system and / or may itself form part of (or the entirety of) the client device's operating system. The automated assistant application includes and / or has access to on-device speech recognition, on-device natural language understanding, and on-device fulfillment. For example, on-device speech recognition may be implemented using an on-device speech recognition module that processes audio data (detected by the microphone) using an end-to-end speech recognition machine learning model stored locally on the client device. The on-device speech recognition generates recognized text for spoken utterances (if any) present in the audio data. Also, for example, on-device natural language understanding (NLU) may be implemented using an on-device NLU module that processes the recognized text generated using on-device speech recognition and, optionally, contextual data, to generate NLU data. The NLU data may include an intent that corresponds to the verbal utterance and, optionally, parameters (e.g., slot values) for that intent.

[0019] On-device fulfillment may be implemented using an on-device fulfillment module that utilizes recognized text (from on-device speech recognition) and / or NLU data (from on-device NLU), and optionally other local data, to determine actions to take to analyze the intent of the verbal utterance (and optionally parameters for that intent). This may include determining local and / or remote responses (e.g., answers) to the verbal utterance, interactions with locally installed applications to perform based on the verbal utterance, commands to send to Internet of Things (IoT) devices (directly or via corresponding remote systems) based on the verbal utterance, and / or other analytical actions to perform based on the verbal utterance. The on-device fulfillment may then initiate local and / or remote implementation / execution of the actions determined to analyze the verbal utterance. As described herein, in various implementations, on-device fulfillment utilizes a locally stored proactive cache in response to various user inputs. For example, the on-device fulfillment may utilize action content of a proactive cache entry of a locally stored proactive cache in response to a verbal utterance based on determining that the recognized text and / or NLU data matches the assistant request parameters of the proactive cache entry.

[0020] In various implementations, remote speech processing, remote NLU, and / or remote fulfillment may be utilized at least selectively. For example, recognized text may be at least selectively sent to a remote automated assistant component for remote NLU and / or remote fulfillment. For example, recognized text may optionally be sent for remote fulfillment in parallel with on-device fulfillment or in response to failure of on-device NLU and / or on-device fulfillment. However, on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device execution may be prioritized due to the reduced latency they provide at least when analyzing spoken speech (by not requiring a client-server round trip to analyze the spoken speech). Furthermore, on-device functionality may be the only functionality available in situations where network connectivity is absent or limited.

[0021] In various implementations, on-device speech recognition and / or other on-device processes are activated in response to detecting the occurrence of an explicit Assistant activation cue and / or in response to any occurrence of an implicit activation cue. An explicit activation cue is a cue that, when detected separately, will always cause at least on-device speech recognition to be activated. Some non-limiting examples of explicit activation cues include detecting a spoken hotword with at least a threshold degree of reliability, an activation of an explicit Assistant interface element (e.g., a hardware button or a graphical button on a touchscreen display), a "phone squeeze" of at least a threshold intensity (e.g., as detected by a sensor in a mobile phone bezel), and / or other explicit activation cues. However, other cues are implicit in that on-device speech recognition will only be activated in response to any occurrence of those cues, such as their occurrence in certain content (e.g., following or combined with other implicit cues). For example, on-device voice recognition may optionally not be activated in response to independent voice activity detection, but may be activated in response to detecting a user's presence at the client device and / or detecting voice activity together with detecting a user's presence at the client device within a threshold distance. Also, sensor data from non-microphone sensors, such as a gyro and / or accelerometer, indicating that a user has picked up and / or is currently holding the client device, may optionally not independently activate on-device voice recognition. However, on-device voice recognition may be activated in response to such indications along with the detection of voice activity and / or directed speech (described in more detail herein) in hotword-free audio data. Hotword-free audio data is audio data that lacks any verbal utterances that include an explicit Assistant activation cue, a "hotword."As yet another example, a "phone squeeze" below a threshold intensity may optionally be insufficient to activate on-device voice recognition independently, however, on-device voice recognition may be activated in response to such a low intensity "phone squeeze" in conjunction with the detection of directional speech within voice activity and / or hotword-free audio data.

[0022] Some implementations disclosed herein include one or more computing devices including one or more processors, such as a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU). One or more of the processors are operable to execute instructions stored in associated memory, which instructions are configured to cause any of the methods described herein to be performed. The computing device may include, for example, a client assistant device with a microphone, a display, and / or other sensor components. Some implementations also include one or more non-transitory computer-readable storage media having stored thereon computer instructions executable by the one or more processors to perform any of the methods described herein. [Brief explanation of the drawings]

[0023] [Figure 1] FIG. 1 is a block diagram of an example environment in which implementations disclosed herein may be implemented. [Figure 2] 2A-2C illustrate exemplary process flows illustrating how the various components of FIG. 1 may interact, according to various implementations. [Figure 3] FIG. 2 illustrates some examples of proactive cache entries. [Figure 4] 1 is a flowchart illustrating an exemplary method for prefetching and storing proactive cache entries according to implementations disclosed herein. [Figure 5A]1 is a flowchart illustrating an example method for generating proactive cache entries, provisioning a local subset of proactive cache entries to a given client device, and / or determining a remote subset of proactive cache entries. [Figure 5B] 5B is a flowchart illustrating some implementations of block 510 of FIG. 5A. [Figure 5C] 5B is a flowchart illustrating some additional or alternative implementations of block 510 of FIG. 5A. [Figure 6] FIG. 1 illustrates an exemplary architecture of a computing device. DETAILED DESCRIPTION OF THE INVENTION

[0024] 1 , a client device 160 is shown that at least selectively executes an automated assistant client 170. The term "assistant device" is used herein to refer to a client device 160 that at least selectively executes an automated assistant client 170. Automated assistant client 170, in the example of FIG. 1 , includes an audio capture engine 171, a visual capture engine 172, an on-device speech recognition engine 173, an on-device NLU engine 174, an on-device fulfillment engine 175, an on-device execution engine 176, and a prefetch engine 177.

[0025] One or more remote automated assistant components 180 may optionally be implemented on one or more computing systems communicatively coupled to client device 160 via one or more local and / or wide area networks (e.g., the Internet), shown generally at 190. Remote automated assistant component 180 may be implemented, for example, via a cluster of high performance servers.

[0026] In various implementations, an instance of automated assistant client 170, through its interactions with one or more cloud-based automated assistant components 180, can form what appears from the user's perspective to be a logical instance of automated assistant 195 with which the user can engage in human-to-computer interactions (e.g., verbal interactions, gesture-based interactions, and / or touch-based interactions).

[0027] Client device 160 may be, for example, a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device in the user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker, a smart appliance such as a smart television (or a standard television equipped with a network-connected dongle with automated assistant functionality), and / or a user's wearable equipment including a computing device (e.g., a user's watch with a computing device, a user's glasses with a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices may be provided.

[0028] Client device 160 may optionally be equipped with one or more vision components 163 having one or more fields of view. Vision component 163 may take various forms, such as a monographic camera, a stereo camera, a LIDAR component (or other laser-based component), a radar component, etc. One or more vision components 163 may be used, for example, by vision capture engine 172 to capture visual frames (e.g., image frames, laser-based frames) of the environment in which client device 160 is deployed. In some implementations, such visual frames may be utilized to determine whether a user is present near client device 160 and / or the distance of the user (e.g., the user's face) relative to the client device. Such a determination may be utilized by automated assistant client 170 in determining whether to activate on-device speech recognition engine 173, on-device NLU engine 174, on-device fulfillment engine 175, and / or on-device execution engine 176. The vision engine may additionally or alternatively be utilized to locally detect various user touch-free gestures (e.g., "thumbs up," "hand wave," etc.). Optionally, the detected gesture may be the "assistant request parameter" of a proactive cache entry described herein, and the corresponding action to be taken based on the gesture may be the "action content" of the proactive cache entry.

[0029] Client device 160 may also be equipped with one or more microphones 165. Voice capture engine 171 may be configured to capture the user's voice and / or other audio data captured via microphone 165. As described herein, such audio data may be utilized by on-device voice recognition engine 173.

[0030] Client device 160 may also include one or more presence sensors 167 and / or one or more displays 169 (e.g., touch-sensitive displays). Display 169 may optionally be utilized to render streaming text transcription from on-device speech recognition engine 173. Display 169 may further be one of the user interface output components through which the visual portion of a response is rendered from automated assistant client 170. Presence sensor 167 may include, for example, a PIR and / or other passive presence sensor. In various implementations, one or more components and / or functions of automated assistant client 170 may be initiated in response to the detection of a human presence based on the output from presence sensor 167. In implementations in which the initiation of components and / or functions of automated assistant client 170 is conditional on the first detection of the presence of one or more users, power resources may be conserved.

[0031] The automated assistant client 170 activates the on-device speech recognition engine 173 in response to detecting the occurrence of an explicit activation cue and / or detecting the occurrence of an implicit activation cue. When activated, the on-device speech recognition engine 173 processes the audio data captured by the microphone 165 using an on-device speech recognition model (not shown in FIG. 1 for simplicity) to determine recognized text within the spoken utterance (if any) captured by the audio data. The on-device speech recognition model may optionally be an end-to-end model and may optionally be supplemented by one or more techniques that explore generating additional recognized text hypotheses and selecting the best hypothesis using various considerations. The processed audio data may include audio data captured after on-device speech recognition is activated and, optionally, recent audio data that has been locally buffered (e.g., buffered prior to activation of the on-device speech recognition engine 173). The audio data processed by the on-device speech recognition engine 173 may include raw audio data and / or a representation thereof. Audio data may be provided to the on-device speech recognition engine 173 in a streaming manner as new audio data is detected.

[0032] In some implementations, when the on-device speech recognition engine 173 is activated, a human-perceptible cue is rendered to inform the user that such activation has occurred and / or to render a stream of recognized text as recognition occurs. The visual rendering may also include a selectable “delete” element that, when selected via touch input on a touchscreen display, stops the on-device speech recognition engine 173. As used herein, activating the speech recognition engine 173 or other component means causing that component to perform at least more processing than it was previously performing prior to activation. It may mean activating a component from a completely dormant state.

[0033] On-device NLU engine 174, on-device fulfillment engine 175, and / or on-device execution engine 176 may optionally be activated in response to detecting the occurrence of explicit and / or implicit activation cues. Alternatively, one or more of these engines may be activated only based on an initial analysis of recognized text from on-device speech recognition engine 173 that indicates the recognized text may be an assistant request.

[0034] When the on-device NLU engine 174 is activated, it performs on-device natural language understanding on the recognized text generated by the on-device speech recognition engine 173 to generate NLU data. The NLU engine 174 may optionally utilize one or more on-device NLU models (not shown in FIG. 1 for simplicity) in generating the NLU data. The NLU data may include, for example, an intent corresponding to a verbal utterance and, optionally, parameters (e.g., slot values) for that intent.

[0035] Additionally, when on-device fulfillment engine 145 is activated, on-device fulfillment engine 145 generates fulfillment data using recognized text (from on-device speech recognition engine 173) and / or NLU data (from on-device NLU engine 174). On-device fulfillment engine 145 may access proactive cache 178 stored locally at client device 160 when determining whether the recognized text and / or NLU data matches the assistant request parameters of a proactive cache entry in the proactive cache. If there is a match, on-device fulfillment engine 145 may utilize action content of the matching proactive cache entry as all or part of the generated fulfillment data. The action content may include an action to take based on the request and / or data related to the fulfillment of such request. If no match is found, on-device fulfillment engine 145 can optionally utilize other on-device fulfillment models (if any) when attempting to generate fulfillment data, or can wait for remote fulfillment data from remote automated assistant component 180 (e.g., when automated assistant client 170 is online and provides recognized text and / or other data to remote automated assistant component 180 for generating fulfillment data). The combination of on-device speech recognition and local proactive caching can reduce the need to transmit data to and from a server, thus reducing bandwidth and network resource usage. Furthermore, low-latency responses can be provided to users even in areas with poor or no network connectivity.

[0036] When determining whether the recognized text or NLU data matches the Assistant request parameters of a proactive cache entry in the proactive cache, the on-device fulfillment engine 145 may utilize exact matches and / or soft matches. For example, when determining whether the recognized text matches the Assistant request parameters, the on-device fulfillment engine 145 may require an exact match of the recognized text to the text of the Assistant request parameters, or may allow only minimal differences (e.g., the inclusion or exclusion of certain stop words). Also, for example, when determining whether the NLU data matches the Assistant request parameters, the on-device fulfillment engine 145 may require an exact match of the NLU data to the NLU data of the Assistant request parameters. Although various examples are described herein with respect to recognized text determined by on-device speech recognition, it should be understood that typed text (e.g., typed using a virtual keyboard) and / or NLU data based on the typed text may also be provided and responded to in a similar manner.

[0037] When the fulfillment data is generated by the on-device fulfillment engine 175, the fulfillment data may then be provided to the on-device execution engine 176 for on-device execution based on the fulfillment data (e.g., on-device implementation of an action based on the action content of the proactive cache entry). On-device execution based on the action content of the fulfillment data may include, for example, executing deep links included within the action content, rendering text, graphics, and / or audio included within the action content (or audio data converted from a response using an on-device speech synthesizer), and / or sending commands included within the action content to a peripheral device (e.g., via WiFi and / or Bluetooth) to control the peripheral device.

[0038] Prefetch engine 177 prefetches proactive cache entries from proactive cache system 120 for inclusion in proactive cache 178. One or more proactive cache entries prefetched and stored in proactive cache 178 may be proactive cache entries that have assistant request parameters that do not comply with any user input previously provided to automated assistant client 170 and / or have action content that has never been utilized by automated assistant client 170. Thus, such cache entries may reduce the amount of time to provide assistant responses to various user inputs that have never been received by automated assistant client 170.

[0039] In some implementations and / or situations, proactive cache system 120 may optionally push proactive cache entries to prefetch engine 177. However, in other implementations and / or situations, proactive cache system 120 sends proactive cache entries in response to a prefetch request from prefetch engine 177. In some of those implementations, prefetch engine 177 sends the prefetch request in response to a determination that one or more conditions are met. The conditions may include, for example, one or more of: certain network conditions exist (e.g., a connection to a Wi-Fi network and / or a connection to a Wi-Fi network with certain bandwidth conditions); client device 160 is charging and / or has at least a threshold battery charge state; the client device is not being actively utilized by a user (e.g., based on on-device accelerometer and / or gyroscope data); the client device's current processor usage and / or current memory usage does not exceed certain thresholds; etc. Thus, a proactive cache entry can be retrieved during periods when certain ideal conditions exist, but can be utilized later under any conditions, including when certain ideal conditions do not exist.

[0040] A prefetch request from prefetch engine 177 may include an identifier of client device 160 and / or an account associated with the client device, and / or may include another indication of proactive cache entries already stored in proactive cache 178. Proactive cache system 120 may utilize the indication of proactive cache entries already stored in proactive cache 178 to provide only proactive cache entries not already stored in proactive cache 178 in response to a prefetch request. This may conserve network resources by only sending new / updated proactive cache entries to add to proactive cache 178 instead of sending new / updated proactive cache entries along with existing cache entries to completely replace proactive cache 178. As an example, proactive cache system 120 may maintain a listing of active proactive cache entries in proactive cache 178 against the identifier of client device 160. Proactive cache system 120 may utilize such listing in determining which new / updated proactive cache entries to provide in response to a prefetch request that includes the identifier. As another example, a prefetch request may include a token or other identifier that maps to a set of proactive cache entries stored in proactive cache 178, and proactive cache system 120 may utilize such token in determining which new / updated proactive cache entries are not within the mapped set and should be provided in response to a prefetch request that includes the token.

[0041] Prefetch engine 177 may also selectively remove proactive cache entries from proactive cache 178. For example, a proactive cache entry may include a time to live (TTL) value as part of its metadata. The TTL value may define a duration or threshold time period after which a proactive cache entry may be considered stale and, as a result, unused by on-device fulfillment engine 175 and / or removed from proactive cache 178 by prefetch engine 177. For example, if a proactive cache entry's TTL value indicates that it should live for seven days and the proactive cache entry's timestamp indicates that it was received eight days ago, prefetch engine 177 may remove the proactive cache entry from proactive cache 178. This can free up limited storage resources of client device 160 and create space in proactive cache 178 for other temporally proactive cache entries whose overall size may be constrained by the limited storage resources of client device 160.

[0042] In some implementations, prefetch engine 177 may additionally or alternatively remove proactive cache entries from proactive cache 178, with one or more of the removed proactive cache entries not being indicated as stale based on their TTL values, yet still creating space for new proactive cache entries. For example, proactive cache 178 may have a maximum size, which may be, for example, a user set and / or may be determined based on the storage capacity of client device 160. If a new proactive cache entry from a prefetch request exceeds the maximum size, prefetch engine 177 may remove one or more existing cache entries to create space for the new proactive cache entry. In some implementations, proactive cache entries may be removed based on the cache entry's timestamp, e.g., the oldest proactive cache entry may be removed first by prefetch engine 177. In some implementations, metadata about existing proactive cache entries may include a score, ranking, or other priority data (also more generally referred to herein as ranking criteria), and those with the lowest priority may be removed by prefetch engine 177. Additionally or alternatively, prefetch engine 177 may optionally bias against removal of proactive cache entries that have never been used, also considering their timestamps to bias against proactive cache entries that have never been used and have been in the proactive cache for at least a threshold duration. Additionally or alternatively, proactive cache system 120 may optionally provide an indication of which existing proactive cache entries should be removed in response to a prefetch request.

[0043] Proactive cache system 120 generates proactive cache entries and serves prefetch requests from client device 160 and other client devices with proactive cache entries selected for the requesting client device. Proactive cache system 120 may also generate remote proactive caches 184 that are utilized by remote automated assistant component 180 and each of which may be specific to one or more client devices.

[0044] Proactive cache system 120 may include cache entry generation engine 130, cache assembly engine 140, and entity event engine 150. Generally, cache entry generation engine 130 utilizes various techniques described herein to generate a large number of candidate proactive cache entries (referred to as cache candidates 134 in FIGS. 1 and 2 ). Cache assembly engine 140 determines, for each of a plurality of client devices, a corresponding subset of cache candidates 134 to be provided to the corresponding client device for storage in the corresponding client device's local proactive cache. Cache assembly engine 140 may also optionally generate remote proactive caches 184, each associated with a subset of client devices, each containing a corresponding subset of cache candidates, and each utilized by remote automated assistant component 180 in fulfilling for the corresponding client device. Entity event engine 150 may optionally determine the occurrence of events related to various entities through interactions with one or more remote systems 151. In some implementations, the entity event engine 150 may provide information about those events to the cache entry generation engine 130 to cause the cache entry generation engine 130 to generate one or more corresponding cache candidates 134 for the entity. In some implementations, the entity event engine 150 may additionally or alternatively provide information about those event cache assembly engines 140, which may use that information in determining whether to provide various cache candidates to corresponding client devices and / or for inclusion in the remote proactive cache 184.

[0045] In FIG. 1 , cache entry generation engine 130 includes request parameter module 131, action content module 132, and metadata module 133. Request parameter module 131 generates assistant request parameters for each proactive cache entry. The assistant request parameters of a proactive cache entry represent one or more assistant requests to perform a given action. An assistant request can be a typed or spoken utterance requesting the performance of a given action. For example, multiple assistant requests can each be a request to render a local forecast for the current day, such as the following typed or spoken utterances: "Today's forecast," "Today's local forecast," "What's the weather like today," and "What's the weather like?" Request parameter module 131 seeks to generate a textual and / or NLU representation that captures the multiple assistant requests to perform the same given action. For example, the text of each of the previous utterances can be included as an assistant request parameter and / or an NLU expression common to all of the utterances, such as a structured representation of "{intent=weather; location=local; date=today}."

[0046] The action content module 132 generates action content for each proactive cache entry. The action content may vary depending on the proactive cache entry. The action content for a proactive cache entry may include, for example, a deep link to be executed, text, graphics, and / or audio to be rendered, and / or commands to be sent to a peripheral device.

[0047] Continuing with the example of the current day's local forecast, the action content module 132 may generate different action content for each of a plurality of cache entries, the action content for each being tailored for a different geographic region. For example, a first proactive cache entry may be for a first city and may include first action content specifying a daily forecast for the first city, including assistant request parameters and text, graphics, and / or audio to be rendered. A second proactive cache entry may be for a second city and may include second action content including the same (or similar) assistant request parameters but specifying a different daily forecast for the second city, including text, graphics, and / or audio to be rendered. As described with respect to the cache assembly engine 140, these two different cache entries may be provided to different client devices and / or remote proactive caches 184 based on attributes of the client devices. For example, a first client device in a first city may be provided with the first proactive cache entry but not the second proactive cache entry.

[0048] In some implementations, the action content module 132 may generate different action content for each of multiple cache entries, with the action content for each being tailored to one or more different applications. For example, a first proactive cache entry may include assistant request parameters of “play Hypothetical Artist,” “listen to a Hypothetical Artist,” and / or “{intent=listen to music; artist= Hypothetical Artist}.” A second proactive cache entry may include the same (or similar) assistant request parameters. Nevertheless, the action content module 132 may generate first action content for the first proactive cache entry that includes a deep link that, when executed, causes a first music application to open with music by the Hypothetical Artist beginning to stream. The action content module 132 may generate second action content for the second proactive cache entry that includes a different deep link that, when executed, causes a second music application to open with music by the Hypothetical Artist beginning to stream. As described with respect to cache assembly engine 140, these two disparate cache entries may be provided to different client devices and / or remote proactive cache 184 based on attributes of the client devices, i.e., based on which applications the client devices have installed and / or indicated as preferred applications for music streaming. For example, a first client device that has a first application as its only music streaming application may be provided with the first proactive cache entry but not the second proactive cache entry.

[0049] For some Assistant request parameters, there may be only a single proactive cache entry. For example, for Assistant request parameters for an Assistant request for an image of a Cavalier King Charles Spaniel (a type of dog), a single proactive cache entry may be provided that includes action content with an image of a Cavalier King Charles Spaniel.

[0050] The metadata module 133 optionally generates metadata about the cache entries. Some of the metadata may optionally be used by the proactive cache system 120 without needing to be transmitted to the client device. For example, the metadata module 133 may generate metadata about the proactive cache entry that indicates one or more entities associated with the proactive cache entry, the language of the action content for the proactive cache entry, and / or other data about the proactive cache entry. The cache assembly engine 140 may use such metadata when determining which client devices the proactive cache entry should be provided to. For example, for a proactive cache entry that includes as its action content a local weather forecast for a first city, the metadata module 133 may generate metadata that indicates the first city. The cache assembly engine 140 may use such metadata when selecting proactive cache entries for inclusion in local or remote proactive caches only for client devices that have the first city as their current or preferred location. Also, for example, for a proactive cache entry that includes as its action content graphics and / or text about an actress, the metadata module 133 may generate metadata that indicates the actress. The cache assembly engine 140 may utilize such metadata in selecting proactive cache entries for inclusion in a local or remote proactive cache only for client devices that have an attribute corresponding to that actress, e.g., client devices that have an attribute indicating that actress based on prior viewing of content related to the celebrity and / or have an attribute indicating movies / shows in which that actress starred, based on streaming movies or television shows that include that actress.

[0051] The metadata module 133 may also generate metadata to be sent with the proactive cache entry to the client device and / or utilized in maintaining the proactive cache in a remote proactive cache. For example, the metadata module 133 may generate a timestamp for the proactive cache entry indicating when it was generated and / or last validated (e.g., verifying the accuracy of the action content). The metadata module 133 may also generate a TTL value for the proactive cache entry. The TTL value for a given proactive cache entry may be generated based on various considerations, such as assistant request parameters and / or characteristics of the action content. For example, some action content, such as weather-related action content, is dynamic, and a proactive cache entry with such action content may have a relatively short TTL (e.g., 6 hours, 12 hours). On the other hand, some action content is static, and a proactive cache entry with such action content may have a relatively long TTL (e.g., 7 days, 14 days, 30 days). As another example, a proactive cache entry that includes static content but is provided and / or generated based on event detection by entity event engine 150 may have a shorter TTL than a proactive cache entry that includes static content but is provided and / or generated independently of event detection by entity event engine 150.

[0052] 1 , cache assembly engine 140 includes local module 141 and remote module 142. Local module 141 selects, for each client device, from cache candidates 134, a corresponding subset of cache candidates 134 to provide to the client device. Local module 141 may determine the subset to provide to a given client device based on various considerations, such as a comparison of attributes of the client device against attributes of the proactive cache entries and / or ranking criteria for the proactive cache entries.

[0053] For example, when selecting a subset for a given client device, the local module 141 may filter out any proactive cache entries whose metadata indicates that the action content is applicable only to corresponding applications not installed on the given client device (e.g., deep links related only to that application), that the action content is applicable only to geographic regions not associated with the given client device, that the action content is only in languages ​​that are not set as primary (and optionally secondary) languages ​​for the given client device. The remaining proactive cache entries may be selected based on a comparison of their attributes against the attributes of the given client device, ranking criteria (which may also be considered attributes) for the remaining proactive cache entries, and / or other considerations.

[0054] For example, the local module 141 may be more likely to select a proactive cache entry having metadata indicating one or more entities that correspond to one or more entities determined to have been previously interacted with by a given client device compared to a proactive cache entry having metadata indicating an alternative entity that fails to accommodate the one or more entities determined to have been previously interacted with by a given client device. Also, for example, the ranking criteria for a given proactive cache entry may indicate how frequently their corresponding assistant requests are submitted via the assistant interface (either overall or for a given client device) and / or how frequently their corresponding action content is rendered (either overall or for a given client device). Rendering the action content may include causing any associated textual / graphical / audible content of the action content to be rendered on the client device to perform the action. Executing the deep link may include automatically performing the associated action (e.g., opening an application to a particular state) or may include preparing the client device to perform the associated action, where the performance may be in response to, for example, a user interface input. The ranking criteria may be based on recent event detection for a given proactive cache entry by entity event engine 150. For example, assume that a given proactive cache contains action content that includes assistant request parameters and installation instructions for installing a new peripheral device (e.g., a new smart thermostat).If the entity event engine 150 determines a significant increase in requests (assistant requests, search engine requests, or other requests) for the new peripheral device and / or a significant increase in content (e.g., web pages, social media comments) for the new peripheral device, the corresponding ranking criteria may indicate a higher ranking, making the proactive cache entry more likely to be selected.

[0055] When determining a proactive cache entry for a given client device, the local module 141 may also consider the storage space allocated to the proactive cache for the given client device. Additionally, when determining whether to provide a new proactive cache entry for a given client device when the given client device already includes an existing proactive cache entry, the existing proactive cache entry may be considered. For example, ranking criteria for existing proactive cache entries may be considered, and / or the unoccupied storage space (if any) of the given client device's existing proactive cache may be considered.

[0056] The remote module 142 optionally generates and maintains one or more remote proactive caches 184. Each remote proactive cache 184 is for a subset of client devices. The subset of remote proactive cache 184 can be a single client device or a collection of client devices that share one or more (e.g., all) attributes in common. For example, a collection of client devices for a remote proactive cache can be client devices that are in the same geographic region, have the same applications installed, and / or whose past interactions indicate at least a threshold amount of common interest in the same entities. For a remote proactive cache 184 for a client device, the remote module 142 may select one or more cache candidates 134 that have not been filtered out and have not already been offered for storage in the local proactive cache. For example, assume that for client device 160, proactive cache 178 has a 500 MB limit. Further assume that local module 141 has already selected and offered 500 MB worth of proactive cache entries for proactive cache 178. Remote module 142 may then select additional proactive cache entries using the same considerations as local cache module 141 for inclusion in remote proactive cache 184 for client device 160. For example, remote proactive cache 184 for client device 160 may have a 2GB limit, and remote module 142 may select 2GB worth of remaining proactive cache entries by comparing attributes of the proactive cache entries and client device 160 and / or considering ranking criteria for the proactive cache entries.

[0057] The entity event engine 150 interacts with one or more remote systems 151 when monitoring event occurrences related to various entities. Some examples of determining events for an entity are determining an increase in demand for the entity, determining an increase in Internet content related to the entity, and / or predicting an increase in demand for the entity. For example, the entity event engine 150 may determine whether the volume of assistance requests (and / or legacy search requests) has spiked for a particular router and / or whether there is a spike in Internet content for that particular router. In response, the entity event engine 150 can provide criteria related to the spike to the cache assembly engine 140, which may be more likely (than before the spike) to provide proactive cache entries related to that particular router to various client devices for local proactive caching and / or remote proactive caching. Providing proactive cache entries to various client devices may further be based on determining that the various client devices have corresponding attributes related to that router (e.g., past searches for that particular router or routers in general).

[0058] The entity event engine 150 may also provide an indication of the spike to the cache entry generation engine 130. In response, the cache entry generation engine 130 may optionally generate one or more proactive cache entries for a particular router. The cache entry generation engine 130 may optionally generate one or more proactive cache entries for a particular router based on determining that there are no current cache candidates 134 for the particular router and / or that there are less than a threshold amount of cache candidates 134 for the particular router. For example, the cache generation engine 130 may determine a class of a particular router (e.g., a general class of routers) and determine templates for frequent queries for entities of that class. For example, a template for "what is the maximum bandwidth for [router alias]" (based on related queries for other particular router aliases) or "what is the default IP address for [router alias]" (based on related queries for other particular router aliases). Assistance request parameters for the proactive cache entry may then be generated based on replacing "router alias" with the alias for the particular parameter. Additionally, action content for proactive cache entries may be generated based on snippets from the top search results for queries that replace "router alias" with the alias for a particular router, and / or using other techniques. For example, an assistant request for "what is the default IP address for a particular router" may generate action content for "192.168.1.1."

[0059] In some implementations, entity event engine 150 may determine that an event indicating action content in an existing cache entry is stale and provide instructions to cache entry generation engine 130 to cause cache entry generation engine 130 to generate a new cache entry to reflect the updated action content and / or remove the existing cache entry with the stale content. As used herein, generating a new proactive cache entry to reflect updated action content may include updating an existing proactive cache entry to reflect the new action content (and optionally, updated metadata) while retaining the assistant request parameters of the proactive cache entry. It may also include removing the existing proactive cache entry entirely and generating a new proactive cache entry with the same assistant request parameters but with updated action content (and optionally, updated metadata). As an example, entity event engine 150 may determine that the weather forecast for a geographic region has changed by at least a threshold amount, resulting in the generation of a new corresponding proactive cache entry.

[0060] In some implementations, the remote automated assistant component 180 may include a remote ARS engine 181 that performs speech recognition, a remote NLU engine 182 that performs natural language understanding, and / or a remote fulfillment engine 183 that generates fulfillment data, optionally utilizing a remote proactive cache 184 as described herein. Optionally, a remote execution module that performs remote execution based on fulfillment data determined locally or remotely may be included. Additionally and / or alternatively, a remote engine may be included. As described herein, in various implementations, on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device execution may be prioritized due to at least the latency and / or network usage reduction they offer when analyzing spoken and / or typed utterances (because a client-server round trip is not required to analyze the spoken utterance). However, one or more cloud-based automated assistant components 180 may be utilized at least selectively. For example, such components may be utilized in parallel with on-device components, with output from such components utilized when local components fail. For example, the on-device fulfillment engine 175 may fail in certain circumstances (e.g., when the size-constrained proactive cache 178 fails to contain a matching proactive cache entry), and the remote fulfillment engine 183 may utilize the more robust remote proactive cache 184 (or additional resources when the remote proactive cache does not have a match) to generate fulfillment data in such circumstances. The remote fulfillment engine 184 can operate in parallel with the on-device fulfillment engine 175, and its results can be utilized when on-device fulfillment fails or can be invoked in response to a determination of on-device fulfillment failure.

[0061] In various implementations, an NLU engine (on-device and / or remote) may generate annotated output that includes one or more annotations of the recognized text and one or more (e.g., all) of the terms of the natural language input. In some implementations, the NLU engine is configured to identify and annotate various types of grammatical information in the natural language input. For example, the NLU engine may include a morphological module that can separate individual words into morphemes and / or annotate the morphemes, for example, with their classes. The NLU engine may also include a portion of a phonetic tagger that is configured to annotate terms with their grammatical roles. Also, for example, in some implementations, the NLU engine may additionally and / or alternatively include a dependency parser configured to determine syntactic relationships between terms in the natural language input.

[0062] In some implementations, the NLU engine may additionally and / or alternatively include an entity tagger configured to annotate entity references in one or more segments, such as references to people (e.g., including literary figures, celebrities, public figures, etc.), organizations, locations (real and fictional), etc. In some implementations, the NLU engine may additionally and / or alternatively include a coreference resolver (not shown) configured to group or "cluster" references to the same entity based on one or more contextual cues. In some implementations, one or more components of the NLU engine may rely on annotations from one or more components of the NLU engine.

[0063] The NLU engine may also include an intent matcher configured to determine the intent of users involved in interactions with the automated assistant 195. The intent matcher may use various techniques to determine the user's intent. In some implementations, the intent matcher may have access to one or more local and / or remote data structures that include, for example, multiple mappings between grammars and response intents. For example, grammars included in the mappings may be selected and / or learned over time and represent common intents of users. In addition to or instead of grammars, in some implementations, the intent matcher may employ one or more trained machine learning models, alone or in combination with one or more grammars. These trained machine learning models may be trained to identify intents, for example, by embedding text recognized from spoken utterances into a reduced-dimensional space and then determining which other embeddings (and therefore intents) are closest, using techniques such as Euclidean distance, cosine similarity, etc. Some grammars have slots (e.g., <artist>) that can be filled with slot values. Slot values ​​may be determined in various ways. Users will often proactively provide slot values. For example, for the grammar "Order me a <topping> pizza," a user may speak the phrase "Order me a sausage pizza," in which case the slot <topping> is automatically filled. Other slot values ​​can be inferred based on, for example, user location, currently rendered content, user preferences, and / or other cues. Use of an intent manager as described herein, which may be implemented locally, may allow a proactive cache entry to be retrieved even if a user interface input to a client device is not a match (optionally, an exact match) to one or more assistant request parameters of the cache entry. This may improve the utility of the device.

[0064] Referring now to FIG. 2, an exemplary process flow is shown illustrating how the various components of FIG. 1 may interact according to various implementations.

[0065] 2 , prefetch engine 177 sends request 221 to proactive cache system 120. Proactive cache system 120 responds to the request with proactive cache entry 222. As described herein, proactive cache entry 222 may be a subset of cache candidates and may be selected for client device 160 based on attributes of client device 160, attributes of cache entry 222, ranking criteria for cache entry 222, and / or based on proactive cache entries already in proactive cache 178. Prefetch engine 177 optionally removes one or more existing proactive cache entries to make room for cache entry 222 and stores cache entry 222 in proactive cache 178.

[0066] At some point after storing cache entry 222 in proactive cache 178 (e.g., several minutes or hours), audio data 223 is detected via microphone 165 (FIG. 1) of client device 160 (FIG. 1). The detected audio data 223 is an example of a user interface input to the client device. An on-device speech recognition module processes audio data 223 to generate recognized text 171A.

[0067] Recognized text 171A may optionally be provided to on-device fulfillment engine 175 and / or to remote fulfillment engine 183. When recognized text 171A is provided to on-device fulfillment engine 175 and on-device fulfillment engine 175 determines that the text matches an assistant request parameter of a proactive cache entry in proactive cache 178, on-device fulfillment engine 175 may generate fulfillment data 175A that includes at least the action content of the matching proactive cache entry.

[0068] In addition to or instead of considering recognized text 171A, on-device fulfillment engine 175 may consider NLU data 174A generated by on-device NLU engine 174 based on processing of recognized text 171A (and optionally based on context data). When NLU data 174A is provided to on-device fulfillment engine 175 and on-device fulfillment engine 175 determines that the text matches assistant request parameters of a proactive cache entry in proactive cache 178, on-device fulfillment engine 175 may generate fulfillment data 175A that includes at least the action content of the matching proactive cache entry. Fulfillment data may be generated if the intent of the NLU data matches the assistant request parameters.

[0069] The on-device execution engine 176 may process the on-device fulfillment data 175A, including (or limited to) the action content of the matching proactive cache entry, and perform the corresponding action, which may include generating an audible, visual, and / or haptic response based on the action content, deep linking the action content, and / or transmitting (e.g., via Bluetooth or Wi-Fi) commands contained within the action content.

[0070] In some implementations, the recognized text 171A is provided to the remote fulfillment engine 183. The recognized text 171A may be provided to the remote fulfillment engine 183 in parallel with provisioning to the on-device fulfillment engine 175, or optionally, solely in response to a determination by the on-device fulfillment engine 175 (based on the recognized text 171A and / or NLU data 174A) that no matching entry exists in the proactive cache 178. The remote fulfillment engine 183 may access the remote proactive cache 184A assigned to the client device 160 and determine whether the remote proactive cache includes a proactive cache entry with assistant request parameters that match the recognized text 171A and / or the remotely determined NLU data for the recognized text 171A. If so, the remote fulfillment engine 183 may optionally provide the on-device execution engine 176 with remote fulfillment data 183A including action content from the matching remote proactive cache entry. Optionally, remote fulfillment engine 183 provides remote fulfillment data 183A only in response to an indication (by on-device fulfillment engine 175) that local fulfillment has failed and / or in response to lack of receipt of a “stop” command from client device 160 (which may be provided when local fulfillment is successful). Remote fulfillment engine 183 may optionally utilize other techniques for generating fulfillment data 183A (e.g., action content) “on the fly.” This may be done in parallel with accessing remote proactive cache 184A to determine whether a machine remote proactive cache entry exists and / or may be performed in response to a determination that no matching remote proactive cache entry exists.Because the action content in the remote proactive cache 184A has already been pre-generated, the remote fulfillment data may be retrieved more quickly and using fewer resources than if the content were generated on the fly by the remote fulfillment engine 183. Thus, various implementations may provide remote fulfillment data based solely on the remote proactive cache 184A when a matching remote proactive cache entry is identified, optionally without attempting to generate fulfillment data on the fly and / or ceasing on-the-fly generation when a match is determined.

[0071] 3 shows some non-limiting examples of proactive cache entries 310, 320, and 330. Such proactive cache entries, along with a large number of additional entries, may be stored within local proactive cache 178 (FIG. 1) or in one of remote proactive caches 184 (FIG. 1).

[0072] Proactive cache entry 310 includes request parameters 310A that represent various assistant requests for performing a given action to obtain tomorrow's local weather forecast. Request parameters 310A include textual representations of "tomorrow's weather," "weather tomorrow," and a structured NLU data representation that specifies the intent of "weather" and slot values ​​of "tomorrow" for the "day" slot and "local" for the "location" slot. Action content 310B of proactive cache entry 310 includes text describing tomorrow's local weather and a graphic conveying tomorrow's local weather. Both the text and the graphic can be rendered in response to a determination that user input (e.g., spoken or typed speech) matches request parameters 310A. Optionally, synthesized speech based on the text can also be rendered in response. Metadata 310C of proactive cache entry 310 includes a 12-hour TTL and a timestamp. As described herein, proactive cache entry 310 can be removed from the proactive cache (or at least can no longer be utilized) once it is determined that the TTL has expired. As also described herein, a proactive cache entry 310 may be provided to a given client device in a given geographic area, while other proactive cache entries having the same request parameters but different action content may be provided to other client devices in other geographic areas.

[0073] Proactive cache entry 320 includes request parameters 320A that represent various Assistant requests for performing a given action to access the thermostat schedule adjustment state of the corresponding application. Request parameters 320A include textual representations of "adjust thermostat schedule," "change schedule for thermostat," and structured NLU data representations that specify the intent of "thermostat" and the slot value of "change / adjust schedule" for the "settings" slot. Action content 320B of proactive cache entry 320 includes a deep link to a specific application. The deep link can be executed in response to a determination that user input (e.g., a spoken or typed utterance) matches request parameters 320A. Executing the deep link opens the specific application in a state where the thermostat schedule setting can be adjusted (i.e., the application is in a state where the next user input can cause the action to be performed). Metadata 320C of proactive cache entry 320 includes a 30-day TTL and a timestamp. As described herein, a proactive cache entry 320 may be removed from the proactive cache (or at least may no longer be utilized) once the TTL is determined to have expired. As also described herein, a proactive cache entry 320 may be provided to a given client device based on determining that the client device has a particular application (corresponding to a deep link in action content 320B) installed and / or indicated as a primary thermostat application, while other proactive cache entries having the same request parameters but different action content (e.g., different deep links) may be provided to other client devices that do not have the particular application installed.

[0074] Proactive cache entry 330 includes request parameters 330A that represent various assistant requests to perform a given action to obtain an estimated net worth for "John Doe" (a hypothetical person). Request parameters 330A include textual representations of "John Doe's net worth" and "How much is John Doe's wealth?" Action content 330B of proactive cache entry 330 includes text describing John Doe's net worth. The text may be rendered in response to a determination that user input (e.g., spoken or typed speech) matches request parameters 330A. Optionally, synthesized speech based on the text may also be rendered in response. Metadata 330C of proactive cache entry 330 includes a 7-day TTL and a timestamp. As described herein, proactive cache entry 330 may be removed from the proactive cache (or at least may no longer be utilized) once it is determined that the TTL has expired. As also described herein, a proactive cache entry 330 may be provided for a given client device based on determining that attributes of the given client device are related to John Doe. This may be based, for example, on past user searches for John Doe, visits to Internet content related to John Doe, and / or searches for other entities that have a strong relationship with John Doe (e.g., as determined based on a knowledge graph or other data structure). In various implementations, the proactive cache entry 330 may be generated and / or provided based at least in part on a determination of an event related to John Doe, such as an increase in requests and / or Internet content related to John Doe. For example, a proactive cache entry for John Doe's net worth may be generated based on determining that John Doe is a celebrity (class) and that frequent queries for celebrities have the template "what is [celebrity alias]'s net worth."Also, for example, proactive cache entries may be provided for storage in local or remote proactive cache entries based on ranking criteria that are influenced by the growth of requests and / or internet content related to John Doe.

[0075] 4 depicts a flowchart illustrating an example method 400 for prefetching and storing proactive cache entries according to implementations disclosed herein. For convenience, the operations of method 400 are described with reference to a system that performs those operations. The system may include various components of various computer systems, such as one or more components of a client device (e.g., prefetch engine 177 of FIG. 1). Furthermore, although the operations of method 400 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0076] At block 410, the system determines whether one or more prefetch conditions have occurred. If not, the system continues to determine whether a prefetch condition has occurred. If so, the system proceeds to block 420. The prefetch conditions may include, for example, one or more of: certain network conditions exist for the client device; the client device is charging and / or has at least a threshold battery charge state; the client device is not actively being utilized by a user; the client device's current processor utilization and / or current memory utilization does not exceed certain thresholds; and / or a certain amount of time (e.g., at least one hour) has elapsed since the most recent prefetch request.

[0077] The system sends a prefetch request at block 420. The prefetch request may optionally include an identifier for the client device and / or a token or other indication of a proactive cache entry already stored locally at the client device.

[0078] In block 430, the system receives the proactive assistant cache entry in response to the request in block 420.

[0079] In block 440, the system stores the received proactive assistant cache entry in the local proactive cache. Block 440 may optionally include block 440A, in which the system removes one or more existing proactive cache entries from the local proactive cache to make room for the received proactive cache entry. After block 440, the system may optionally proceed again to block 410 after a threshold amount of time has elapsed.

[0080] FIG. 5A shows a flowchart illustrating an example method 500 for generating proactive cache entries, provisioning a local subset of proactive cache entries for a given client device, and / or determining a remote subset of proactive cache entries. FIG. 5B shows a flowchart illustrating several implementations of block 510 of FIG. 5A. FIG. 5C shows a flowchart illustrating several additional or alternative implementations of block 510 of FIG. 5A. For convenience, the operations of method 500 are described with reference to a system that performs those operations. This system may include various components of various computer systems, such as one or more components of a remote server (e.g., the proactive cache system of FIG. 1). Furthermore, although the operations of method 500 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0081] 5A , in block 510, the system generates a proactive assistant cache entry. In generating each proactive assistant cache entry, block 510 may include sub-blocks 512, 514, and / or 516. In block 512, the system generates assistant request parameters for the proactive assistant cache entry. In block 514, the system generates action content for the proactive assistant cache entry. In block 516, the system generates metadata for the proactive assistant cache entry. Additional description of implementations of block 510 is provided below with respect to FIGS. 5B and 5C, and elsewhere herein.

[0082] In block 520, the system selects a client device.

[0083] In block 530, the system determines a local subset of proactive assistant cache entries for the selected client device. The system determines the local subset based on attributes of the client device and attributes of the proactive assistant cache entries. For example, some of the proactive assistant cache entries may be selected based on having attributes corresponding to attributes of the client device. Such attributes may include an application (e.g., installed on the client device and corresponding to a deep link in the proactive assistant cache entry), a geographic location (e.g., of the client device and corresponding to the proactive assistant cache entry), an entity interacted with via the client device (e.g., through a search or visit of content) compared to the entity corresponding to the proactive assistant cache entry, and / or other attributes. Also, for example, some of the proactive cache entries may additionally or alternatively be selected based on ranking criteria for the proactive cache entries. The amount of proactive assistant cache entries included in the local subset may be affected by the size of the proactive assistant cache on the client device.

[0084] At block 540, the system optionally determines a remote subset of proactive cache entries for the selected client device. The remote subset may include (or be limited to) proactive cache entries that are not included in the local subset of block 530. In some implementations, block 540 may consider attributes of the client device and attributes of the proactive assistant cache entries when determining the remote subset of block 540. The amount of proactive assistant cache entries included in the remote subset may also be affected by the size of the remote proactive assistant cache. The remote proactive assistant cache may be specific to the client device or may be specific to a restricted group of client devices that includes that client device and other similar client devices. The remote proactive assistant cache may be utilized by the remote automated assistant component in reducing latency in provisioning responses to requests originating from the client device, as described herein.

[0085] The system may then return to block 520, select another client device, and perform blocks 530 and 540 for the other client device. It should be understood that multiple iterations of blocks 520, 530, and 540 may each be performed in parallel for different client devices. It should further be understood that blocks 520, 530, and 540 may be repeated for various client devices at regular or irregular intervals to maintain an up-to-date local and / or remote proactive assistant cache and to account for newly generated proactive assistant cache entries that may be generated through multiple iterations of block 510.

[0086] At block 550, the system optionally receives a prefetch request from a given client device. For example, the given client device may send a prefetch request as described with respect to method 400.

[0087] In block 560, the system provides to the given client device one or more proactive assistant cache entries of the local subset determined for the given client device in the iteration of block 530. Which entries of the local subset are provided may be based on determining which, if any, are already stored in the local proactive cache of the given client device. For example, optionally, only entries that are not already in the local proactive cache may be provided in block 560. The provided proactive assistant cache entries may be stored in the local proactive cache by the automated assistant application of the given client device for use in locally fulfilling future user interface inputs provided at the given client device.

[0088] When block 550 is performed, block 560 may be performed in response to a prefetch request. When block 550 is not performed, block 560 may include proactively pushing a proactive assistant cache entry independent of an explicit request from a given client device. It should be understood that block 560 (and optionally, block 550) would be performed for each of a large number of client devices and would be performed at multiple points in time for each of the client devices. It should further be understood that throughout method 500, heterogeneous proactive cache entries would be provided to different client devices (and / or for storage in corresponding remote proactive caches) and updated over time.

[0089] Referring to FIG. 5B, a flowchart 510B illustrating some implementations of block 510 of FIG. 5A is provided. In block 511B, the system determines whether an event exists for the entity. If not, the system continues to monitor events in block 511B. If so, the system optionally generates a proactive cache entry based on the event by proceeding to block 512B, generating assistant request parameters for the entity, proceeding to block 514B, generating action content for the assistant request parameters, and proceeding to block 516B, generating metadata for the proactive cache entry. The proactive cache entry will include the assistant request parameters, the action content, and the metadata. In block 517B, the system determines whether to generate a further proactive assistant cache entry for the entity. If so, the system optionally performs another iteration of block 512B and another iteration of blocks 514B and 516B to generate another proactive assistant cache entry for the entity. If not, the system returns to block 511B to monitor for another event for the entity and / or another entity.

[0090] As an example, in block 511B, the event may be a change to the weather forecast for a geographic area. Continuing with this example, a proactive cache entry may be generated that reflects new action content (describing the new weather forecast) in block 514B and new metadata (e.g., a timestamp) in block 516B, but maintains the same Assistant request parameters as the previous entry for the weather forecast for the geographic area.

[0091] As another example, in block 511B, the event may be an increase in music streaming requests for a particular musical artist. Continuing with this example, a proactive cache entry may be generated that includes the Assistant request parameters generated in 512B for an Assistant request to stream music from a particular musical artist, action content generated in block 514B that includes a deep link for streaming the particular artist for the first application, and metadata (e.g., a TTL value) generated in block 516B. Further, in block 517B, it may be determined to generate another proactive cache entry that includes the same Assistant request parameters and / or metadata but includes a different deep link in the action content. The different deep link is generated in another iteration of block 514B and includes a deep link for streaming the particular musical artist for the second application.

[0092] As yet another example, in block 511B, the event may be an increase in requests and / or content for a particular city. Continuing with this example, a proactive cache entry may be generated that includes the Assistant request parameters generated in 512B based on frequent Assistant requests for other cities (e.g., requesting the population of a particular city), the action content generated in block 514B including a visual and / or text response (e.g., a visual and / or textual representation of the population), and the metadata (e.g., a TTL value) generated in block 516B. Further, in block 517B, it may be determined to generate an additional proactive cache entry that is based on other frequent Assistant requests for other cities. For example, an additional proactive cache entry may be generated that includes the Assistant request parameters generated in 512B based on other frequent Assistant requests for other cities (e.g., requesting weather information for a particular city), the action content generated in block 514B including a visual and / or textual response (e.g., a visual and / or textual representation of weather information), and the metadata (e.g., a TTL value) generated in block 516B.

[0093] 5C, a flowchart 510C is provided illustrating some additional or alternative implementations of block 510 of FIG. 5A. For example, some iterations of block 510 may be performed based on flowchart 510B, and other iterations may be performed based on flowchart 510C. In block 512C, the system generates Assistant request parameters. As an example, Assistant request parameters for streaming Bluegrass music may be generated, such as parameters for "play some Bluegrass," "play some Bluegrass," and / or "{Intent=stream music; genre=bluegrass}."

[0094] In block 514C, the system generates action content by determining N applications for the request parameters in sub-block 514C1, where N is an integer greater than 1. For example, the system may determine 15 applications for streaming bluegrass music. Further, in sub-block 514C2, the system determines action content for each of the N applications. For example, the system determines, as action content for each of the N applications, a corresponding deep link that, when executed, causes the corresponding application to stream bluegrass music.

[0095] In block 516C, the system generates the metadata.

[0096] In block 518C, the system generates N proactive assistant cache entries. Each of the N generated proactive assistant cache entries has the same assistant request parameters of 512C and, optionally, the same metadata of 516C, but includes different action content (i.e., each may include action content with only a corresponding deep link to a single application among the N applications).

[0097] 6 is a block diagram of an example computing device 610 that may optionally be utilized to implement one or more aspects of the techniques described herein. In some implementations, one or more of the client device, cloud-based automated assistant component, and / or other component may comprise one or more components of the example computing device 610.

[0098] The computing device 610 typically includes at least one processor 614 that communicates with several peripheral devices via a bus subsystem 612. These peripheral devices may include, for example, a storage subsystem 624 including a memory subsystem 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices enable user interaction with the computing device 610. The network interface subsystem 616 provides an interface to external networks and is coupled to corresponding interface devices in other computing devices.

[0099] The user interface input devices 622 may include pointing devices such as a keyboard, a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touchscreen integrated into a display, a voice input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to include all conceivable types of devices and methods for inputting information into the computing device 610 or over a communications network.

[0100] The user interface output devices 620 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visual image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to include all conceivable types of devices and methods that output information from the computing device 610 to a user or to another machine or computing device.

[0101] Storage subsystem 624 stores programming and data configurations that provide the functionality of some or all of the modules described herein. For example, storage subsystem 624 may include logic for performing selected aspects of the methods described herein and for implementing the various components shown herein.

[0102] These software modules are generally executed by the processor 614, alone or in combination with other processors. The memory 625 used within the storage subsystem 624 may include several memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution, and a read-only memory (ROM) 632 in which fixed instructions are stored. The file storage subsystem 626 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of some implementations may be stored by the file storage subsystem 626, within the storage subsystem 624, or in other machines accessible by the processor 614.

[0103] The bus subsystem 612 provides a mechanism for allowing the various components and subsystems of the computing device 610 to communicate with each other, as intended. Although the bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0104] The computing device 610 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computing device 610 shown in Figure 6 is intended only as a specific example to illustrate some implementations. Many other configurations of the computing device 610 are possible, having more or fewer components than the computing device shown in Figure 6.

[0105] In situations where the systems described herein may collect or possibly monitor personal information about a user or utilize personal and / or monitored information, the user may be given the opportunity to control whether a program or feature collects user information (e.g., the user's social network, social actions or activities, occupation, user preferences, or the user's current geographic location) or to control whether and / or how content is received from content servers that may be more relevant to the user. Also, some data may be treated in one or more ways before being stored or used so that personally identifiable information is removed. For example, the user's identity may be treated so that personally identifiable information cannot be determined about the user, or so that if geographic location information is obtained, the user's geographic location may be generalized (e.g., to the city level, ZIP code level, or state level) so that the user's specific geographic location cannot be determined. Thus, the user can control how information is collected and / or used about the user.

[0106] In some implementations, a method implemented by one or more processors is provided, the method including determining assistant request parameters representing one or more assistant requests to perform a given action. The assistant request parameters define one or more textual representations of the assistant request and / or one or more semantic representations of the assistant request. The method further includes determining that the given action can be performed using a first application and can also be performed using a second application. The method further includes generating first action content for the first application and generating second action content for the second application. The first action content includes a first deep link to the first application. The first deep link is locally executable by an assistant client application of a client device that has the first application installed, and local execution of the first deep link causes the first application to open in a first state for performing the given action. The second action content includes a second deep link to the second application. A second deep link different from the first deep link is locally executable by an assistant client application of a client device that has the second application installed, and local execution of the second deep link causes the second application to open in a second state for performing a given action. The method further includes generating a first proactive assistant cache entry including assistant request parameters and first action content, and generating a second proactive assistant cache entry including assistant request parameters and second action content. The method further includes generating a proactive cache entry for the given client device.Generating the proactive cache entry includes including a first proactive cache entry but not a second proactive cache entry based on the given client device having the first application installed but not the second application installed. Optionally, the method further includes sending the proactive cache entry to the given client device in response to receiving a proactive cache request sent by the given client device. The automated assistant application of the given client device stores the proactive cache entry in a local proactive cache for use by the automated assistant application in locally fulfilling future user interface inputs provided at the given client device.

[0107] These and other implementations of the techniques disclosed herein may include one or more of the following features.

[0108] In some implementations, the method further includes determining an increase in demand for the given action, and generating a first proactive cache entry and a second proactive cache entry in response to determining the increase in demand for the given action.

[0109] In some implementations, the method further includes determining an increase in demand for the given action, and including the first proactive cache entry in the proactive cache entries is further based on the increase in demand for the given action. In some of these implementations, including the first proactive cache entry in the proactive cache entries further based on the increase in demand for the given action includes determining that an entity targeted by the given action corresponds to one or more attributes for a given client device for which the proactive cache entry is generated. In some versions of these implementations, the first application is a music streaming application, the given action is streaming music for the entity, and the attributes for the given client device include association of the given client device with the entity. These requests may include, for example, automated assistant requests and / or additional requests. The automated assistant requests are each generated for a corresponding automated assistant application in response to a corresponding user interface input. The additional requests are in addition to the automated assistant request and arise from one or more additional applications in addition to the automated assistant application.

[0110] In some implementations, the method further includes predicting an increase in demand for the given action, and generating a first proactive cache entry and a second proactive cache entry in response to determining the increase in demand for the given action. In some additional or alternative implementations, the method further includes predicting an increase in demand for the given action, and including the first proactive cache entry in the proactive cache entry is further based on the predicted increase in demand for the given action. Predicting an increase in demand for the given action may include determining an increase in Internet content related to the entity targeted by the given action and / or determining future events related to the entity targeted by the given action.

[0111] In some implementations, the method further includes generating a lifetime value for the first proactive assistant cache entry and including the lifetime value in the first proactive assistant cache entry. The lifetime value causes the given client device to remove the first proactive assistant cache entry from the local proactive cache in response to expiration of a duration defined by the lifetime value. In some of these implementations, the method further includes removing the first proactive assistant cache entry from the proactive cache entries based on an assistant client application of the given client device comparing the lifetime value to a timestamp for the first proactive assistant cache entry.

[0112] In some implementations, the method further includes, by an assistant client application of the given client device, receiving a proactive cache entry in response to transmitting a proactive cache request, and storing the proactive cache entry in a local proactive cache of the given client device. In some of these implementations, the method further includes, after the assistant client application of the given client device stores the proactive cache entry in the local proactive cache, generating recognized text using on-device speech recognition based on a spoken utterance captured in audio data detected as a user interface input by one or more microphones of the client device, determining, based on accessing the proactive cache, that assistant request parameters of a first proactive assistant cache entry match the recognized text and / or natural language understanding data generated based on the recognized text, and, in response to determining a match, locally executing a first deep link to open the first application in a first state to perform the given action.

[0113] In some implementations, the method further includes the steps of: an assistant client application of a given client device determining that a network status of the given client device and / or a computational load status of the given client device satisfies one or more conditions; sending a proactive cache request in response to determining that the network status and / or the computational load status satisfies the one or more conditions; receiving a proactive cache entry in response to sending the proactive cache request; and storing the proactive cache entry in a local proactive cache of the given client device.

[0114] In some implementations, a method implemented by one or more processors is provided, the method including determining the occurrence of an event related to a particular entity. The method further includes generating one or more proactive assistant cache entries for the particular entity in response to determining the occurrence of the event related to the particular entity. Each of the proactive assistant cache entries defines a respective assistant request parameter and a respective assistant action content. The respective assistant request parameters each represent one or more respective assistant requests related to the particular entity and define one or more textual representations of the assistant requests and / or one or more semantic representations of the assistant requests. The respective assistant action content is locally interpretable by an assistant client application on the client device to cause the assistant client application to locally perform an assistant action with respect to the particular entity and in response to the one or more respective assistant requests. The method further includes selecting a subset of client devices based on determining that the client devices in the subset each have one or more corresponding attributes corresponding to the particular entity. The method further includes sending a proactive assistant cache entry for the particular entity to a plurality of client devices in the subset without sending the proactive assistant cache entry to other client devices not in the selected subset. The step of transmitting the proactive assistant cache entry causes the corresponding automated assistant application of each of the client devices to locally cache the proactive assistant cache entry in a local proactive cache for use by the automated assistant application in locally fulfilling future spoken utterances provided at the given client device.

[0115] These and other implementations of the techniques disclosed herein may include one or more of the following features.

[0116] In some implementations, determining the occurrence of an event related to the particular entity includes determining an increase in demand for the particular entity and / or an increase in Internet content related to the particular entity.

[0117] In some implementations, generating a given one of the one or more proactive assistant cache entries for the particular entity includes generating one or more respective assistant requests for the particular entity based on one or more attributes of the particular entity, and generating respective assistant request parameters based on the one or more respective assistant requests.

[0118] In some implementations, the particular entity is a particular person or a particular organization.

[0119] In some implementations, generating one or more respective assistant requests based on one or more attributes of a particular entity includes determining a class of the entity, determining templates for most frequent queries for entities of the class, and generating at least one of the respective assistant requests using the templates and aliases of the entity.

[0120] In some implementations, the event related to the particular entity is a changed attribute for the particular entity or a new attribute for the particular entity. In some of these implementations, generating one or more proactive assistant cache entries for the particular entity includes generating the given one of the proactive assistant cache entries by modifying respective action content for a previously generated proactive assistant cache entry to cause the attribute to be rendered during local performance of an action for the given one of the proactive assistant cache entries. In some versions of these implementations, the particular entity is weather in a geographic area, and the attributes are high temperature in the geographic area, low temperature in the geographic area, and / or probability of precipitation in the geographic area. In some other versions of these implementations, the particular entity is an event, and the attributes are a start time of the event, an end time of the event, and / or a location of the event.

[0121] In some implementations, the method further includes selecting a second subset of client devices that do not overlap with the client devices of the subset. The method further includes storing the proactive assistant cache entries in one or more remote proactive caches that are utilized in response to assistance requests from the client devices of the second subset. In some of these implementations, the one or more remote proactive caches include a corresponding one of the remote proactive caches for each of the client devices of the second subset.

[0122] In some implementations, the method further includes: an assistant client application of a given client device of the subset of client devices sending a proactive cache request; and, in response to receiving the proactive cache request, sending a proactive assistant cache entry to the given client device.

[0123] In some implementations, the method further includes: an assistant client application of a given client device of the subset of client devices receiving a proactive cache entry; and storing the proactive cache entry in a given local proactive cache of the given client device. In some of these implementations, after storing the proactive cache entry in the local proactive cache, the assistant client application of the given client device uses on-device speech recognition to generate recognized text based on spoken utterances captured in audio data detected by one or more microphones of the client device; and, based on accessing the proactive cache, determining that the assistant request parameters of each given one of the proactive assistant cache entries match the recognized text and / or natural language understanding data generated based on the recognized text. In response to the determination of the match, the method further includes locally interpreting the respective assistant action content of the given one of the proactive assistant cache entries. In some versions of these implementations, locally interpreting the respective assistant action content of a given one of the proactive assistant cache entries includes rendering the text and / or graphical content of the assistant action content on the client device.

[0124] In some implementations, a method implemented by one or more processors is provided, the method including determining the occurrence of an event related to a particular entity. The method further includes selecting a subset of client devices based on determining that the client devices of the subset each have one or more corresponding attributes corresponding to the particular entity based on determining the occurrence of the event related to the particular entity. The method further includes sending one or more proactive assistant cache entries for the particular entity to a plurality of client devices of the subset without sending the proactive assistant cache entry to other client devices not in the selected subset. Each proactive assistant cache entry for the particular entity defines respective assistant request parameters representing one or more respective assistant requests regarding the particular entity and respective assistant action content that is locally interpretable by an assistant client application of the client device to cause the assistant client application to locally perform an assistant action regarding the particular entity and in response to the one or more respective assistant requests. Sending the proactive assistant cache entry causes the corresponding automated assistant application of each of the client devices to locally cache the proactive assistant cache entry in a local proactive cache for use by the automated assistant application in locally fulfilling future spoken utterances provided at the given client device.

[0125] These and other implementations of the techniques disclosed herein may include one or more of the following features.

[0126] In some implementations, determining the occurrence of an event related to the particular entity includes determining an increase in internet content related to the particular entity, determining an increase in requests for the particular entity, including requests originating from one or more non-automated assistant applications, and / or determining a change in attributes of the particular entity through interactions with one or more remote servers. [Explanation of symbols]

[0127] 120 Proactive Cache System 130 Cache Entry Generation Engine 131 Request Parameter Module 132 Action Content Module 133 Metadata Module 134 cache candidates 140 Cache Assembly Engine, Event Cache Assembly Engine 141 Local Module 142 Remote Module 145 On-Device Fulfillment Engine 150 Entity Event Engine 151 Remote Systems 160 client devices 163 Visual Components 165 microphones 167 Presence Sensor 169 Display 170 Automated Assistant Clients 171 Audio Capture Engine 171A Recognized Text 172 Visual Capture Engine 173 On-device speech recognition engine 174 On-device NLU engine 174A NLU Data 175 On-Device Fulfillment Engine 175A Performance Data 176 On-Device Execution Engine 177 Prefetch Engine 178 Proactive Cache 180 Remote Automated Assistant Components 181 Remote ARS Engine 182 Remote NLU Engine 183 Remote Fulfillment Engine 183A Remote Performance Data, Performance Data 184 Remote Proactive Cache 184A Remote Proactive Cache 195 Automated Assistants 221 Request 222 Proactive Cache Entries, Cache Entries 223 Audio Data 310 Proactive Cache Entries 310A Required Parameters 310B Action Content 310C Metadata 320 proactive cache entries 320A Required Parameters 320B Action Content 320C Metadata 330 Proactive Cache Entries 330A Required Parameters 330B Action Content 330C Metadata 400 ways 500 ways 610 Computing Devices 612 Bus Subsystem 614 processor 616 Network Interface Subsystem 620 User Interface Output Device 622 User Interface Input Devices 624 Storage Subsystem 625 Memory Subsystem, Memory 626 File Storage Subsystem

Claims

1. A method implemented by one or more processors, comprising: determining the occurrence of an event associated with a particular entity, the event associated with a music streaming request; determining that a particular client device is associated with the event associated with the particular entity; In response to determining an occurrence of the event associated with the particular entity and determining that the particular client device is associated with the event associated with the particular entity, sending a proactive assistant cache entry for the particular entity to the particular client device; The proactive assistant cache entries define respective assistant action content, wherein the respective assistant action content includes text and is interpretable locally by an assistant client application of the particular client device to cause an assistant action to be performed locally by the assistant client application with respect to a particular entity and in response to an assistant request; The assistant action includes rendering audio data generated by an on-device speech synthesizer using the text of the assistant action content; The method, wherein the step of sending the proactive assistant cache entry causes the assistant client application of the particular client device to locally cache the proactive assistant cache entry for use in performing the assistant action in response to future verbal utterances provided at the given client device and determined to correspond to the assistant request.

2. Receiving the proactive assistant cache entry at the particular client device. The method of claim 1 further comprising:

3. The method of claim 2, further comprising: processing, at the particular client device, speech detected via one or more microphones of the client device; determining at the particular client device that the proactive assistant cache entry is utterance-responsive; In response to determining that the proactive assistant cache entry is responsive to an utterance, using the proactive assistant cache entry at the particular client device to generate the audio data using the on-device speech synthesizer; The method of claim 2 further comprising:

4. The method further comprises selecting the proactive assistant cache entry from a plurality of proactive assistant cache entries for the particular entity, wherein the step of selecting the proactive assistant cache entry is based on attributes of the particular client device and on determining that the particular client device is associated with the event related to the particular entity, 2. The method of claim 1, wherein the step of sending the proactive assistant cache entry to the particular client device is in response to the step of selecting the proactive assistant cache entry.

5. The method of claim 4, wherein the attribute is a particular music streaming application designated as preferred for the particular client device.

6. The method of claim 1, wherein the specific entity is a specific musical artist.

7. The method described in claim 6, wherein the event is an increase in music streaming requests related to the particular musical artist.

8. The method of claim 1, wherein the proactive assistant cache entry further defines a lifetime value, and the step of sending the proactive assistant cache entry causes the assistant client application of the particular client device to locally cache the proactive assistant cache entry for a duration based on the lifetime value.

9. The method of claim 8, further comprising: determining an occurrence of an event associated with a particular entity, the event associated with a music streaming request; determining that a particular client device is associated with the event associated with the particular entity; In response to determining an occurrence of the event associated with the particular entity and determining that the particular client device is associated with the event associated with the particular entity, and transmitting a proactive assistant cache entry for the particular entity to the particular client device, The proactive assistant cache entries define respective assistant action content, wherein the respective assistant action content includes text and is interpretable locally by an assistant client application of the particular client device to cause an assistant action to be performed locally by the assistant client application with respect to a particular entity and in response to an assistant request; The assistant action includes rendering audio data generated by an on-device speech synthesizer using the text of the assistant action content; Sending the proactive assistant cache entry causes the assistant client application of the particular client device to locally cache the proactive assistant cache entry for use in performing the assistant action in response to future verbal utterances provided at the given client device and determined to correspond to the assistant request.

10. The system described in claim 9, further comprising one or more client device processors of the particular client device operable to execute client instructions for receiving the proactive assistant cache entry.

11. The client device processor according to claim 1, wherein one or more of the client device processors: processing speech detected via one or more microphones of the client device; determining that the proactive assistant cache entry is responsive to an utterance; In response to determining that the proactive assistant cache entry is responsive to an utterance, using the proactive assistant cache entry to generate the audio data using the on-device speech synthesizer; and 11. The system of claim 10, further operable to execute the client instructions to:

12. The server processors according to claim 1, wherein one or more of the server processors: further operable to execute the instructions for selecting the proactive assistant cache entry from a plurality of proactive assistant cache entries for the particular entity, wherein in selecting the proactive assistant cache entry, one or more of the server processors select the proactive assistant cache entry based on attributes of the particular client device and determining that the particular client device is associated with the event related to the particular entity; 10. The system of claim 9, wherein when sending the proactive assistant cache entry to the particular client device, one or more of the server processors select the proactive assistant cache entry and send the proactive assistant cache entry accordingly.

13. The system of claim 12, wherein the attribute is a particular music streaming application designated as preferred for the particular client device.

14. The system of claim 9, wherein the particular entity is a particular musical artist.

15. The system of claim 14, wherein the event is an increase in music streaming requests related to the particular musical artist.

16. The system of claim 9, wherein the proactive assistant cache entry further defines a lifetime value, and sending the proactive assistant cache entry causes the assistant client application of the particular client device to locally cache the proactive assistant cache entry for a duration based on the lifetime value.

Citation Information

Patent Citations

  • Local and remote speech processing

    JP2016531375A

  • Optimizing user interface data caching for future actions

    WO2018125276A1