Using a large language model in generating automatic assistant responses

By employing a large language model to process audio data and generate contextually relevant responses, the automatic assistant enhances natural conversation flow and user engagement, addressing the limitations of existing systems.

JP7691523B2Active Publication Date: 2025-06-11GOOGLE LLC

Patent Information

Application Number
JP2023569990
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-22
Filing Date
2021-11-30
Publication Date
2025-06-11
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

Existing automatic assistants struggle to maintain natural conversation flow with users, often providing responses that do not engage the user or require additional user input for context information.

Method used

The implementation uses a large language model (LLM) to process a stream of audio data and generate modified assistant outputs that are contextually relevant and engaging, allowing the automatic assistant to proactively provide information and adapt its personality based on the conversation context.

Benefits of technology

This approach enables the automatic assistant to engage in more natural and conversational interactions with users, reducing the need for additional user input and improving the overall user experience by providing contextually relevant information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691523000001
    Figure 0007691523000001
  • Figure 0007691523000002
    Figure 0007691523000002
  • Figure 0007691523000003
    Figure 0007691523000003
Patent Text Reader

Abstract

As part of an interaction session between a user and an automated assistant, an implementation may receive a stream of audio data capturing an utterance including an assistant query, and based on processing the stream of audio data, determine a set of assistant outputs each predicted to be responsive to the assistant query, and process the assistant outputs and the context of the interaction session using a large-scale language model (LLM) output to generate a set of modified assistant outputs, from among the set of modified assistant outputs, such that a given modified assistant output is provided for presentation to the user in response to the utterance. In some implementations, the LLM output may be generated in an offline manner for later use in an online manner. In additional or alternative implementations, the LLM output may be generated in an online manner as the utterance is received.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Humans may engage in human-computer conversations with a dialog software application, referred to herein as an "automatic assistant" (also called a "chatbot", "conversational personal assistant", "intelligent personal assistant", "personal voice assistant", "conversational agent", etc.). Automatic assistants typically rely on a pipeline of components when interpreting and responding to utterances. For example, an automatic speech recognition (ASR) engine can process audio data corresponding to a user's utterance to generate an ASR output such as an ASR hypothesis of the utterance (i.e., a sequence of terms and / or other tokens). Further, a natural language understanding (NLU) engine can process the ASR output (or the touch / typed input) to generate an NLU output such as the request (e.g., intent) expressed by the user when providing the utterance (or the touch / typed input), and optionally, slot values for parameters related to that intent. Finally, the NLU output can be processed by various fulfillment components to generate a fulfillment output such as response content for responding to the utterance and / or one or more actions that can be performed in response to the utterance.

[0002] Generally, a conversation session with an automatic assistant is initiated by a user providing an utterance, and the automatic assistant can respond to the utterance using the pipeline of the aforementioned components. The user can continue the conversation session by providing additional utterances, and the automatic assistant can respond to the additional utterances again using the pipeline of the aforementioned components. In other words, these conversation sessions are generally turn-based in that the turn to provide an utterance in the conversation session is with the user, the turn to respond to an utterance in the conversation session is with the automatic assistant, the additional turn to provide an additional utterance in the conversation session is with the user, the additional turn to respond to the additional utterance in the conversation session is with the automatic assistant, and so on. However, from the user's perspective, these turn-based conversation sessions may not be natural because they do not reflect the way humans actually talk to each other.

[0003] For example, if a first person provides an utterance to convey the first thought to a second person during a conversation session (e.g., "I'm going to the beach today"), the second person can consider that utterance in the context of the conversation session when stating a response to the first person (e.g., "sounds fun, what are you going to do at the beach?", "nice, have you looked at the weather?", etc.). In particular, the second person can provide an utterance that keeps the first person involved in the conversation session in a natural way when responding to the first person. In other words, during the conversation session, rather than one person leading the conversation session, both the first person and the second person can provide utterances to facilitate a natural conversation.

[0004] However, in the above example, when the second person is replaced by an automatic assistant, the automatic assistant may not provide a response that keeps the first person engaged in the conversation session. For example, in response to the first person providing the utterance "I'm going to the beach today", the automatic assistant may simply respond with "sound fun" or "nice" without providing any additional response to prompt the conversation session, such as taking the initiative to ask the first person what they plan to do at the beach, taking the initiative to check the weather forecast for the beach that the first person often visits and including the weather forecast in the response, or making some speculation based on the weather forecast. As a result, the response provided by the automatic assistant in response to the first person's utterance may not reflect a natural conversation among multiple people and may not resonate with the first person. Furthermore, the first person may have to provide additional utterances to explicitly request some information (such as the weather forecast for the beach) that the automatic assistant could have provided proactively, increasing the amount of utterances directed at the automatic assistant and wasting the computing resources of the client device utilized when processing these utterances. Summary of the Invention Means for Solving the Problems

[0005] The implementations described in this specification are directed to enabling an automated assistant to conduct a natural conversation with a user during an interaction session. Some implementations can receive a stream of audio data that captures the user's utterance. The stream of audio data may be generated by one or more microphones of a client device, and the utterance may include an assistant query. Some implementations can further process the stream of audio data to determine a set of automated assistants and process a set of assistant outputs and the context of the interaction session to generate a set of assistant outputs modified using one or more LLM outputs generated using a large language model (LLM). Each of the one or more LLM outputs can be determined based on at least a portion of the context of the interaction session and one or more of the assistant outputs included in the set of assistant outputs. Some implementations can further select, from the set of modified assistant outputs, a given modified assistant output to provide to the user for presentation. Additionally, each of the one or more LLM outputs can include, for example, a probability distribution over a sequence of one or more words and / or phrases spanning one or more vocabularies, and one or more of the sequence of words and / or phrases can be selected as one or more LLM outputs based on the probability distribution. Moreover, the context of the interaction session can be determined based on one or more context signals including, for example, time, day of the week, location of the client device, ambient noise detected in the environment of the client device, user profile data, software application data, environmental data about a known environment of the user of the client device, interaction history of the interaction session between the user and the automated assistant, and / or other context signals.

[0006] In some implementations, the set of assistant outputs can be determined based on processing a stream of audio data using a streaming ASR model to generate a stream of ASR outputs, such as one or more recognized terms or phrases predicted to correspond to the utterance, one or more phonemes predicted to correspond to the utterance, one or more predicted metrics associated with each of the one or more recognized terms or phrases and / or one or more predicted phonemes, and / or other automatic speech recognition (ASR) outputs. Further, the ASR output can be processed using an NLU model to generate a stream of natural language understanding (NLU) outputs, such as one or more predicted intents of the user in providing the utterance, and one or more corresponding slot values for one or more parameters associated with each of the one or more predicted intents. Additionally, the stream of NLU data can be processed by one or more first party (1P) and / or third party (3P) systems to generate the set of assistant outputs. As used herein, one or more 1P systems include systems developed and / or maintained by the same entity (e.g., a common publisher) that develops and / or maintains the automatic assistant described herein, while one or more 3P systems include systems developed and / or maintained by an entity separate from the entity that develops and / or maintains the automatic assistant described herein. In particular, the set of assistant outputs described herein includes the assistant outputs typically considered in response to an utterance. However, by using the claimed techniques, the set of assistant outputs generated in the manner described above can further be processed to generate a modified set of assistant outputs. Specifically, the set of assistant outputs may be modified using one or more LLM outputs, and a given modified assistant output can be selected from the set of modified assistant outputs as the one to be provided to the user in response to receiving the utterance for presentation.

[0007] For example, assume that a user of a client device provides an utterance such as "Hey Assistant, I'm thinking about going surfing today". In this example, the automatic assistant can process the utterance in the manner described above to generate a set of assistant outputs and a set of modified assistant outputs. The assistant outputs included in the set of assistant outputs in this example may include, for example, "That sounds like fun!", "Sounds fun!", etc. Further, the assistant outputs included in the set of modified assistant outputs in this example may include, for example, "That sounds like fun, how long have you been surfing?", "Enjoy it, but if you're going to Example Beach again, be prepared for some light showers", etc. In particular, the assistant outputs included in the set of assistant outputs do not include any assistant output that pulls the conversation session in such a way as to further engage the user of the client device in the conversation session, but the assistant outputs included in the set of modified assistant outputs include assistant outputs that pull the conversation session in a way that further engages the user of the client device in the conversation session by asking contextually relevant questions (such as "how long have you been surfing?"), assistant outputs that provide contextually relevant information (such as "but if you're going to Example Beach again, be prepared for some light showers"), and / or assistant outputs that otherwise sound different to the user of the client device within the context of the conversation session.

[0008] In some implementations, the set of modified assistant responses can be generated using one or more LLM outputs generated in an online manner. For example, in response to receiving an utterance, an automated assistant can cause the set of assistant outputs to be generated in the manner described above. Further, also in response to receiving an utterance, an automated assistant can cause the set of assistant outputs, the context of the conversation session, and / or the assistant query included in the utterance to be processed using one or more LLMs in order to generate a set of modified assistant outputs based on one or more LLM outputs generated using one or more LLMs.

[0009] In additional or alternative implementations, the set of modified assistant responses can be generated using one or more LLM outputs generated in an offline manner. For example, prior to receiving an utterance, an automated assistant can obtain a plurality of assistant queries and the corresponding context of the corresponding previous conversation session for each of the plurality of assistant queries from an assistant activity database, which can be limited assistant activity of a user of a client device. Further, an automated assistant can cause the set of assistant outputs to be generated in the manner described above and for a given assistant query. Moreover, an automated assistant can cause the set of assistant outputs, the corresponding context of the conversation session, and / or the given assistant query to be processed using one or more LLMs in order to generate a set of modified assistant outputs based on one or more LLM outputs generated using one or more LLMs. This process can be repeated for each of the plurality of queries and the corresponding context of the previous conversation sessions obtained by the automated assistant.

[0010] In addition, the automatic assistant can index one or more LLM outputs in a memory accessible by the user's client device. In some implementations, the automatic assistant can cause one or more LLMs to be indexed in the memory based on one or more terms included in a plurality of assistant queries. In additional or alternative implementations, the automatic assistant can generate corresponding embeddings (e.g., word2vec embeddings, or another lower-dimensional representation) for each of the plurality of assistant queries, map each of the corresponding embeddings into an assistant query embedding space, and index one or more LLM outputs. In additional or alternative implementations, the automatic assistant can cause one or more LLMs to be indexed in the memory based on one or more context signals included in the corresponding previous context. In additional or alternative implementations, the automatic assistant can generate corresponding embeddings for each of the corresponding contexts, map each of the corresponding embeddings into a context embedding space, and index one or more LLM outputs. In additional or alternative implementations, the automatic assistant can cause one or more LLMs to be indexed in the memory based on one or more terms or phrases of an assistant output included in a set of assistant outputs for each of the plurality of assistant queries. In additional or alternative implementations, the automatic assistant can generate corresponding embeddings (e.g., word2vec embeddings, or another lower-dimensional representation) for each of the assistant outputs included in the set of assistant outputs, map each of the corresponding embeddings into an assistant output embedding space, and index one or more LLM outputs.

[0011] Accordingly, when speech continues to be received at the user's client device, the automatic assistant can identify one or more LLM outputs previously generated based on the current assistant query corresponding to one or more of the assistant queries included in the plurality of queries, the current context corresponding to one or more of the corresponding previous contexts, and / or one or more current assistant outputs corresponding to one or more of the previous assistant outputs. For example, in an implementation where one or more LLM outputs are indexed based on the corresponding embeddings for previous assistant queries, the automatic assistant can cause an embedding for the current assistant query to be generated and mapped to the assistant query embedding space. Further, the automatic assistant can determine that the current assistant query corresponds to a previous assistant query based on the distance between the embedding for the current assistant query and the corresponding embedding for the previous assistant query in the query embedding space satisfying a threshold. The automatic assistant can retrieve from memory one or more LLM outputs generated based on processing the previous assistant query and utilize the one or more LLM outputs when generating a set of modified assistant outputs. Also, for example, in an implementation where one or more LLMs are indexed based on one or more terms included in the plurality of assistant queries, the automatic assistant can determine, for example, the edit distance between the current assistant query and a plurality of previous assistant queries to identify a previous assistant query corresponding to the current assistant query. Similarly, the automatic assistant can retrieve from memory one or more LLM outputs generated based on processing the previous assistant query and utilize the one or more LLM outputs when generating a set of modified assistant outputs.

[0012] In some implementations, in addition to one or more LLM outputs, additional assistant queries can be generated based on processing the assistant query and / or the context of the conversation session. For example, when processing the assistant query and / or the context of the conversation session, the automated assistant can determine the intent associated with a given assistant query based on a stream of NLU data. Further, the automated assistant can, based on the intent associated with a given assistant query (e.g., based on a mapping of the intent to at least one related intent in a database or memory accessible by the client device and / or based on processing the intent associated with a given assistant query using rules defined in one or more machine learning (ML) models or heuristics), identify at least one related intent associated with the intent associated with the assistant query. Moreover, the automated assistant can generate an additional assistant query based on the at least one related intent. For example, assume the assistant query indicates that the user plans to go to the beach (e.g., "Hey assistant, I'm going to the beach today"). In this example, the additional assistant query can correspond to, for example, "what's the weather at Example Beach?" (e.g., to proactively determine the weather information at the beach that the user with the name Example Beach usually visits). In particular, the additional assistant query may not be provided for presentation to the user of the client device.

[0013] Rather, in these implementations, additional assistant outputs can be determined based on having processed additional assistant queries. For example, an automated assistant can send a structured request to one or more 1P and / or 3P systems to obtain weather information as an additional assistant output. Further, assume that the weather information indicates that rain is predicted at Example Beach. In some versions of these implementations, the automated assistant can further cause additional assistants to be processed using the LLM output and / or one or more of one or more additional LLM outputs to generate an additional set of modified assistant outputs. Thus, in the initial example given above, a given modified assistant output provided to the user in response to receiving the utterance "Hey Assistant, I'm thinking about going surfing today" from the initial set of modified assistant outputs could be "Enjoy it", and a given additional modified assistant output from the additional set of modified assistant outputs could be "but if you're going to Example Beach again, be prepared for some light showers". In other words, the automated assistant

[0014] In various implementations, when generating a set of modified assistant outputs, each of one or more LLM outputs utilized can be generated using a corresponding set of parameters out of a plurality of separate sets of parameters. Each of the plurality of separate sets of parameters can be associated with a separate personality for the auto-assistant. In some versions of those implementations, a single LLM can be utilized to generate one or more corresponding LLM outputs using the corresponding sets of parameters for each of the separate personalities, while in other versions of those implementations, multiple LLMs can be utilized to generate one or more corresponding LLM outputs using the corresponding sets of parameters for each of the separate personalities. Thus, from the set of modified assistant outputs, when a given modified assistant output is provided for presentation to the user, it can reflect various changing contextual personalities through the prosodic properties of the different personalities (e.g., intonation, pitch, tone, pauses, tempo, stress, rhythm, etc. of these different personalities).

[0015] In particular, the responses of these personalities described herein can reflect not only the rhythmic nature of different personalities, but also the distinct vocabulary of different personalities and / or the distinct ways of speaking of different personalities (e.g., verbose, concise, etc.). For example, a given modified assistant output provided for presentation to a user can be generated using a first set of parameters that reflect a first personality of the automated assistant, with respect to a first vocabulary that would be utilized by the automated assistant and / or a first set of rhythmic properties that would be utilized in rendering the modified assistant output for audible presentation to the user. Alternatively, a given modified assistant output provided for presentation to a user can be generated using a second set of parameters that reflect a second personality of the automated assistant, with respect to a second vocabulary that would be utilized by the automated assistant and / or a second set of rhythmic properties that would be utilized in rendering the modified assistant output for audible presentation to the user.

[0016] Accordingly, the automated assistant can dynamically adapt the personality utilized in providing a modified assistant output for presentation to the user, based on both the vocabulary utilized by the automated assistant and the rhythmic properties utilized in rendering the modified assistant output for audible presentation to the user. In particular, the automated assistant can dynamically adapt these personalities utilized in providing a modified assistant output based on the context of the conversation session, including previous utterances received from the user, as well as previous assistant outputs provided by the automated assistant and / or any other context signals described herein. As a result, the modified assistant output provided by the automated assistant may resonate more with the user of the client device. Note also that the personality utilized throughout a given conversation session can be dynamically adapted as the context of the given conversation session is updated.

[0017] In some implementations, the automatic assistant may rank assistant outputs included in a set of assistant outputs (i.e., not generated using one or more LLM outputs) and a set of modified assistant outputs (i.e., generated using one or more LLM outputs) according to one or more ranking criteria. Thus, when selecting a given assistant output to provide to the user, the automatic assistant can select from both the set of assistant outputs and the set of modified assistant outputs. The one or more ranking criteria can include, for example, one or more predicted metrics indicating how responsive each of the assistant outputs included in the set of assistant outputs and the set of modified assistant outputs is predicted to be to an assistant query included in the utterance (e.g., an ASR metric generated when generating a stream of ASR outputs, an NLU metric generated when generating a stream of NLU outputs, a fulfillment metric generated when generating the set of assistant outputs), one or more intents included in the stream of NLU outputs, and / or other ranking criteria. For example, if the intent of the user of the client device indicates that the user desires an answer regarding a fact (e.g., based on providing an utterance including an assistant query such as "why is the sky blue?"), the user is likely to desire a simple answer to the assistant query, so the automatic assistant can prioritize one or more of the assistant outputs included in the set of one or more assistant outputs. However, if the intent of the user of the client device indicates that the user has provided a free-form input (e.g., based on providing an utterance including an assistant query such as "what time is it?"), the user is likely to prefer a more conversational aspect, so the automatic assistant can prioritize one or more of the assistant outputs included in the set of modified assistant outputs.

[0018] In some implementations, before generating a set of modified assistant outputs, the automatic assistant may even determine whether to generate a set of modified assistant outputs. In some versions of those implementations, when providing an utterance as indicated by a stream of NLU data, the automatic assistant may even determine whether to generate a set of modified assistant outputs based on one or more of the user's predicted intents. For example, in an implementation where an utterance requests the automatic assistant to perform a search (e.g., an assistant query such as "Why is the sky blue?"), the automatic assistant may determine not to generate a set of modified assistant outputs because the user is seeking an answer regarding a fact. In additional or alternative versions of those implementations, the automatic assistant may even determine whether to generate a set of modified assistant outputs based on one or more computational costs associated with modifying one or more of the assistant outputs. The one or more computational costs associated with modifying one or more of the assistant outputs may include, for example, one or more of battery consumption, processor consumption associated with modifying one or more of the assistant outputs, or latency associated with modifying one or more of the assistant outputs. For example, when the client device is in a low-power mode, the automatic assistant may determine not to generate a set of modified assistant outputs in order to reduce the battery consumption of the client device.

[0019] By using the techniques described herein, one or more technical advantages can be achieved. In one non-limiting example, the techniques described herein enable an automated assistant to engage in a natural conversation with a user during an interaction session. For example, the automated assistant can generate a modified assistant output using one or more LLM outputs of a more conversational nature. Thus, the automated assistant can proactively provide (e.g., by generating additional assistant queries as described herein and providing additional assistant outputs determined based on the additional assistant queries) context information related to the interaction session that was not directly requested by the user, thereby making the modified assistant output resonate with the user. Further, the modified assistant output may be generated with various personalities with respect to both the contextually adapted vocabulary throughout the interaction session and the prosodic properties utilized to aurally render the modified assistant output, thereby making the modified assistant output resonate more with the user. This provides various technical advantages that conserve computational resources on the client device, enable the interaction session to conclude in a more rapid and efficient manner, and / or reduce the amount of the interaction session. For example, by enabling contextually relevant information related to the interaction session to be proactively provided by the automated assistant for presentation to the user, the amount of situations in which the user has to request such information can be reduced, thereby reducing the amount of user input received at the client device. Also, for example, in implementations where one or more LLM outputs are generated in an offline manner and subsequently utilized in an online manner, latency can be reduced at runtime.

[0020] As used herein, an "interaction session" may include a logically self - contained exchange between a user and an automated assistant (and in some cases, other human participants). The automated assistant can distinguish multiple interaction sessions with the user based on various signals such as the passage of time between sessions, changes in the user's situation between sessions (e.g., location, before / during / after a scheduled meeting, etc.), detection of one or more interfering interactions between the user and the client device other than the interaction between the user and the automated assistant (e.g., the user switches applications for a while, the user moves away from and then returns to a product that operates by standalone voice), locking / sleeping of the client device between sessions, changes in the client device used to interface with the automated assistant, etc. In particular, during a given interaction session, the user can interact with the automated assistant using various input modalities including, but not limited to, spoken input, typed input, and / or touch input.

[0021] The above description is provided as an overview of only some of the implementations disclosed in this specification. Those implementations and other implementations are described in additional detail in this specification.

[0022] It should be understood that the techniques disclosed herein may be implemented locally on a client device, remotely by a server connected to the client device via one or more networks, and / or both.

Brief Description of the Drawings

[0023]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

[0024] Referring now to FIG. 1, a block diagram of an exemplary environment 100 is shown that illustrates various aspects of the present disclosure and in which implementations disclosed herein may be implemented. The exemplary environment 100 includes a client device 110 and a natural language conversation system 120. In some implementations, the natural language conversation system 120 may be implemented locally on the client device 110. In additional or alternative implementations, the natural language conversation system 120 may be implemented remotely from the client device 110 (e.g., on a remote server), as shown in FIG. 1. In these implementations, the client device 110 and the natural language conversation system 120 may be communicatively coupled to each other via one or more networks 199, such as one or more wired or wireless local area networks (LANs, including Wi-Fi LAN, mesh networks, Bluetooth, near field communication, etc.) or wide area networks (WANs, including the Internet).

[0025] The client device 110 can be one or more of, for example, a desktop computer, a laptop computer, a tablet, a mobile phone, a vehicle computing device (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a stand-alone interactive speaker (optionally having a display), a smart appliance such as a smart TV, and / or a user's wearable device including a computing device (e.g., a user's wristwatch having a computing device, a user's glasses having a computing device, a virtual or augmented reality computing device). Additional and / or alternative client devices may be provided.

[0026] Client device 110 can execute an auto assistant client 114. An example of the auto assistant client 114 may be an application separate from the operating system of the client device 110 (e.g., installed "on top of" the operating system), or alternatively, may be implemented directly by the operating system of the client device 110. The auto assistant client 114 can interact with a natural conversation system 120 that is implemented locally on the client device 110 or remotely implemented and invoked via one or more of the networks 199 as shown in FIG. 1. The auto assistant client 114 (and optionally by interaction with other remote systems (e.g., servers)) can form what appears to the user's perspective to be a logical instance of an auto assistant 115 with which the user can engage in a human-computer dialogue. Examples of the auto assistant 115 are shown in FIG. 1 and are surrounded by a dashed line including the auto assistant client 114 and the natural conversation system 120 of the client device 110. Thus, it should be understood that a user interacting with the auto assistant client 114 running on the client device 110 can substantially interact with a logical instance of the auto assistant 115 itself (or a logical instance of the auto assistant 115 shared among a household or other group of users). For brevity and simplicity, the auto assistant 115 as used herein refers to the auto assistant client 114 that is executed locally on the client device 110 and / or remotely executed on one or more remote servers that may implement the natural conversation system 120.

[0027] In various implementations, the client device 110 may include a user input engine 111 configured to detect user input provided by a user of the client device 110 using one or more user interface input devices. For example, the client device 110 may be equipped with one or more microphones configured to capture audio data, such as audio data corresponding to a user's speech or other sounds in the environment of the client device 110. Additionally or alternatively, the client device 110 may be equipped with one or more visual components configured to capture visual data corresponding to images and / or motion (e.g., gestures) detected in one or more fields of view of the visual components. Additionally or alternatively, the client device 110 may be equipped with one or more touch sensing components (e.g., keyboard and mouse, stylus, touch screen, touch panel, one or more hardware buttons, etc.) configured to capture signals corresponding to touch input directed to the client device 110.

[0028] In various implementations, the client device 110 may include a rendering engine 112 configured to provide content for audible and / or visual presentation to a user of the client device 110 using one or more user interface output devices. For example, the client device 110 may be equipped with one or more speakers configured to enable the provision of content for audible presentation to a user via the client device 110. Additionally or alternatively, the client device 110 may be equipped with a display or projector configured to enable the provision of content for visual presentation to a user via the client device 110.

[0029] In various implementations, client device 110 can include one or more presence sensors 113 configured to provide a signal indicating a detected presence, particularly a human presence, upon approval from the corresponding user. In some of those implementations, the automatic assistant 115 can identify the client device 110 (or another computing device associated with the user of the client device 110) for which an utterance should be satisfied, based at least in part on the presence of the user in the client device 110 (or another computing device associated with the user of the client device 110). By rendering response content in the client device 110 and / or another computing device associated with the user of the client device 110 (e.g., via the rendering engine 112), by causing the client device 110 and / or another computing device associated with the user of the client device 110 to be controlled, and / or by causing any other action to be performed in the client device 110 and / or another computing device associated with the user of the client device 110 to satisfy the utterance, the utterance can be satisfied. As described herein, the automatic assistant 115 can utilize data determined based on the presence sensors 113 when determining the client device 110 (or other computing device) and provide the corresponding command only to that client device 110 (or those other computing devices), based on where the user is or was recently nearby.In some additional or alternative implementations, the automatic assistant 115 can utilize data determined based on the presence sensor 113 when determining whether any user (any user or a specific user) is currently near the client device 110 (or other computing device), and can optionally suppress the provision of data to and / or from the client device 110 (or other computing device) based on the user near the client device 110 (or other computing device).

[0030] The presence sensor 113 can be in various forms. For example, the client device 110 can utilize one or more of the user interface input components described above with respect to the user input engine 111 to detect the presence of a user. Additionally or alternatively, the client device 110 can be equipped with other types of light-based presence sensors 113, such as a passive infrared (PIR) sensor that measures infrared (IR) light emitted from objects within the field of view.

[0031] Additionally or alternatively, in some implementations, the presence sensor 113 can be configured to detect other phenomena related to the presence of a person or a device. For example, in some embodiments, the client device 110 can be equipped with a presence sensor 113 that detects various types of wireless signals (such as waves like radio waves, ultrasonic waves, electromagnetic waves, etc.) emitted by other computing devices (such as mobile devices, wearable computing devices) carried / operated by a user and / or other computing devices. For example, the client device 110 can be configured to emit waves that are imperceptible to humans, such as ultrasonic or infrared waves, which can be detected by other computing devices (such as via an ultrasonic / infrared receiver such as an ultrasonic microphone).

[0032] Additionally or alternatively, the client device 110 may emit other types of waves that are imperceptible to humans, such as radio waves (e.g., Wi-Fi, Bluetooth, cellular, etc.) that can be detected by other computing devices (e.g., mobile devices, wearable computing devices, etc.) carried / operated by the user and used to determine the user's specific location. In some implementations, GPS and / or Wi-Fi triangulation may be used to detect a person's location based on, for example, GPS and / or Wi-Fi signals to / from the client device 110. In other implementations, other wireless signal characteristics such as time of flight, signal strength, etc. may be used alone or collectively by the client device 110 to determine the location of a particular person based on signals emitted by other computing devices carried / operated by the user. Additionally or alternatively, in some implementations, the client device 110 may perform speaker identification (SID) to recognize the user from the user's voice and / or face identification (FID) to recognize the user from visual data capturing the user's face.

[0033] In some implementations, the movement of the speaker can then be determined, for example, by the presence sensor 113 of the client device 110 (and optionally the GPS sensor, Soli chip, and / or accelerometer of the client device 110). In some implementations, based on such detected movement, the user's location may be predicted, which may be assumed to be the user's location when any content is rendered in the client device 110 and / or other computing devices, at least in part based on the proximity of the client device 110 and / or other computing devices to the user's location. In some implementations, the user may simply be assumed to be at the last location where they interacted with the assistant 115, especially if not much time has elapsed since that last interaction.

[0034] Furthermore, client device 110 and / or natural conversation system 120 may include one or more memories for storing data and / or software applications, one or more processors for accessing the data and executing the software applications, and / or other components that facilitate communication via one or more of network 199. In some implementations, one or more of the software applications may be installed locally on client device 110, while in other implementations, one or more of the software applications may be hosted remotely (e.g., by one or more servers) and may be accessible by client device 110 via one or more of network 199.

[0035] In some implementations, the operations performed by automatic assistant 115 may be implemented locally on client device 110 via automatic assistant client 114. As shown in FIG. 1, automatic assistant client 114 may include an automatic speech recognition (ASR) engine 130A1, a natural language understanding (NLU) engine 140A1, a large language model (LLM) engine 150A1, and a text-to-speech (TTS) engine 160A1. In some implementations, the operations performed by automatic assistant 115 may be distributed across multiple computer systems, such as when natural conversation system 120 is implemented remotely from client device 110 as shown in FIG. 1. In these implementations, automatic assistant 115 may additionally or alternatively utilize the ASR engine 130A2, NLU engine 140A2, LLM engine 150A2, and TTS engine 160A2 of natural conversation system 120.

[0036] Each of these engines can be configured to perform one or more functions. For example, the ASR engines 130A1 and / or 130A2 can use a streaming ASR model stored in the machine learning (ML) model database 115A (e.g., a recurrent neural network (RNN) model, a transformer model, and / or any other type of ML model capable of performing ASR) to capture speech and process a stream of audio data generated by a microphone of the client device 110 to generate a stream of ASR output. In particular, the streaming ASR model can be utilized to generate a stream of ASR output as the stream of audio data is generated. Further, the NLU engines 140A1 and / or 140A2 can use an NLU model (e.g., long short-term memory (LSTM), gated recurrent unit (GRU), and / or any other type of RNN or other ML model capable of performing NLU) stored in the ML model database 115A and / or grammar-based rules to process the stream of ASR output. Moreover, the automatic assistant 115 can cause the NLU output to be processed to generate a stream of fulfillment data. For example, the automatic assistant 115 can send one or more structured requests to one or more first-party (1P) systems 191 via one or more of the network 199 (or one or more application programming interfaces (APIs)) and / or to one or more third-party (3P) systems 192 via one or more of the network and receive fulfillment data from one or more of the 1P systems 191 and / or 3P systems 192 to generate a stream of fulfillment data. The one or more structured requests can include, for example, NLU data included in the stream of fulfillment data.The stream of fulfillment data may correspond to a set of assistant outputs that are predicted to respond to assistant queries included in utterances captured in, for example, a stream of audio data processed by ASR engines 130A1 and / or 130A2.

[0037] Furthermore, the LLM engines 150A1 and / or 150A2 can process a set of assistant outputs that are predicted to respond to assistant queries included in utterances captured in a stream of audio data processed by the ASR engines 130A1 and / or 130A2. (For example, with respect to FIGS. 2 - 6) As described herein, in some implementations, the LLM engines 150A1 and / or 150A2 can use one or more LLM outputs to modify a set of assistant outputs to generate a modified set of assistant outputs. In some versions of those implementations (for example, as described with respect to FIG. 3), the automatic assistant 115 can be such that one or more LLM outputs are generated offline (for example, without responding to an utterance being received during a dialogue session) and subsequently utilized online (for example, when an utterance is received during a dialogue session) to generate a modified set of assistant outputs. In additional or alternative implementations of those implementations (for example, as described with respect to FIGS. 4 and 5), the automatic assistant 115 can be such that one or more LLM outputs are generated online (for example, when an utterance is received during a dialogue session). In these implementations, one or more LLM outputs can be generated based on processing a set of assistant outputs (for example, a stream of fulfillment data), the context of the dialogue session in which the utterance is received (for example, based on one or more context signals stored in the context database 110A), the recognized text corresponding to the assistant query included in the utterance, and / or other information that the automatic assistant 115 can utilize when generating one or more LLM outputs, using one or more LLMs (for example, one or more transformer models such as Meena, RNN, and / or any other LLM) stored in the model database 115A.Each of the one or more LLM outputs can include, for example, a probability distribution over a sequence of one or more words and / or phrases spanning one or more vocabularies, and one or more of the sequence of words and / or phrases can be selected as one or more LLM outputs based on the probability distribution. In various implementations, one or more of the LLM outputs can be stored in the LLM output database 150A for later use in modifying one or more of the assistant outputs included in the set of assistant outputs.

[0038] In addition, in some implementations, the TTS engines 160A1 and / or 160A2 use the TTS models stored in the ML model database 115A to process text data (e.g., the text spoken by the assistant 115) to generate synthetic audio data including computer-generated synthetic speech. The text data can correspond to, for example, one or more assistant outputs from a set of assistant outputs included in a stream of fulfillment data, one or more modified assistant outputs from a set of modified assistant outputs, and / or any other text data described herein. In particular, the ML models stored in the ML model database 115A may be on-device ML models stored locally in the client device 110, or may be shared ML models accessible to both the client device 110 and / or a remote system when the natural conversation system 120 is not implemented locally in the client device 110. In additional or alternative implementations, the assistant does not need to use the TTS engines 160A1 and / or 160A2 to generate any synthetic audio data so that the audio data is provided for audible presentation to the user, since audio data corresponding to one or more assistant outputs from a set of assistant outputs included in a stream of fulfillment data, one or more modified assistant outputs from a set of modified assistant outputs, and / or any other text data described herein can be stored in a memory or one or more databases accessible by the client device 110.

[0039] In various implementations, the stream of ASR output can include, for example, a stream of ASR hypotheses (e.g., term hypotheses and / or transcription hypotheses) predicted to correspond to a user's utterance captured in a stream of audio data, one or more corresponding predicted values (e.g., probabilities, log-likelihoods, and / or other values) for each of the ASR hypotheses, a plurality of phonemes predicted to correspond to the user's utterance captured in the stream of audio data, and / or other ASR output. In some versions of those implementations, the ASR engines 130A1 and / or 130A2 can select one or more of the ASR hypotheses as the recognized text corresponding to the utterance (e.g., based on the corresponding predicted values).

[0040] In various implementations, the stream of NLU outputs can include a stream of annotated recognized text that includes one or more (e.g., all) annotations of the recognized text for one or more of the terms of the recognized text. For example, NLU engines 140A1 and / or 140A2 can include part of a voice tagger (not shown) configured to annotate terms with their grammatical role in the term. Additionally or alternatively, NLU engines 140A1 and / or 140A2 can include an entity tagger (not shown) configured to annotate entity references in one or more segments of the recognized text, such as references to people (including, e.g., literary characters, celebrities, public figures, etc.), organizations, locations (real and fictional), etc. In some implementations, data about entities can be stored in one or more databases, such as in a known graph (not shown). In some implementations, the known graph can include nodes representing known entities (and optionally, entity attributes), as well as edges connecting the nodes to represent relationships between the entities. The entity tagger can annotate references to entities at a high level of granularity (e.g., to enable identification of all references to entity classes such as people) and / or at a low level of granularity (e.g., to enable identification of all references to a particular entity such as a particular person). The entity tagger may depend on the content of the natural language input to resolve a particular entity and / or, optionally, may optionally communicate with a known graph or other entity database to resolve a particular entity.

[0041] Additionally or alternatively, the NLU engines 140A1 and / or 140A2 may include a coreference resolver (not shown) configured to group or "cluster" references to the same entity based on one or more context cues. For example, the coreference resolver may be used to resolve the term "them" in the input "buy them" to "buy theatre tickets" based on the fact that "theatre tickets" is mentioned in a client device notification rendered immediately prior to receiving the natural language input "buy them". In some implementations, one or more components of the NLU engines 140A1 and / or 140A2 may depend on annotations from one or more other components of the NLU engines 140A1 and / or 140A2. For example, in some implementations, the entity tagger may depend on annotations from the coreference resolver when annotating all references to a particular entity. Also, for example, in some implementations, the coreference resolver may depend on annotations from the entity tagger when clustering references to the same entity.

[0042] Although FIG. 1 is described with respect to a single client device having a single user, this is for illustration purposes only and it should be understood that no limitation is intended. For example, one or more additional client devices of the user may also implement the techniques described herein. For example, client device 110, one or more additional client devices, and / or any other computing device of the user may form an ecosystem of devices that can utilize the techniques described herein. These additional client devices and / or computing devices may communicate with client device 110 (e.g., via network 199). As another example, a given client device may be utilized by multiple users in a shared setting (e.g., a user group, household).

[0043] As described herein, the automatic assistant 115 can determine whether to modify a set of assistant responses using one or more of the LLM outputs and / or determine one or more sets of assistant outputs modified based on one or more of the LLM outputs. The automatic assistant 115 can make these determinations using the natural conversation system 120. In various implementations, as shown in FIG. 1, the natural conversation system 120 can include, additionally or alternatively, an offline output correction engine 170, an online output correction engine 180, and / or a ranking engine 190. The offline output correction engine 170 can include, for example, an assistant activity engine 171 and an indexing engine 172. Further, the online output correction engine 180 can include, for example, an assistant query engine 181 and an assistant personality engine 182. These various engines of the natural conversation system 120 are described in more detail with respect to FIGS. 2 through 5.

[0044] Referring to FIG. 2 here, an exemplary process flow 200 for using an LLM to generate assistant output is shown. A stream of audio data 201 generated by one or more microphones of the client device 110 in FIG. 1 can be processed by the ASR engines 130A1 and / or 130A2 to generate a stream of ASR output 203. Further, the ASR output 203 can be processed by the NLU engines 140A1 and / or 140A2 to generate a stream of NLU output 204. In some implementations, the NLU engines 140A1 and / or 140A2 can process the context 202 of the conversation session between the user of the client device 110 and the automatic assistant 115, which is at least partially executed on the user's client device 110. In some versions of those implementations, the context 202 of the conversation session can be determined based on one or more context signals generated by the client device 110 (e.g., time, day of the week, location of the client device 110, ambient noise detected in the environment of the client device 110, and / or other context signals generated by the client device 110). In additional or alternative versions of those implementations, the context 202 of the conversation session can be determined based on one or more context signals stored in the context database 110A accessible on the client device 110 (e.g., user profile data, software application data, environmental data about the known environment of the user of the client device 110, the conversation history of the ongoing conversation session between the user and the automatic assistant 115 and / or the past conversation history of one or more previous conversation sessions between the user and the automatic assistant 115, and / or other context data stored in the context database 110A).In addition, the stream of NLU outputs 204 may be processed by one or more of the 1P system 191 and / or the 3P system to generate a stream of fulfillment data that includes a set of one or more assistant outputs 205, and each of the one or more assistant outputs included in the set of one or more assistant outputs 205 is predicted to respond to an utterance captured in the stream of audio data 201.

[0045] Typically, in an order-based dialogue session that does not utilize an LLM, the ranking engine 190 may process a set of one or more assistant outputs 205 to rank each of the one or more assistant outputs included in the set of one or more assistant outputs 205 according to one or more ranking criteria, and its automated assistant 115 may select one or more given assistant outputs 207 from the set of one or more assistant outputs 205 to be provided for presentation to the user of the client device 110 in response to receiving the utterance. In some implementations, the selected one or more given assistant outputs 207 may be processed by the TTS engines 160A1 and / or 160A2 to generate synthetic audio data including a synthetic voice corresponding to the selected one or more given assistant outputs 207, and the rendering engine 112 may cause the synthetic audio data to be audibly rendered by the speaker of the client device 110 for audible presentation to the user of the client device 110. In additional or alternative implementations, the rendering engine 112 may cause text data corresponding to the selected one or more given assistant outputs 207 to be visually rendered by the display of the client device 110 for visual presentation to the user of the client device 110.

[0046] However, when using the claimed technique, the automatic assistant 115 can further cause a set of one or more assistant outputs 205 to be processed by the LLM engines 150A1 and / or 150A2 to generate a set of one or more modified assistant outputs 206. In some implementations, one or more LLM outputs may be generated offline (e.g., before receiving a stream of audio data 201 generated by one or more of the microphones of the client device 110) using an offline output modification engine 170, and one or more LLM outputs may be stored in the LLM output database 150A. As described with respect to FIG. 3, one or more LLM outputs may be pre-indexed in the LLM output database 150A based on the corresponding assistant query and / or the corresponding context of the corresponding dialogue session in which the corresponding assistant query was received. Further, the LLM engines 150A1 and / or 150A2 can determine that an assistant query included in an utterance captured in the stream of audio data 201 matches the corresponding assistant query and / or that the context 202 of the dialogue session in which the assistant query is received matches the corresponding context of the corresponding dialogue session in which the corresponding assistant query was received. The LLM engines 150A1 and / or 150A2 can obtain one or more LLM outputs indexed by the corresponding assistant query and / or the corresponding context that matches the assistant query to modify the set of one or more assistant outputs 205. Moreover, the set of one or more assistant outputs 205 may be modified based on one or more of the LLM outputs, thereby obtaining a set of one or more modified assistant outputs 206.

[0047] In additional or alternative implementations, one or more LLM outputs can be generated online (e.g., in response to receiving a stream of audio data 201 generated by one or more microphones of the client device 110) using the online output correction engine 180. As described with respect to FIGS. 4 and 5, one or more LLM outputs are used to generate one or more LLM outputs using one or more LLMs stored in the ML model database 115A, a set of one or more assistant outputs 205, recognized text corresponding to an assistant query included in an utterance captured in the stream of audio data 201 (e.g., included in the stream of ASR output 203), and / or based on processing the context 202 of the conversation session between the user of the client device 110 and the automatic assistant 115. Further, the set of one or more assistant outputs 205 may be modified based on one or more of the LLM outputs, thereby obtaining a set of one or more modified assistant outputs 206. In other words, in these implementations, the LLM engines 150A1 and / or 150A2 can directly generate a set of one or more modified assistant outputs 206 using one or more of the LLMs stored in the ML model database 115A.

[0048] In these implementations, in contrast to the typical order-based dialogue sessions described without using an LLM, the ranking engine 190 ranks each of one or more assistant outputs included in both a set of one or more assistant outputs 205 and a set of one or more modified assistant outputs 206 according to one or more ranking criteria, and thus may process the set of one or more assistant outputs 205 and the set of one or more modified assistant outputs 206. Therefore, when selecting one or more given assistant outputs 207, the automatic assistant 207 can select from the set of one or more assistant outputs 205 and the set of one or more modified assistant outputs 206. In particular, the assistant outputs included in the set of one or more modified assistant outputs 206 are generated based on the set of one or more assistant outputs 205 and may convey the same or similar information, but also convey the same or similar information along with additional information related to the dialogue context 202 (such as described with respect to FIG. 4), and / or more natural and fluent, and / or more in line with the personality of the automatic assistant, so that one or more given assistant outputs 207 are more likely to resonate with the user of the client device 110.

[0049] One or more ranking criteria may include, for example, one or more predicted metrics indicating how responsive each of the assistant outputs included in a set of one or more assistant outputs 205 and a set of one or more modified assistant outputs 206 is predicted to be to an assistant query included in an utterance captured in a stream of audio data 201 (e.g., an ASR metric generated by ASR engines 130A1 and / or 130A2 when generating a stream of ASR output 203, an NLU metric generated by NLU engines 140A1 and / or 140A2 when generating a stream of NLU output 204, a fulfillment metric generated by one or more of 1P system 191 and / or 3P system 192), one or more intentions included in a stream of NLU output 204, and a metric derived from a classifier that processes each of the assistant outputs included in a set of one or more assistant outputs 205 and a set of one or more modified assistant outputs 206 to determine how natural, fluent, and / or in character with the automated assistant each of the assistant outputs is when provided for presentation to the user, and / or other ranking criteria. For example, if the user's intention on client device 110 indicates that the user desires an answer regarding a fact (e.g., based on providing an utterance including an assistant query such as "why is the sky blue?"), the user is likely to desire a simple answer to the assistant query, and thus the ranking engine 190 can re-use one or more of the assistant outputs included in the set of one or more assistant outputs 205. However, if the user's intention on client device 110 indicates that the user provided a free-form input (e.g., based on providing an utterance including an assistant query such as "what time is it?"), the user is likely to prefer a more conversational aspect, and thus the ranking engine 190 can re-use one or more of the assistant outputs included in the set of one or more modified assistant outputs 206.

[0050] Although FIGS. 1 and 2 are described herein with respect to voice-based dialogue sessions, this is for illustration purposes and is not intended to be limiting. Rather, it should be understood that the techniques described herein can be utilized regardless of the user's input modality. For example, in some implementations where the user provides input typed as an assistant query and / or touch input, the automatic assistant 115 can process the typed input using the NLU engines 140A1 and / or 140A2 to generate a stream of NLU outputs 204 (e.g., skipping the processing of the stream of audio data 201), and the LLM engines 150A1 and / or 150A2 can utilize the text input corresponding to the assistant query (e.g., derived from the typed input and / or touch input) in the same or a similar manner as described above to generate a set of one or more modified assistant outputs 206.

[0051] Referring now to FIG. 3, a flowchart illustrating an exemplary method 300 of utilizing a large language model to generate assistant output in an offline manner for later use in an online manner is shown. For convenience, the operations of method 300 are described with reference to a system that executes the operations from the process flow 200 of FIG. 2. This system of method 300 includes one or more processors, memories, and / or other components of a computing device (e.g., the client device 110 of FIG. 1, the client device 610 of FIG. 6, and / or the computing device 710 of FIG. 7, one or more servers, and / or other computing devices). Moreover, although the operations of method 300 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.

[0052] In block 352, the system obtains a plurality of assistant queries directed to the automatic assistant and the corresponding context of the corresponding previous dialogue session for each of the plurality of assistant queries. For example, the system can cause the assistant activity engine 171 of the offline output correction engine of FIGS. 1 and 2 to obtain a plurality of assistant queries and the corresponding context of the previous dialogue session in which the plurality of assistant queries were received, for example, from the assistant activity database 170A shown in FIG. 1. In some implementations, the plurality of assistant queries and the corresponding context of the previous dialogue session in which the plurality of assistant queries were received may be limited to those related to the user of the client device (e.g., the user of client device 110 in FIG. 1). In other implementations, the plurality of assistant queries and the corresponding context of the previous dialogue session in which the plurality of assistant queries were received may be limited to those related to multiple users of each client device (which may or may not include, for example, the user of client device 110 in FIG. 1).

[0053] In block 354, the system processes a given assistant query of a plurality of assistant queries using one or more LLMs to generate one or more corresponding LLM outputs, and each of the one or more corresponding LLM outputs is predicted to respond to the given assistant query. Each of the one or more corresponding LLM outputs can include, for example, a probability distribution over a sequence of one or more words and / or phrases spanning one or more vocabularies, and one or more of the sequence of words and / or phrases can be selected as one or more corresponding LLM outputs based on the probability distribution. In various implementations, when generating one or more corresponding LLM outputs for a given assistant query, the system uses one or more of the LLMs to process, along with the assistant query, the corresponding context of the corresponding previous dialogue session in which the given assistant query was received, and / or a set of assistant outputs predicted to respond to the given assistant query (e.g., generated based on processing audio data corresponding to the given assistant query using one or more of ASR engines 130A1 and / or 130A2, NLU engines 140A1 and / or 140A2, and 1P system 191 and / or 3P system 192 as described with respect to FIG. 2). In some implementations, the system can process the recognized text corresponding to the given assistant query, and in additional or alternative implementations, the system can process audio data capturing the utterance including the given assistant query. In some implementations, the system can cause the LLM engine 150A1 to process the given assistant query locally at the client device of the user (e.g., the user of client device 110 in FIG. 1) using one or more of the LLMs, while in other implementations, the system can cause the LLM engine 150A2 to process the given assistant query remotely from the client device of the user (e.g., at a remote server) using one or more of the LLMs.As described herein, one or more corresponding LLM outputs can reflect a more natural conversational output than typical assistant outputs that can be provided by an automated assistant, which enables the automated assistant to more smoothly lead the conversation session, so that assistant outputs modified based on one or more of the corresponding LLM outputs are likely to resonate with users who perceive the modified assistant outputs.

[0054] In some implementations, in addition to one or more corresponding LLM outputs, additional assistant queries can be generated using one or more of the LLM models based on a given assistant query and / or the corresponding context of the corresponding previous dialogue session in which the given assistant query was received. For example, when processing a given assistant query and / or the corresponding context of the corresponding previous dialogue session in which the given assistant query was received, one or more of the LLMs can determine the intent associated with the given assistant query (e.g., based on a stream of NLU outputs 204 generated using, for example, NLU engines 140A1 and / or 140A2 of FIG. 2). Further, one or more of the LLMs can identify at least one relevant intent related to the intent associated with the given assistant query based on the intent associated with the given assistant query (e.g., based on a mapping of the intent to at least one relevant intent in a database or memory accessible by the client device 110 and / or based on processing the intent associated with the given assistant query using rules defined by one or more machine learning (ML) models or heuristics). Moreover, one or more of the LLMs can generate an additional assistant query based on the at least one relevant intent. For example, assume that an assistant query indicates that the user has not yet had dinner (e.g., a given assistant query of "I'm feeling pretty hungry" received in the evening at the user's physical residence, as indicated by the corresponding context of the corresponding previous dialogue session related to the user's intent to indicate that they want to eat).In this example, additional assistant queries may correspond to, for example, "what types of cuisine has the user indicated he / she prefers?" (reflecting an intent for the types of relevant cuisine associated with the user's intent to indicate a desire to eat), "what restaurants nearby are open?" (reflecting an intent to search for relevant restaurants associated with the user's intent to indicate a desire to eat), and / or other additional assistant queries.

[0055] In these implementations, the additional assistant output may be determined based on processing the additional assistant queries. In the above example where the additional assistant query is "what types of cuisine has the user indicated he / she prefers?", user profile data from one or more of the 1P systems 191 stored locally on the client device 110 may be utilized to determine that the user has indicated a preference for Mediterranean and Indian cuisine. Based on the user profile data indicating the user's preference for Mediterranean and Indian cuisine, one or more corresponding LLM outputs may be modified to ask the user whether Mediterranean cuisine and / or Indian cuisine sounds appealing for dinner (e.g., "how does Mediterranean cuisine or Indian cuisine sound for dinner").

[0056] In the above example where the additional assistant query is "what restaurants nearby are open?", restaurant data from one or more of the 1P system 191 and / or 3P system 192 can be utilized (and optionally, limited to restaurants that serve Mediterranean and Indian cuisine based on user profile data) to determine which restaurants near the user's primary residence are open. Based on the results, one or more corresponding LLM outputs can be modified to provide the user with a list of one or more restaurants near the user's primary residence that are open (e.g., "Example Mediterranean Restaurant is open until 9:00 PM and Example Indian Restaurant is open until 10:00 PM"). In particular, additional assistant queries initially generated using the LLM (e.g., "what types of cuisine has the user indicated he / she prefers?" and "what types of cuisine has the user indicated he / she prefers?" in the above example) may not be included in one or more corresponding LLM outputs and, as a result, may not be provided for presentation to the user. Rather, additional assistant outputs determined based on the additional assistant queries (e.g., "how does Mediterranean cuisine or Indian cuisine sound for dinner" and "Example Mediterranean Restaurant is open until 9:00 PM and Example Indian Restaurant is open until 10:00 PM") may be included in one or more corresponding LLM outputs and, as a result, may be provided for presentation to the user.

[0057] In additional or alternative implementations, each of one or more corresponding LLM outputs (and optionally, additional assistant outputs determined based on additional assistant queries) can be generated using a corresponding set of parameters from a plurality of separate sets of one or more parameters of the LLM. Each of the plurality of separate sets of parameters can be associated with a separate personality for the auto-assistant. In some versions of those implementations, a single LLM can be utilized to generate one or more corresponding LLM outputs using the corresponding sets of parameters for each of the separate personalities, while in other versions of those implementations, multiple LLMs can be utilized to generate one or more corresponding LLM outputs using the corresponding sets of parameters for each of the separate personalities. For example, a first LLM output can be generated using a first set of parameters that reflect a first personality (e.g., the chef personality in the above example where a given assistant query corresponds to "I'm feeling pretty hungry"), a second LLM output can be generated using a second set of parameters that reflect a second personality (e.g., the butler personality in the above example where a given assistant query corresponds to "I'm feeling pretty hungry"), and a single LLM can be utilized to do the same for a plurality of other separate personalities. Also, for example, a first LLM can be utilized to generate a first LLM output using a first set of parameters that reflect a first personality (e.g., the chef personality in the above example where a given assistant query corresponds to "I'm feeling pretty hungry"), a second LLM can be utilized to generate a second LLM output using a second set of parameters that reflect a second personality (e.g., the butler personality in the above example where a given assistant query corresponds to "I'm feeling pretty hungry"), and the same can be true for a plurality of other separate personalities.Thus, when the corresponding LLM output is provided for presentation to the user, it can reflect various changing contextual personalities through the rhythmic nature of different personalities (e.g., the intonation, pitch, tone, pauses, tempo, stress, rhythm, etc. of these different personalities). Additionally or alternatively, the user can define one or more personalities to be utilized by the auto - assistant in a consistent manner (e.g., always using the butler personality) and / or in a context - dependent manner (e.g., using the butler personality in the morning and evening but a different personality during the day), (e.g., via the settings of the auto - assistant application related to the auto - assistant described herein).

[0058] In particular, the responses of these personalities described herein can reflect not only the rhythmic nature of different personalities, but also the vocabulary of different personalities and / or the distinct ways of speaking of different personalities (e.g., verbose ways of speaking, concise ways of speaking, friendly personalities, sarcastic personalities, etc.). For example, since the personality of the chef described above may have the vocabulary of a particular chef, the probability distribution over one or more words and / or phrases for one or more corresponding LLM outputs generated using a set of parameters for the chef personality can re - weight the set of words and / or phrases used by the chef more than other sets of words and / or phrases for other personalities (e.g., the personality of a scientist, the personality of a librarian). Thus, when one or more of the corresponding LLM outputs are provided for presentation to the user, it can reflect the personality according to various changing contexts not only with respect to the rhythmic nature of different personalities, but also with respect to the accurate and realistic vocabulary of different personalities, so that one or more of the corresponding LLM outputs will sound right to the user in various context scenarios. Moreover, it should be understood that the vocabulary and / or ways of speaking of different personalities can be defined with varying degrees of granularity. Continuing with the above example, the personality of the chef described above may have the unique vocabulary of a Mediterranean chef when asked about Mediterranean cuisine based on additional assistant queries being related to Mediterranean cuisine, and may have the unique vocabulary of an Indian chef when asked about Indian cuisine based on additional assistant queries being related to Indian cuisine, and so on.

[0059] In block 356, the system indexes one or more corresponding LLM outputs in a memory accessible in the client device (e.g., the LLM output database 150A of FIG. 1) based on a given assistant query and / or the corresponding context of the corresponding previous dialogue session for the given assistant query. For example, the system can cause the indexing engine 172 of the offline output correction engines of FIGS. 1 and 2 to index one or more corresponding LLM outputs in the LLM output database 150A. In some implementations, the indexing engine 172 can index one or more corresponding LLM outputs based on one or more terms included in a given assistant query and / or one or more context signals included in the corresponding context of the corresponding previous dialogue session in which the given assistant query was received. In additional or alternative implementations, the indexing engine 172 can cause an embedding (e.g., a word2vec embedding or any other low-dimensional representation) of the given assistant query to be generated and / or an embedding of one or more context signals included in the corresponding context of the corresponding previous dialogue session in which the given assistant query was received to be generated, and one or more of these embeddings can be mapped to an embedding space (e.g., a low-dimensional space). In these implementations, one or more corresponding LLM outputs indexed in the LLM output database 150A can be later utilized by the online output correction engine 180 to modify a set of assistant outputs (as described below with respect to blocks 362, 364, and 366 of method 300 of FIG. 3).In various implementations, the system can, additionally or alternatively, generate one or more assistant outputs (e.g., one or more assistant outputs 205 as described with respect to FIG. 2) that are predicted to respond to a given assistant query (but not generated using the LLM engines 150A1 and / or 150A2), and one or more corresponding LLM outputs can, additionally or alternatively, be indexed by the one or more assistant outputs.

[0060] In some implementations, as shown at block 358, the system can optionally receive user input for evaluating and / or modifying one or more of the corresponding LLM outputs. For example, a human evaluator can analyze one or more corresponding LLM outputs generated using one or more LLM models and modify one or more of the corresponding LLM outputs by changing one or more of the terms and / or phrases included in the one or more corresponding LLM outputs. Also, for example, a human evaluator can re-index, discard, and / or otherwise modify the indices of one or more of the corresponding LLM outputs. Thus, in these implementations, one or more of the corresponding LLM outputs generated using one or more LLMs can be selected by a human evaluator to ensure the quality of the one or more corresponding LLM outputs. Moreover, any non-discarded, re-indexed, and / or selected LLM outputs can be utilized to modify or retrain the LLM in an offline manner.

[0061] In block 360, the system determines whether there are additional assistant queries included in the plurality of assistant queries obtained in block 352 that have not been processed using one or more of the LLMs. In an iteration of block 360, if the system determines that there are additional assistant queries included in the plurality of assistant queries obtained in block 352 that have not been processed using one or more of the LLMs, the system returns to block 354 and performs additional iterations of blocks 354 and 356 with respect to the additional assistants rather than the given assistant query. These operations can be repeated for each of the assistant queries included in the plurality of assistant queries obtained in block 352. In other words, the system can index the assistant query and / or one or more corresponding LLM outputs for each of the corresponding contexts of the corresponding previous conversation sessions in which the corresponding one of the plurality of assistant queries is received, before those one or more corresponding LLM outputs are made available.

[0062] In the iteration of block 360, if the system determines that there are no additional assistant queries included in the plurality of assistant queries obtained in block 352 that have not been processed using one or more of the LLMs, the system may proceed to block 362. In block 362, the system can monitor a stream of audio data generated by one or more microphones of the client device to determine whether to capture the utterance of the user of the client device to which the stream of audio data is directed to the automatic assistant. For example, the system may monitor one or more specific words or phrases included in the stream of audio data (e.g., one or more specific words or phrases that call the automatic assistant using a hotword detection model). Also, for example, the system may optionally monitor the utterance directed to the client device in addition to one or more other signals (e.g., one or more gestures captured by the visual sensor of the client device, a gaze directed to the client device, etc.). In the iteration of block 362, if the system determines that it has not captured the utterance of the user of the client device to which the stream of audio data is directed to the automatic assistant, the system may continue to monitor the stream of audio data in block 362. In the iteration of block 362, if the system determines that it has captured the utterance of the user of the client device to which the stream of audio data is directed to the automatic assistant, the system may proceed to block 364.

[0063] In block 364, the system determines that the utterance includes the current assistant query corresponding to one of the plurality of assistant queries and / or that the utterance is received in the current context of the current dialogue session corresponding to the corresponding context of the corresponding previous dialogue session for one of the plurality of assistant queries, based on processing a stream of audio data. For example, the system can process a stream of audio data (e.g., the stream of audio data 201 in FIG. 2) using ASR engines 130A1 and / or 130A2 to generate a stream of ASR output (e.g., the stream of ASR output 203 in FIG. 2). Further, the system can process the stream of ASR output using NLU engines 140A1 and / or 140A2 to generate a stream of NLU output (e.g., the stream of NLU output 204 in FIG. 2). Moreover, based on the stream of ASR output and / or the stream of NUL output, the system can identify the current assistant query. In some implementations, the system can further determine a set of one or more assistant outputs (e.g., one or more assistant outputs 205 in FIG. 2) by causing the 1P system 191 and / or one or more of the 3P systems to process the stream of ASR output and / or the stream of NLU output.

[0064] In some implementations of method 300 of FIG. 3, the system can use the online output correction engine 180 to determine that one or more terms or phrases of the current assistant query correspond to one or more terms or phrases of one of the plurality of assistant queries for which one or more of the corresponding LLM outputs are indexed therefor (e.g., using any known technique for determining whether terms or phrases correspond to each other, such as exact matching techniques, soft matching techniques, edit distance techniques, phonetic similarity techniques, embedding techniques, etc.). In response to determining that one or more terms or phrases of the current assistant query correspond to one or more terms or phrases of one of the plurality of assistant queries, the system can obtain one or more of the corresponding LLM outputs that are indexed and associated with one of the plurality of assistant queries (e.g., from the LLM output database 150A) in the iteration of block 356. For example, if the current query includes the phrase "I'm hungry", the system can obtain one or more of the corresponding LLM outputs generated in the example described above for a given assistant query. For example, the system can determine that both the current query and a given assistant query described above include the phrase "I'm hungry" based on comparing the edit distance between the terms of the current query and the terms of the given assistant query. Also, for example, the system can generate an embedding of the current assistant query and map the embedding of the current query to the embedding space described above with respect to block 356. Further, the system can determine that the current assistant query corresponds to a given assistant query based on the distance between the generated embedding of the current query and the previously generated embedding for the given assistant query in the embedding space satisfying a distance threshold.

[0065] In an additional or alternative implementation of method 300 of FIG. 3, the system can use the online output modification engine 180 to determine that one or more context signals detected when the current assistant query is received correspond to one or more corresponding context signals when one of the plurality of assistant queries is received (e.g., received on the same day of the week, received at the same time, received at the same location, the same ambient noise exists in the environment of the client device, received in a specific series of utterances during the conversation session, etc.). In response to determining that one or more of the context signals associated with the current assistant query correspond to one or more of the context signals of one of the plurality of assistant queries, the system can obtain one or more of the corresponding LLM outputs (e.g., from the LLM output database 150A) that are indexed in the iteration of block 356 and associated with one of the plurality of assistant queries. For example, if the current query is received in the evening at the user's primary residence, the system can obtain one or more of the corresponding LLM outputs generated in the example described above for a given assistant query. For example, the system can determine that both the current query and a given assistant query described are associated with a temporal context signal of "evening" and a location context signal of "primary residence". Also, for example, the system can generate an embedding of one or more context signals associated with the current assistant query and map the embedding of one or more context signals associated with the current query to the embedding space described above with respect to block 356. Further, the system can determine that one or more context signals associated with the current assistant query correspond to one or more context signals associated with a given assistant query based on the distance between the generated embedding of one or more context signals associated with the current query and the previously generated embedding for one or more context signals associated with the given assistant query in the embedding space satisfying a threshold distance.

[0066] In particular, the system can leverage one or both of the current assistant query and the context of the dialogue session in which the current assistant query is received (e.g., one or more detected context signals) when determining one or more corresponding LLM outputs to be utilized in generating one or more current assistant outputs to be provided for presentation to the user in response to the current assistant query. In various implementations, the system can additionally or alternatively utilize one or more of the assistant outputs generated for the current assistant query when determining one or more corresponding LLM outputs to be utilized in generating one or more current assistant outputs to be provided for presentation to the user in response to the current assistant query. For example, the system can additionally or alternatively generate one or more assistant outputs (e.g., one or more assistant outputs 205 as described with respect to FIG. 2) that are predicted to respond to the current assistant query (but not generated using LLM engines 150A1 and / or 150A2), and it can be determined that one or more of the assistant outputs predicted to respond to the current assistant query correspond to one or more previously generated assistant outputs for one of a plurality of assistant queries using the various techniques described above.

[0067] In block 366, the system causes the automatic assistant to utilize one or more of the corresponding LLM outputs when generating one or more current assistant outputs to be provided for presentation to a user of a client device. For example, the system can rank one or more assistant outputs (e.g., one or more assistant outputs 205 as described with respect to FIG. 2), and one or more corresponding LLM outputs (e.g., one or more modified assistant outputs 206 as described with respect to FIG. 2) according to one or more ranking criteria. Further, the system can select one or more current assistant outputs from among the one or more assistant outputs and the one or more corresponding LLM outputs. Moreover, the system can cause the one or more current assistant outputs to be rendered visually and / or audibly for presentation to a user of the client device.

[0068] An implementation of method 300 of FIG. 3 is described as generating one or more corresponding LLM outputs offline (e.g., by using one or more of the LLM's to generate corresponding LLM outputs using an offline output modification engine 170 and indexing the one or more corresponding LLM outputs in an LLM output database 150A), and then utilizing the one or more corresponding LLM outputs online (e.g., by using an online output modification engine 180 to determine the LLM outputs to utilize from the LLM output database 150A based on a current assistant query), but this is for illustration and is not intended to be limiting. For example, as described below with respect to FIGS. 4 and 5, the online output modification engine 180 can, additionally or alternatively, in a preferred implementation, utilize the LLM engines 150A1 and / or 150A2 online.

[0069] Referring to FIG. 4 here, a flowchart showing an exemplary method 400 of utilizing a large language model when generating an assistant output based on generating an assistant query is shown. For convenience, the operations of method 400 are described with reference to a system that executes the operations from the process flow 200 of FIG. 2. This system of method 400 includes one or more processors, memories, and / or other components of a computing device (e.g., the client device 110 of FIG. 1, the client device 610 of FIG. 6, and / or the computing devices 710 of FIG. 7, one or more servers, and / or other computing devices). Moreover, although the operations of method 400 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.

[0070] In block 452, the system receives a stream of audio data that captures the user's utterance, the utterance includes an assistant query directed to the automatic assistant, and the utterance is received during a dialogue session between the user and the automatic assistant. In some implementations, the system may only process the stream of audio data to determine that the system captures the assistant query in response to determining that one or more conditions are met. For example, the system may monitor one or more specific words or phrases included in the stream of audio data (e.g., monitor one or more specific words or phrases that call the automatic assistant using a hotword detection model). Also, for example, the system may optionally monitor an utterance directed to the client device in addition to one or more other signals (e.g., one or more gestures captured by a visual sensor of the client device, a gaze directed to the client device, etc.).

[0071] In block 454, the system determines a set of assistant outputs based on processing a stream of audio data, and each assistant output included in the set responds to an assistant query included in the utterance. For example, the system can process a stream of audio data (e.g., a stream of audio data 201) using ASR engines 130A1 and / or 130A2 to generate a stream of ASR outputs (e.g., ASR output 203). Further, the system can process a stream of ASR outputs (e.g., a stream of audio data 201) using NLU engines 140A1 and / or 140A2 to generate a stream of NLU outputs (e.g., a stream of NLU output 204). Moreover, the system can cause one or more 1P systems 191 and / or 3P systems 192 to process a stream of NLU outputs (e.g., a stream of NLU output 204) to generate a set of assistant outputs (e.g., a set of assistant outputs) 205. In particular, the set of assistant outputs can correspond to one or more candidate assistant outputs that an automatic assistant might consider using in responding to the utterance without the techniques described herein (i.e., techniques that do not utilize LLM engines 150A1 and / or 150A2 when modifying assistant outputs as described herein).

[0072] In block 456, the system processes a set of assistant outputs and the context of the conversation session to (1) generate a set of assistant outputs modified using one or more LLM outputs, where each of the one or more LLM outputs is determined based at least on the context of the conversation session and / or one or more assistant outputs included in the set of assistant outputs, and (2) generate additional assistant queries related to the utterance based at least in part on the context of the conversation session and at least in part on the assistant query included in the utterance. In various implementations, each of the LLM outputs may be further determined based on the assistant query included in the utterance captured in the audio data stream. In some implementations, when generating the set of modified assistant outputs, one or more LLM outputs may have been generated beforehand in an offline manner (e.g., using the offline output modification engine 170 as described above with respect to FIG. 3, before receiving the utterance). In these implementations, the system can determine that the assistant query included in the utterance captured in the audio data stream corresponds to a previous assistant query for which one or more LLM outputs were previously generated therefor, that the context of the conversation session corresponds to the previous context of the previous conversation session in which the previous assistant query was received, and / or that one or more of the assistant outputs included in the set of assistant outputs determined in block 454 corresponds to one or more previous assistant outputs determined based on the previous query. Further, the system can obtain (e.g., using the online output modification engine 180) one or more LLM outputs indexed based on previous assistant queries, previous context, and / or one or more previous assistant outputs corresponding to the assistant query, context, and / or one or more of the assistant outputs included in the set of assistant outputs, respectively, as described with respect to method 300 of FIG. 3, and utilize the one or more LLM outputs as the set of modified assistant outputs.

[0073] In additional or alternative implementations, when generating a set of modified assistant outputs, the system can at least process the context of the conversation session and / or one or more assistant outputs included in the set of assistant outputs in an online manner (e.g., in response to receiving an utterance, using the online output modification engine 180) to generate one or more LLM outputs. For example, the system can cause the LLM engines 150A1 and / or 150A2 to process the context of the conversation session, the assistant query, and / or one or more assistant outputs included in the set of assistant outputs using one or more LLMs to generate a set of modified assistant outputs. The one or more LLM outputs can be generated in the same or a similar manner as described above with respect to block 354 of method 300 in FIG. 3 for generating one or more LLM outputs in an offline manner, but in response to an utterance being received at the client device, can be generated in an online manner. For example, the system can use one or more LLMs to process one or more assistant outputs included in the set of assistant outputs to generate one or more personality responses for each of the one or more assistant outputs included in the set of assistant outputs, as described above with respect to block 354 of method 300 in FIG. 3. In other words, each of the assistant outputs included in the set of assistant outputs can have a limited vocabulary and a consistent personality with respect to the rhythmic nature associated with each of the assistant outputs. However, when processing each of the assistant outputs included in the set of assistant outputs to generate a set of modified assistant outputs, each of the modified assistant outputs can have a much larger vocabulary depending on the use of one or more LLMs when generating the one or more modified assistant outputs, and at this time, the variations in the rhythmic nature associated with each of the modified assistant outputs are much larger.As a result, each of the modified assistant outputs can correspond to a contextually relevant assistant output that resonates with the user involved in the conversation session with the automated assistant.

[0074] Similarly, in some implementations, when generating additional assistant queries, additional assistants may be pre-generated in an offline manner (e.g., using the offline output modification engine 170 before receiving the utterance, as described above with respect to FIG. 3). In these implementations, similar to what was described above regarding obtaining one or more LLM outputs pre-generated in an offline manner, the system can, as described with respect to method 300 of FIG. 3, obtain additional assistant queries that are indexed based on previous assistant queries, previous context, and / or one or more of the assistant outputs included in the set of assistant outputs that correspond to each of the assistant queries, previous context, and / or one or more previous assistant outputs corresponding to the assistant query, and can utilize the previously generated additional assistant queries as the additional assistant queries.

[0075] Also, similarly, in additional or alternative implementations, when generating additional assistant queries, the system can process at least the context of the conversation session and / or one or more assistant outputs included in a set of assistant outputs in an online manner (e.g., in response to receiving an utterance, using the online output modification engine 180) to generate the additional assistant queries. For example, the system can cause the LLM engines 150A1 and / or 150A2 to process one or more assistant outputs included in a set of the context of the conversation session, assistant queries, and / or assistant outputs using one or more LLMs to generate the additional assistant queries. The additional assistant queries can be generated in an online manner, but in response to an utterance being received at the client device, in the same or a similar manner as described above with respect to block 354 of method 300 in FIG. 3 for generating additional assistant queries offline. In some implementations, one or more of the LLMs described herein can have multiple separate layers dedicated to performing some functions. For example, one or more first layers of one or more LLMs may be utilized when generating the personality responses described herein, and one or more second layers of one or more LLMs may be utilized when generating the additional assistant queries described herein. In additional or alternative implementations, one or more LLMs can communicate with one or more additional layers not included in the one or more LLMs when generating the additional assistant queries described herein. For example, one or more layers of one or more LLMs may be utilized when generating the personality responses described herein, and one or more additional layers of another ML model communicating with the one or more LLMs may be utilized when generating the additional assistant queries described herein. Non-limiting examples of personality responses and additional assistant queries are described in more detail below with respect to FIG. 6.

[0076] In block 458, the system determines additional assistant output responsive to an additional assistant query, based on the additional assistant query. In some implementations, the system may cause the additional assistant query to be processed by one or more of the 1P system 191 and / or the 3P system 192 in the same or a similar manner as described with respect to processing the assistant query in FIG. 2, to generate the additional assistant output. In some implementations, the additional assistant output may be a single additional assistant output, but in other implementations, the additional assistant output may be included in a set of additional assistant outputs (e.g., similar to the set of additional assistant outputs 205 in FIG. 2). In additional or alternative implementations, the additional assistant query may be directly mapped to the additional assistant output based on, for example, the user's user profile data that provided the utterance and / or any other data accessible by the automated assistant.

[0077] In block 460, the system processes a set of modified assistant outputs based on additional assistant outputs in response to additional assistant queries to generate an additional set of modified assistant outputs. In some implementations, the system may add additional assistant outputs at the beginning or end of one or more of the modified assistant outputs included in the set of modified assistant outputs generated in block 456 to each of the one or more assistant outputs. In additional or alternative implementations, as shown in block 460A, the system uses one or more additional LLM outputs generated based at least in part on the context of the conversation session in addition to the LLM output used in block 456 and / or one or more LLM outputs used in block 456, and at least in part on one or more additional assistant outputs, to process the additional assistant outputs and the context of the conversation session to generate an additional set of modified assistant outputs. The additional set of modified assistant outputs may be generated in the same or a similar manner as described above with respect to generating the set of modified assistant outputs, but based on the additional assistant outputs rather than the set of assistant outputs (e.g., using one or more LLM outputs generated in an offline manner and / or using LLM engines 150A1 and / or 150A2 in an online manner).

[0078] In block 462, the system causes a given modified assistant output from a set of modified assistant outputs and / or a given additional modified assistant output from a set of additional modified assistant outputs to be provided for presentation to the user. In some implementations, the system causes the ranking engine 190 to rank each of one or more modified assistant outputs included in the set of modified assistant outputs (and optionally each of one or more assistant outputs included in the set of assistant outputs) according to one or more ranking criteria and to select a given modified assistant output from the set of modified assistant outputs (or a given assistant output from the set of assistant outputs). Further, the system further causes the ranking engine 190 to rank each of one or more additional modified assistant outputs included in the set of additional modified assistant outputs (and optionally additional assistant outputs) according to one or more ranking criteria and to select a given additional modified assistant output from the set of additional modified assistant outputs (or a given additional assistant output as an additional assistant output). In these implementations, the system can combine a given modified assistant output and a given additional assistant output such that the given modified assistant output and the given additional assistant output are provided for visual and / or audible presentation to a user of a client device involved in an interaction session with the automated assistant.

[0079] Referring to FIG. 5 here, a flowchart showing an exemplary method 500 of utilizing a large language model when generating an assistant output based on generating an assistant personality response is shown. For convenience, the operations of method 500 are described with reference to a system that executes the operations from the process flow 200 of FIG. 2. This system of method 500 includes one or more processors, memories, and / or other components of a computing device (e.g., the client device 110 of FIG. 1, the client device 610 of FIG. 6, and / or the computing devices 710 of FIG. 7, one or more servers, and / or other computing devices). Moreover, although the operations of method 500 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.

[0080] In block 552, the system receives a stream of audio data that captures the user's utterance, the utterance includes an assistant query directed to the automatic assistant, and the utterance is received during a dialogue session between the user and the automatic assistant. In block 554, the system determines a set of assistant outputs based on processing the stream of audio data, and each of the assistant outputs included in the set responds to the assistant query included in the utterance. The operations of blocks 552 and 554 of method 500 of FIG. 5 may be performed in the same or similar manner as described with respect to blocks 452 and 454 of method 400 of FIG. 4, respectively.

[0081] In block 556, the system determines whether to modify one or more assistant outputs included in a set of assistant outputs. The system can determine whether to modify one or more assistant outputs based on, for example, the user's intent when providing an utterance (e.g., included in the stream of NLU output 204), one or more assistant outputs included in the set of assistant outputs (e.g., the set of assistant outputs 205), one or more computational costs associated with modifying one or more of the assistant outputs included in the set of assistant outputs (e.g., battery consumption, processor consumption, latency, etc.), the length of time interacting with the automated assistant, and / or other considerations. For example, if the user's intent indicates that the user providing the utterance anticipates a quick and / or fact-based response (e.g., "why is the sky blue?", "what's the weather?", "what time is it", etc.), in some cases the system may determine not to modify one or more of the assistant outputs in order to reduce the latency and consumption of computational resources when providing content in response to the utterance. Also, for example, if the user's client device is in a power-saving mode, the system may determine not to modify one or more of the assistant outputs in order to conserve battery power. Also, for example, if the user is engaged in a conversation with a session length that exceeds a threshold length of time (e.g., 30 seconds, 1 minute, etc.), the system may determine not to modify one or more of the assistant outputs in an attempt to end the conversation session in a more rapid and efficient manner.

[0082] In the iteration of block 556, if the system determines not to modify one or more of the assistant outputs included in the set of assistant outputs, the system may proceed to block 558. In block 558, the system causes a given assistant output from the set of assistant outputs to be provided for presentation to the user. For example, the system may cause the ranking engine 190 to rank each of the assistant outputs included in the set of assistant outputs according to one or more ranking criteria and, based on the ranking, select a given assistant output to be provided for visual and / or audible presentation to the user.

[0083] In the iteration of block 556, if the system determines to modify one or more of the assistant outputs included in the set of assistant outputs, the system may proceed to block 560. In block 560, the system processes the set of assistant outputs and the context of the dialogue session to generate a set of assistant outputs modified using one or more LLM outputs, where each of the one or more LLM outputs is determined based on the context of the dialogue session and / or one or more of the assistant outputs included in the set of assistant outputs, and each of the one or more LLM outputs reflects the corresponding personality of the automated assistant from among a plurality of distinct personalities. As described with respect to block 354 of method 300 of FIG. 3, the one or more LLM outputs may be generated using various distinct parameters to reflect different personalities of the automated assistant. In various implementations, each of the LLM outputs may be further determined based on an assistant query included in an utterance captured in a stream of audio data. In some implementations, the system may process the set of assistant outputs and the context of the dialogue session (and optionally the assistant query) to generate a set of assistant outputs modified using one or more LLM outputs pre-generated in an offline manner as described herein, but in additional or alternative implementations, the system may process the set of assistant outputs and the context of the dialogue session (and optionally the assistant query) to generate a set of assistant outputs modified in an online manner as described herein. As described above with respect to method 400 of FIG. 4, non-limiting examples of personality responses and additional assistant queries are described in more detail below with respect to FIG. 6.

[0084] In block 562, the system causes a given assistant output from a set of assistant outputs to be provided for presentation to the user. In some implementations, the system causes the ranking engine 190 to rank each of one or more modified assistant outputs included in the set of modified assistant outputs (and optionally, each of one or more assistant outputs included in the set of assistant outputs) according to one or more ranking criteria, and may cause a given modified assistant output (or a given assistant output from the set of assistant outputs) to be selected from the set of modified assistant outputs. Further, the system may cause a given modified assistant output to be provided for visual and / or audible presentation to the user.

[0085] Although FIG. 5 is not described with respect to generating additional assistant queries, it should be understood that this is for illustration and is not intended to be limiting. Rather, generating additional assistant queries in the method 400 of FIG. 4 can be based on determining that there are additional contextually relevant assistant outputs that can be provided to facilitate the dialogue session with respect to providing a more natural conversation experience for the user, whereby any assistant output provided for presentation to the user becomes more reflective of a person-to-person dialogue session and the dialogue session between the user and the automated assistant becomes more appealing to the user. Moreover, although FIGS. 3 and 4 are not described with respect to determining whether to modify the set of assistant outputs, it should also be understood that this is for illustration and is not intended to be limiting. Rather, it should be understood that determining whether to cause modification of the set of assistant responses using one or more LLMs can be performed in any of the exemplary methods of FIGS. 3, 4, and 5.

[0086] Referring now to FIG. 6, a non-limiting example of a dialogue session between a user and an automated assistant is shown where the automated assistant utilizes one or more LLMs to generate assistant output. As described herein, in some implementations, the automated assistant can utilize one or more pre-generated LLM outputs in an offline manner to generate a set of modified assistant outputs (e.g., as described above with respect to method 300 of FIG. 3). For example, the automated assistant can determine that one or more previous assistant queries for which one or more LLM outputs have been pre-generated therefor correspond to the assistant queries included in the utterance, that the previous context of a previous dialogue session corresponds to the context of the dialogue session between the user receiving the utterance and the automated assistant, and / or that one or more previous assistant outputs correspond to one or more of the assistant outputs included in the set of assistant outputs for the assistant queries included in the utterance. Further, the automated assistant can obtain one or more LLM outputs that are indexed (e.g., in LLM output database 150A) according to one or more of the previous assistant queries, the previous context, and / or the previous assistant outputs included in the previous set of assistant outputs, and utilize the one or more LLM outputs as the set of modified assistant outputs. In additional or alternative implementations, the automated assistant can cause one or more assistant queries, the context of the dialogue session, and / or one or more of the assistant outputs included in the set of assistant outputs to be processed using one or more LLMs to generate one or more LLM outputs to be utilized as the set of modified assistant outputs in an online manner (e.g., as described above with respect to method 400 of FIG. 4 and method 500 of FIG. 5). Thus, the non-limiting example of FIG. 6 is provided to show how the utilization of LLMs by the techniques described herein can result in an improved natural conversation between a user and an automated assistant.

[0087] The client device 610 (e.g., an example of the client device 110 in FIG. 1) may include various user interface components, such as a microphone for generating audio data based on, for example, speech and / or other audible inputs, a speaker for audibly rendering synthesized speech and / or other audible outputs, and / or a display 680 for visually rendering visual outputs. Further, the display 680 of the client device 610 may include various system interface elements 681, 682, and 683 (e.g., hardware and / or software interface elements) with which a user of the client device 610 can interact to cause the client device 610 to perform one or more actions. The display 680 of the client device 610 enables interaction between the user and the content rendered on the display 680 by touch input (e.g., by directing user input to the display 680 or a portion thereof (e.g., a text input box (not shown), a keyboard (not shown), or another portion of the display 680)) and / or by spoken input (e.g., by selecting the microphone interface element 684 or simply by speaking without necessarily selecting the microphone interface element 684 in the client device 610 (i.e., the automatic assistant enables one or more terms or phrases, gestures, gazes, mouth movements, lip movements, and / or other conditions for enabling spoken input)). It should be understood that the client device 610 shown in FIG. 6 is a mobile phone, but this is for illustration purposes and is not intended to be limiting. For example, the client device 610 may be a stand-alone speaker with a display, a stand-alone speaker without a display, a home automation device, an in-vehicle system, a laptop, a desktop computer, and / or any other device capable of running an automatic assistant to participate in a human-computer dialogue session with a user of the client device 610.

[0088] For example, assume that a user of the client device 610 provides an utterance 652 of "Hey Assistant, what time is it?". In this example, the automatic assistant can cause the audio data capturing the utterance 652 to be processed using the ASR engines 130A1 and / or 130A2 to generate a stream of ASR output. Further, the automatic assistant can cause the stream of ASR output to be processed using the NLU engines 140A1 and / or 140A2 to generate a stream of NLU output. Moreover, the automatic assistant can cause the stream of NLU output to be processed by one or more of the 1P system 191 and / or 3P system 192 to generate a set of one or more assistant outputs. The set of assistant outputs can include, for example, "8:30 AM", "Good morning, it's 8:30 AM", and / or any other output that conveys the current time to the user of the client device 610.

[0089] In the example of FIG. 6, the automatic assistant determines to modify one or more of the assistant outputs included in the set of assistant outputs to generate a set of modified assistant outputs. For example, the automatic assistant can determine that the assistant query included in utterance 652 requests the automatic assistant to provide the current time for presentation to the user. The automatic assistant can determine that there are one or more LLM outputs previously generated that request the automatic assistant to provide the current time for presentation to the user, that an example of a previous assistant query corresponding to the assistant query included in utterance 652 of FIG. 6, that the previous context of the previous dialogue session corresponds to the context of the dialogue session between the user and the automatic assistant in FIG. 6 (e.g., the user requests the automatic assistant to provide the current time in the morning (and optionally makes the request by starting a dialogue session), the client device 610 is located at a specific location, and / or other context signals), and / or that one or more of the previous assistant outputs correspond to one or more of the assistant outputs included in the set of assistant outputs for the assistant query included in utterance 652 of FIG. 6. Further, the automatic assistant can obtain one or more LLM outputs indexed according to one or more of the previous assistant queries, the previous context, and / or one or more of the previous assistant outputs included in the previous set of assistant outputs (e.g., in the LLM output database 150A) and utilize the one or more LLM outputs as the set of modified assistant outputs. Also, for example, the automatic assistant can cause one or more of the assistant queries, the context of the dialogue session, and / or one or more of the assistant outputs included in the set of assistant outputs to be processed using one or more LLMs to generate one or more LLM outputs to be utilized as the set of modified assistant outputs in an online manner.

[0090] In the example of FIG. 6, assume that the automated assistant determines to provide a modified assistant output 654 of "Good morning [User]! It's 8:30 AM. Any fun plans today?" for presentation to the user, and that the modified assistant output 654 is determined based on one or more of the LLM outputs. The modified assistant output 654 provided for presentation to the user is personalized or adapted to the user of the client device 610 and the context of the conversation session in that the modified assistant output 654 greets the user with a contextually appropriate greeting (e.g., "Good Morning") and addresses the user of the client device 610 by name (e.g., "[User]"). In particular, one or more of the LLM outputs may include one or more corresponding substitute terms (e.g., as indicated by "[User]!" in the modified assistant output 654) that may be populated with user profile data accessible to the automated assistant. The example of FIG. 6 includes a corresponding substitute term for the name of the user of the client device 610, but it should be understood that this is for illustration purposes and is not intended to be limiting. For example, one or more of the corresponding substitute terms may be populated with any data accessible to the automated assistant, such as a smart network connection device identifier (e.g., smart lighting, smart TV, smart appliance, smart speaker, smart door lock, etc.), a known location associated with the user of the client device 610 (e.g., city, state, county, region, area, country, workplace, the user's office, or the physical address of the user's primary residence of the client device 610), an entity criterion (e.g., a criterion of people, places, things, etc.), a software application accessible on the user's client device 610, and / or any other data accessible to the automated assistant.

[0091] Moreover, the modified assistant output 654 functions with respect to responding to the assistant query included in the utterance 652 (e.g., "It's 8:30 AM"). However, the modified assistant output 654 is not only personalized or adapted to the user and functions with respect to responding to the assistant query, but the modified assistant output 654 also helps to drive the conversation session between the user and the automated assistant by further engaging with the user in the conversation session (e.g., "Any fun plans today?"). Without using the techniques described herein with respect to modifying the originally generated set of assistant outputs based on processing the utterance 652, the automated assistant may not greet the user of the client device 610 (e.g., "Good morning"), may not address the user of the client device 610 by name (e.g., "[User]"), and may not further engage with the user of the client device 610 in the conversation session (e.g., "Any fun plans today?"), and may simply respond with "It's 8:30 AM". Thus, the modified assistant output 654 may resonate more with the user of the client device 610 than any of the assistant outputs included in the originally generated set of assistant outputs that do not utilize one or more LLM outputs.

[0092] In the example of FIG. 6, it is further assumed that the user of client device 610 provides utterance 656, “Yes, I'm thinking about going to the beach”. In this example, the automatic assistant can cause the audio data capturing utterance 656 to be processed to generate a set of assistant outputs that are generated without using one or more LLM outputs. Further, the automatic assistant can process the assistant queries, the set of assistant outputs, and / or the context of the conversation session included in utterance 656 to generate a set of modified assistant outputs (e.g., in an offline and / or online manner) that are determined using one or more LLM outputs and, optionally, to generate additional assistant queries based on the assistant queries.

[0093] In this example, the assistant outputs included in the set of assistant outputs (i.e., generated without using one or more LLM outputs) may be limited because the assistant queries included in utterance 656 do not request the automatic assistant to take any action. For example, the assistant outputs included in the set of assistant outputs may include “Sounds fun!”, “Surf's up!”, “That sounds like fun!”, and / or other assistant outputs that respond to utterance 656 but do not further engage with the user of client device 610 in the conversation session. In other words, the assistant outputs included in the set of assistant outputs may have limited vocabulary diversity because they are not generated using one or more LLM outputs as described herein. Nevertheless, the automatic assistant can utilize the assistant outputs included in the set of assistant outputs to determine how to modify one or more of the assistant outputs using one or more of the LLM outputs.

[0094] Furthermore, as described with respect to FIGS. 3 and 4, the automatic assistant can generate additional assistant queries based on the assistant queries and using one or more of the LLMs or a separate ML model communicating with one or more of the LLMs. For example, in the example of FIG. 6, the utterance 656 provided by the user of the client device 610 indicates that the user plans to go to the beach. Based on identifying that the utterance 656 indicates an intention associated with the user planning to go to the beach, the automatic assistant can determine an associated intention related to checking the weather at a beach that the user of the client device 610 frequently visits (e.g., an exemplary beach named "Half Moon Bay"). Based on identifying the associated intention, the automatic assistant generates an additional assistant query of "What's the weather?" along with the location parameter of "Half Moon Bay" and can query one or more of the 1P system 191 and / or the 3P system 192 to obtain an additional assistant output including "weather" of "Half Moon Bay" indicating, for example, that it rains all day and the temperature is low in "Half Moon Bay". In some implementations, the automatic assistant can cause additional assistant output and / or the context of the conversation session to be processed to generate a set of additional modified assistant outputs determined using one or more of the LLM outputs and / or one or more additional LLM outputs.

[0095] In the example of FIG. 6, the automatic assistant can rank the assistant outputs and the set of modified assistant outputs included in the set of assistant outputs according to one or more ranking criteria, and select one or more of the assistant outputs based on the ranking (for example, select a given assistant output such as "Sounds fun!"). Further, the automatic assistant can rank the assistant outputs and additional assistant outputs included in the set of additional modified assistant outputs according to one or more ranking criteria, and select one or more of the assistant outputs based on the ranking (for example, select a given additional assistant output such as "But if you're going to Half Moon Bay again, expect rain and chilly temps"). Moreover, the automatic assistant can combine the selected given assistant output and the selected given additional assistant output to yield a modified assistant output 658 such as "Sounds fun! But if you're going to Half Moon Bay again, expect rain and chilly temps", and the modified assistant output 658 can be provided for visual and / or audible presentation to the user of the client device 610. Thus, in this example, the automatic assistant can process the utterance 656 and provide additional context information related to the utterance (such as the weather at the beach where the user of the client device 610 is likely to visit) to further engage with the user of the client device 610 in the dialogue session.Without the techniques described herein, a user of client device 610 may be required to proactively request weather information from the automatic assistant, even though the automatic assistant is capable of determining and providing the weather information, thereby increasing the amount of user input, wasting computing resources at client device 610 in processing the increased amount of user input, and increasing the cognitive load on the user of client device 610.

[0096] In the example of FIG. 6, it is further assumed that the user of the client device 610 provides an utterance 660 of "Oh no... thanks for the heads up, can you remind me to check the weather again in two hours?". In this example, the automatic assistant can cause the audio data capturing the utterance 660 to be processed to generate a set of assistant outputs that are generated without using one or more LLM outputs. Further, the automatic assistant can cause the assistant query, the set of assistant outputs, and / or the context of the conversation session included in the utterance 660 to be processed to generate a set of modified assistant outputs (e.g., in an offline and / or online manner) that are determined using one or more LLM outputs. Based on processing the utterance 660, the automatic assistant can decide to set a reminder to remind the user of the client device 610 to check the weather in Half Moon Bay at 10:30 AM (e.g., two hours after the conversation session), or to proactively provide the user of the client device 610 with the weather in Half Moon Bay at 10:30 AM. Further, the automatic assistant can cause a modified assistant output 662 of "Sure thing, I set the reminder and hope the weather clears up for you" to be provided for visual and / or audible presentation to the user of the client device 610 from a set of modified assistant outputs (i.e., generated using one or more LLM outputs). In particular, the modified assistant output 662 in the example of FIG. 6 is contextual with respect to the conversation session in that it indicates to the user that it is hoped that the weather will clear up. In contrast, the assistant output included in the set of assistant outputs (i.e., generated without using one or more LLM outputs) may simply provide an indication that the reminder has been set without considering the context of the conversation session.

[0097] FIG. 6 is described with respect to using one or more LLM outputs to generate certain modified assistant outputs based on the context of a particular utterance and dialogue session and selecting certain modified assistant outputs to be provided for presentation to a user, but this is for illustration purposes and is not intended to be limiting. Rather, the techniques described herein can be utilized for any dialogue session between any user and the corresponding instance of an automated assistant. Further, a transcription corresponding to a dialogue session between a user and an automated assistant is shown on display 680 of client device 610, but this is for illustration purposes and it should also be understood that this is not intended to be limiting. For example, it should be understood that the dialogue session can be executed on any device capable of executing an automated assistant, whether or not the client device includes a display.

[0098] Moreover, it should be understood that the assistant output provided for presentation to the user in the dialog session of FIG. 6 may include various personality responses described herein. For example, the modified assistant output 654 may be generated using a first set of parameters that reflect a first personality of the automated assistant, with respect to a first vocabulary that will be utilized by the automated assistant and / or a first set of prosodic properties that will be utilized in providing the modified assistant output 654 for audible presentation to the user. Further, the modified assistant output 658 may be generated using a second set of parameters that reflect a second personality of the automated assistant, with respect to a second vocabulary that will be utilized by the automated assistant and / or a second set of prosodic properties that will be utilized in providing the modified assistant output 658 for audible presentation to the user. In this example, the first personality may reflect, for example, the personality of a butler or maid utilized to provide a morning greeting, respond to a user who requests the current time, and ask the user if they have any plans for the day. Further, the second personality may reflect, for example, the personality of a weather forecaster, surfer, or monitor (or a combination thereof) utilized to indicate that it might be fun to go to the beach, but that the weather that day may not be ideal at that beach. For example, since the user indicated that they plan to go to the beach, the surfer personality may be utilized to provide the "Sounds fun!" portion of the modified assistant output 658, and since the automated assistant is providing weather information to the user, the weather forecaster personality may be utilized to provide "But if you're going to Half Moon Bay again, expect rain and chilly temps".

[0099] Accordingly, the automatic assistant can dynamically adapt the personality used to provide the assistant output modified for presentation to the user based on both the vocabulary used by the automatic assistant and the prosodic properties used when rendering the assistant output modified for audible presentation to the user. In particular, the automatic assistant can dynamically adapt these personalities used to provide the modified assistant output based on the context of the conversation session, including previous utterances received from the user, as well as previous assistant outputs provided by the automatic assistant and / or any other context signals described herein. As a result, the modified assistant output provided by the automatic assistant may resonate more with the user of the client device.

[0100] Moreover, FIG. 6 is described herein in connection with a user providing utterances throughout a conversation session, which is for illustration purposes and is not intended to be limiting. For example, the user can, additionally or alternatively, provide typed input and / or touch input throughout the conversation session. In these implementations, the automatic assistant can process the typed input (e.g., using NLU engines 140A1 and / or 140A2) to generate a stream of NLU outputs (e.g., and optionally skipping any processing using ASR engines 130A1 and / or 130A2), and can process a stream of NLU data and text input corresponding to an assistant query derived from the typed input and / or touch input (e.g., using LLM engines 150A1 and / or 150A2) in the same or a similar manner as described above to generate a set of one or more modified assistant outputs.

[0101] Referring now to FIG. 7, a block diagram of an exemplary computing device 710 that may optionally be utilized to execute one or more aspects of the techniques described herein is shown. In some implementations, one or more of a client device, a cloud-based auto assistant component, and / or other components may include one or more components of the exemplary computing device 710.

[0102] The computing device 710 generally includes at least one processor 714 that communicates with several peripheral devices via a bus subsystem 712. These peripheral devices may include, for example, a storage subsystem 724 that includes a memory subsystem 725 and a file storage subsystem 726, a user interface output device 720, a user interface input device 722, and a network interface subsystem 716. The input and output devices enable user interaction with the computing device 710. The network interface subsystem 716 provides an interface to an external network and is coupled to a corresponding interface device in other computing devices.

[0103] The user interface input device 722 may include a keyboard, and pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated in a display, audio input devices such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 710 or into a communication network.

[0104] The user interface output device 720 may include a non-visual display such as a display subsystem, a printer, a fax machine, or an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 710 to the user, or to another machine or computing device.

[0105] The storage subsystem 724 stores programming and data structures that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 724 may include logic for performing selected aspects of the methods disclosed herein and for implementing the various components shown in FIGS. 1 and 2.

[0106] These software modules are generally executed by processor 714 alone or in combination with other processors. The memory 725 used in storage subsystem 724 may include several memories, including main random access memory (RAM) 730 for storing instructions and data during program execution, and read-only memory (ROM) 732 in which fixed instructions are stored. File storage subsystem 726 can provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of some implementations may be stored by file storage subsystem 726, in storage subsystem 724, or in other machines accessible by processor 714.

[0107] Bus subsystem 712 provides a mechanism for enabling the various components and subsystems of computing device 710 to communicate with each other as intended. Bus subsystem 712 is shown schematically as a single bus, although alternative implementations of bus subsystem 712 may use multiple buses.

[0108] Computing device 710 can be of different types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the constantly changing nature of computers and networks, the description of computing device 710 shown in FIG. 7 is intended only as an example for illustrating some implementations. Numerous other configurations of computing device 710 are possible that have more or fewer components than the computing device shown in FIG. 7.

[0109] In situations where the systems described in this specification collect or otherwise monitor personal information about a user, or may use personal and / or monitored information, the user may be provided with the opportunity to control whether a program or function collects user information (e.g., the user's social network, social actions or activities, occupation, user preferences, or the user's current geographical location), or to control whether and / or how the user receives content from a content server that may be relevant to the user. Also, some data may be handled in one or more ways before being stored or used so that information that can identify an individual is removed. For example, a user's identifying information may be handled so that it cannot be used to determine information that can identify an individual about the user, or a user's geographical location may be generalized, in which case the geographical location information is obtained (e.g., to the city, zip code, or state level) so that the specific geographical location of the user cannot be determined. Thus, the user can control how information is collected and / or used about the user.

[0110] In some implementations, a method implemented by one or more processors is provided as part of an interaction session between a user of a client device and an auto assistant implemented by the client device, the method including receiving a stream of audio data that captures the user's utterance, the stream of audio data being generated by one or more microphones of the client device and the utterance including an assistant query; determining a set of assistant outputs based on processing the stream of audio data, each assistant output in the set of assistant outputs responding to the assistant query included in the utterance; generating a set of assistant outputs modified using one or more LLM outputs generated using a large language model (LLM), each of the one or more LLM outputs being determined based on at least a portion of the context of the interaction session and one or more of the assistant outputs included in the set of assistant outputs; processing the set of assistant outputs and the context of the interaction session to generate additional assistant queries related to the utterance based on at least a portion of the context of the interaction session and at least a portion of the assistant query; determining additional assistant outputs in response to the additional assistant queries based on the additional assistant queries; processing the additional assistant outputs and the context of the interaction session to generate a set of additional modified assistant outputs using one or more of the LLM outputs or one or more additional LLM outputs generated using the LLM, each of the additional LLM outputs being determined based on the context of the interaction session and at least a portion of the additional assistant outputs; and providing a given modified assistant output from the set of modified assistant outputs and a given additional modified assistant output from the set of additional modified assistant outputs for presentation to the user.

[0111] These and other implementations of the technology disclosed herein are optional and may include one or more of the following features.

[0112] In some implementations, determining assistant output that responds to an assistant query included in an utterance based on processing a stream of audio data may include processing the stream of audio data using an ASR model to generate a stream of ASR output, processing the stream of ASR output using an NLU model to generate a stream of natural language understanding (NLU) data, and determining the set of assistant output based on the NLU stream.

[0113] In some versions of those implementations, processing a set of assistant outputs and the context of a conversation session to generate a set of assistant outputs modified using one or more of the LLM outputs generated using an LLM can include processing the set of assistant outputs and the context of the conversation session using the LLM to generate one or more of the LLM outputs and determining a set of assistant outputs modified based on one or more of the LLM outputs. In some further versions of those implementations, processing the set of assistant outputs and the context of the conversation session using the LLM to generate one or more of the LLM outputs can include processing the set of assistant outputs and the context of the conversation session using a first set of LLM parameters of a plurality of separate sets of LLM parameters to determine one or more of the LLM outputs having a first personality of a plurality of separate personalities. The set of modified assistant outputs can include one or more first personality assistant outputs reflecting the first personality. In yet further versions of those implementations, processing the set of assistant outputs and the context of the conversation session using the LLM to generate one or more of the LLM outputs can include processing the set of assistant outputs and the context of the conversation session using a second set of LLM parameters of a plurality of separate sets of LLM parameters to determine one or more of the LLM outputs having a second personality of a plurality of separate personalities. The set of modified assistant outputs may include one or more second personality assistant outputs reflecting the second personality, and the second personality may be distinct from the first personality.In yet further versions of those implementations, one or more first personality assistant outputs, included in a set of modified assistant outputs and reflecting a first personality, may be determined using a first vocabulary associated with the first personality, and one or more second personality assistant outputs, included in the set of modified assistant outputs and reflecting a second personality, may be determined using a second vocabulary associated with the second personality, where the second personality is distinct from the first personality based on the second vocabulary being distinct from the first vocabulary. In yet additional or alternative versions of those implementations, one or more first personality assistant outputs, included in the set of modified assistant outputs and reflecting a first personality, may be associated with a first set of prosodic properties utilized in providing a given modified assistant output for audible presentation to a user, and one or more second personality assistant outputs, included in the set of modified assistant outputs and reflecting a second personality, may be associated with a second set of prosodic properties utilized in providing a given modified assistant output for audible presentation to a user, where the second personality may be distinct from the first personality based on the second set of prosodic properties being distinct from the first set of prosodic properties.

[0114] In some versions of those implementations, to generate a set of assistant outputs modified using one or more of the LLM outputs generated using an LLM, processing the set of assistant outputs and the context of the dialogue session is based on one or more of the LLM outputs having been previously generated based on previous assistant queries of a previous dialogue session corresponding to the assistant query of the dialogue session, and / or based on one or more of the LLM outputs having been previously generated for the previous context of a previous dialogue session corresponding to the context of the dialogue session, identifying one or more of the LLM outputs previously generated using the LLM model, and making the set of assistant outputs be modified using one or more of the LLM outputs to determine the set of modified assistant outputs. In some further versions of those implementations, identifying one or more of the LLM outputs previously generated using the LLM model may include identifying one or more first LLM outputs of one or more of the LLM outputs that reflect a first personality of a plurality of distinct personalities. The set of modified assistant outputs may include one or more first personality assistant outputs that reflect the first personality. In some still further versions of those implementations, identifying one or more of the LLM outputs previously generated using the LLM model may include identifying one or more second LLM outputs of one or more of the LLM outputs that reflect a second personality of a plurality of distinct personalities. The set of modified assistant outputs may include one or more second personality assistant outputs that reflect the second personality, and the second personality may be distinct from the first personality.In yet further additional or alternative versions of those implementations, the method may further include determining that a previous assistant query of a previous dialogue session corresponds to an assistant query of the dialogue session based on the ASR output including one or more terms of an assistant query corresponding to one or more terms of a previous assistant query of a previous dialogue session. In yet further additional or alternative versions of those implementations, the method may further include generating an embedding of the assistant query based on one or more terms in the ASR output corresponding to the assistant query, and determining that a previous assistant query of a previous dialogue session corresponds to an assistant query of the dialogue session based on comparing the embedding of the assistant query with a previously generated embedding of a previous assistant query of a previous dialogue session. In yet further additional or alternative versions of those implementations, the method may further include determining that a previous context of a previous dialogue session corresponds to a context of the dialogue session based on one or more context signals of the dialogue session corresponding to one or more context signals of a previous dialogue session. In yet further additional versions of those implementations, the one or more context signals may include one or more of time, day of the week, location of the client device, and ambient noise in the environment of the client device. In yet further additional or alternative versions of those implementations, the method may further include generating an embedding of the context of the dialogue session based on the context signal of the dialogue session, and determining that a previous context of a previous dialogue session corresponds to a context of the dialogue session based on comparing the embedding of the one or more context signals with a previously generated embedding of a previous context of a previous dialogue session.

[0115] In some versions of those implementations, to generate additional assistant queries related to an utterance, based at least in part on the context of the dialogue session and at least in part on the assistant queries, processing a set of assistant outputs and the context of the dialogue session may include determining, based on the NLU output, an intention related to the assistant queries included in the utterance; identifying at least one relevant intention related to the intention related to the assistant queries included in the utterance, based on the intention related to the assistant queries included in the utterance; and generating additional assistant queries related to the utterance, based on the at least one relevant intention. In some further versions of those implementations, determining additional assistant output in response to the additional assistant queries, based on the additional assistant queries, may include causing the additional assistant queries to be sent to one or more first-party systems via an application programming interface (API) to generate additional assistant output in response to the additional assistant queries. In some additional or alternative further versions of those implementations, determining additional assistant output in response to the additional assistant queries, based on the additional assistant queries, may include causing the additional assistant queries to be sent to one or more third-party systems via one or more networks and receiving additional assistant output in response to the additional assistant queries in response to the additional assistant queries being sent to one or more of the third-party systems.In some additional or alternative further versions of those implementations, to generate a set of additional modified assistant outputs using one or more additional LLM outputs determined using an LLM or one or more additional LLM outputs, processing the additional assistant outputs and the context of the dialogue session may include processing a set of additional assistant outputs and the context of the dialogue session using the LLM to determine one or more additional LLM outputs and determining a set of additional modified assistant outputs based on the one or more additional LLM outputs. In some additional or alternative further versions of those implementations, to generate a set of additional modified assistant outputs using one or more LLM outputs determined using an LLM or one or more additional LLM outputs, processing the additional assistant outputs and the context of the dialogue session is based on one or more previous assistant queries of a previous dialogue session corresponding to an additional assistant query of the dialogue session having been previously generated and / or based on one or more previous LLM outputs having been previously generated for a previous context of a previous dialogue session corresponding to the context of the dialogue session, identifying one or more of the previously generated additional LLM outputs using the LLM model, and may include causing the set of additional assistant outputs to be modified using the one or more additional LLM outputs to determine the set of additional modified assistant outputs.

[0116] In some implementations, the method may further include ranking a top set of assistant outputs based on one or more ranking criteria, where the top set of assistant outputs includes at least the set of assistant outputs and the set of modified assistant outputs; and selecting a given modified assistant output from the set of modified assistant outputs based on the ranking. In some versions of those implementations, the method may further include ranking a top set of additional assistant outputs based on one or more of the ranking criteria, where the top set of assistant outputs includes at least the additional assistant outputs and the set of additional modified assistant outputs; and selecting a given additional modified assistant output from the set of additional modified assistant outputs based on the ranking. In some further versions of those implementations, causing the given modified assistant output and the given additional modified assistant output to be provided for presentation to the user may include combining the given modified assistant output and the given additional modified assistant output; processing the given modified assistant output and the given additional modified assistant output using a text-to-speech (TTS) model to generate synthetic audio data including a synthetic voice capturing the given modified assistant output and the given additional modified assistant output; and rendering the synthetic audio data audibly for presentation to the user via a speaker of a client device.

[0117] In some implementations, the method may further include ranking a top set of assistant outputs based on one or more ranking criteria, where the top set of assistant outputs includes a set of assistant outputs, a set of modified assistant outputs, additional assistant outputs, and a set of additional modified assistant outputs; and selecting a given modified assistant output from the set of modified assistant outputs and a given additional modified assistant output from the set of additional modified assistant outputs based on the ranking. In some further versions of those implementations, causing the given modified assistant output and the given additional modified assistant output to be provided for presentation to the user may include processing the given modified assistant output and the given additional modified assistant output using a text-to-speech (TTS) model to generate synthetic audio data including a synthetic voice capturing the given modified assistant output and the given additional modified assistant output; and causing the synthetic audio data to be audibly rendered for presentation to the user via a speaker of a client device.

[0118] In some implementations, generating the set of modified assistant outputs using one or more of the LLM outputs may further be based on processing at least a portion of the assistant queries included in the utterance.

[0119] In some implementations, a method implemented by one or more processors is provided as part of an interaction session between a user of a client device and an automated assistant implemented by the client device, the method comprising: receiving a stream of audio data capturing the user's utterance, the stream of audio data being generated by one or more microphones of the client device and the utterance including an assistant query; determining a set of assistant outputs based on processing the stream of audio data, each assistant output in the set of assistant outputs responding to the assistant query included in the utterance; generating a set of assistant outputs modified using one or more LLM outputs generated using a large language model (LLM), each of the one or more LLM outputs being determined based at least in part on the context of the interaction session and one or more of the assistant outputs included in the set of assistant outputs; processing the set of assistant outputs and the context of the interaction session to generate additional assistant queries related to the utterance, at least in part based on the context of the interaction session and at least in part based on the assistant query; determining additional assistant outputs responding to the additional assistant queries based on the additional assistant queries; processing the set of modified assistant outputs based on the additional assistant outputs responding to the additional assistant queries to generate an additional set of modified assistant outputs; and selecting a given additional modified assistant output from the additional set of modified assistant outputs to be provided for presentation to the user.

[0120] In some implementations, a method performed by one or more processors is provided as part of an interaction session between a user of a client device and an auto - assistant implemented by the client device, the method comprising: receiving a stream of audio data capturing the user's utterance, the stream of audio data being generated by one or more microphones of the client device and the utterance including an assistant query; determining a set of assistant outputs based on processing the stream of audio data, each assistant output in the set of assistant outputs responding to the assistant query included in the utterance; processing the set of assistant outputs and the context of the interaction session to generate a set of assistant outputs modified using one or more LLM outputs generated using a large - language model (LLM), each of the one or more LLM outputs being determined based at least in part on the context of the interaction session and one or more of the assistant outputs included in the set of assistant outputs. Generating a set of assistant outputs modified using one or more LLM outputs includes generating a set of first - personality responses based on (i) the set of assistant outputs, (ii) the context of the interaction session, and (iii) one or more first LLM outputs of the one or more LLM outputs that reflect a first personality of a plurality of distinct personalities. The method further includes selecting, from the set of modified assistant outputs, a given modified assistant output for presentation to the user.

[0121] In some implementations, the method performed by one or more processors is provided as part of an interaction session between a user of a client device and an automated assistant performed by the client device, the method comprising: receiving a stream of audio data capturing the user's utterance, the stream of audio data being generated by one or more microphones of the client device and the utterance including an assistant query; determining a set of assistant outputs based on processing the stream of audio data, each assistant output in the set of assistant outputs responding to the assistant query included in the utterance; processing the set of assistant outputs and the context of the interaction session to generate a set of assistant outputs modified using one or more LLM outputs generated using a large language model (LLM), each of the one or more LLM outputs being determined based at least in part on the context of the interaction session and one or more of the assistant outputs included in the set of assistant outputs; and selecting, from the set of modified assistant outputs, a given modified assistant output to be provided for presentation to the user.

[0122] In some implementations, the method implemented by one or more processors is provided as part of an interaction session between a user of a client device and an auto - assistant implemented by the client device, the method including: receiving a stream of audio data that captures the user's utterance, the stream of audio data being generated by one or more microphones of the client device and the utterance including an assistant query; determining a set of assistant outputs based on processing the stream of audio data, each assistant output in the set of assistant outputs responding to the assistant query included in the utterance; determining whether to modify one or more of the assistant outputs included in the set of assistant outputs based on processing the utterance; in response to determining to modify one or more of the assistant outputs included in the set of assistant outputs, processing the set of assistant outputs and the context of the interaction session to generate a set of modified assistant outputs using one or more LLM outputs generated using a large - language model (LLM), each of the one or more LLM outputs being determined based at least in part on the context of the interaction session and one or more of the assistant outputs included in the set of assistant outputs; and selecting, from the set of modified assistant outputs, a given modified assistant output to be provided for presentation to the user.

[0123] These and other implementations of the technology disclosed herein may optionally include one or more of the following features.

[0124] In some implementations, determining whether to modify one or more of the assistant outputs included in a set of assistant outputs may include using an ASR model to process a stream of audio data to generate a stream of automatic speech recognition (ASR) outputs, using an NLU model to process the stream of ASR outputs to generate a stream of natural language understanding (NLU) data, identifying a user's intent in providing the utterance based on the stream of NLU data, and determining whether to modify the assistant output based on the user's intent in providing the utterance, in order to process the utterance.

[0125] In some implementations, determining whether to modify one or more of the assistant outputs included in a set of assistant outputs may further be based on one or more computational costs associated with modifying one or more of the assistant outputs. In some versions of those implementations, the one or more computational costs associated with modifying one or more of the assistant outputs may include one or more of battery consumption, processor consumption associated with modifying one or more of the assistant outputs, or latency associated with modifying one or more of the assistant outputs.

[0126] In some implementations, a method is provided that is performed by one or more processors, the method including obtaining a plurality of assistant queries directed to an automatic assistant and corresponding contexts of corresponding previous dialogue sessions for each of the plurality of assistant queries; processing, using one or more large language models (LLMs), a given assistant query of the plurality of assistant queries to generate a corresponding LLM output for responding to the given assistant query; indexing the corresponding LLM output in a memory accessible in a client device based on the given assistant query and / or the corresponding context of the corresponding previous dialogue session for the given assistant query; after indexing the corresponding LLM output in the memory accessible in the client device, receiving a stream of audio data that captures an utterance of a user as part of a current dialogue session between the user of the client device and the automatic assistant implemented by the client device, the stream of audio data being generated by one or more microphones of the client device; determining, based on processing the stream of audio data, that the utterance includes a current assistant query corresponding to the given assistant query and / or that the utterance is received in a current context of the current dialogue session corresponding to the dialogue context of the corresponding previous dialogue session for the given assistant query; and causing the corresponding LLM output to be utilized in generating an assistant output to be provided to the user in response to the utterance by the automatic assistant.

[0127] These and other implementations of the techniques disclosed herein may optionally include one or more of the following features.

[0128] In some implementations, multiple assistant queries directed to the automatic assistant may have been previously issued by the user via the client device. In some implementations, multiple assistant queries directed to the automatic assistant may have been previously issued by multiple additional users via their respective client devices, in addition to the user of the client device.

[0129] In some implementations, indexing the corresponding LLM output in memory accessible on the client device may be based on the embedding of a given assistant query generated when processing the given assistant query. In some implementations, indexing the corresponding LLM output in memory accessible on the client device may be based on one or more terms or phrases included in the given assistant query generated when processing the given assistant query. In some implementations, indexing the corresponding LLM output in memory accessible on the client device may be based on the embedding of the corresponding context of the corresponding previous dialogue session for the given assistant query. In some implementations, indexing the corresponding LLM output in memory accessible on the client device may be based on one or more context signals included in the corresponding context of the corresponding previous dialogue session for the given assistant query.

[0130] In addition, some implementations include one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) of one or more computing devices, where the one or more processors are operable to execute instructions stored in associated memory, where the instructions are configured to cause execution of any of the methods described above. Some implementations also include one or more non-transitory computer-readable storage media storing computer instructions executable by one or more processors for executing any of the methods described above. Some implementations also include a computer program product including instructions executable by one or more processors for executing any of the methods described above.

Explanation of Signs

[0131] 110 Client device 111 User input engine 112 Rendering engine 113 Presence sensor 114 Auto assistant client 115 Auto assistant 120 Natural conversation system 130 ASR engine 140 NLU engine 150 LLM engine 160 TTS engine 170 Offline output correction engine 171 Assistant activity engine 172 Indexing engine 180 Online output correction engine 181 Assistant query engine 182 Assistant personality engine 190 Ranking engine 191 1P system 192 3P system 199 Network 201 Audio data stream 202 Context 203 ASR output 204 NLU output 205 Assistant output 206 Modified assistant output 207 Given assistant output 710 Computing device 712 Bus subsystem 714 Processor 716 Network interface 720 User interface output device 722 User interface input device 724 Storage subsystem 725 Memory subsystem 726 File storage subsystem 730 RAM 732 ROM

Claims

1. A method implemented by one or more processors, comprising: Receiving, as part of a dialogue session between a user of a client device and an auto-assistant implemented by the client device, a stream of audio data capturing the user's utterance, wherein the stream of audio data is generated by one or more microphones of the client device and the utterance includes an assistant query; Based on processing the stream of audio data, determining a set of assistant outputs, wherein each assistant output in the set of assistant outputs responds to the assistant query included in the utterance; Generating a set of modified assistant outputs modified using a plurality of LLM outputs generated using a large language model (LLM), wherein each of the plurality of LLM outputs is determined based at least in part on the context of the dialogue session and one or more of the assistant outputs included in the set of assistant outputs; Processing the set of assistant outputs and the context of the dialogue session to generate additional assistant queries related to the utterance, based at least in part on the context of the dialogue session and at least in part on the assistant query; Determining additional assistant outputs responsive to the additional assistant queries, based on the additional assistant queries; Processing the additional assistant outputs and the context of the dialogue session to generate a set of additional modified assistant outputs using a plurality of additional LLM outputs generated using the LLM, wherein each of the plurality of additional LLM outputs is determined based at least in part on the context of the dialogue session and the additional assistant outputs; Causing a given modified assistant output from the set of modified assistant outputs and a given additional modified assistant output from the set of additional modified assistant outputs to be provided for presentation to the user. A method, comprising: Receiving, as part of a dialogue session between a user of a client device and an auto-assistant implemented by the client device, a stream of audio data capturing the user's utterance, wherein the stream of audio data is generated by one or more microphones of the client device and the utterance includes an assistant query; Based on processing the stream of audio data, determining a set of assistant outputs, wherein each assistant output in the set of assistant outputs responds to the assistant query included in the utterance; Generating a set of modified assistant outputs modified using a plurality of LLM outputs generated using a large language model (LLM), wherein each of the plurality of LLM outputs is determined based at least in part on the context of the dialogue session and one or more of the assistant outputs included in the set of assistant outputs; Processing the set of assistant outputs and the context of the dialogue session to generate additional assistant queries related to the utterance, based at least in part on the context of the dialogue session and at least in part on the assistant query; Determining additional assistant outputs responsive to the additional assistant queries, based on the additional assistant queries; Processing the additional assistant outputs and the context of the dialogue session to generate a set of additional modified assistant outputs using a plurality of additional LLM outputs generated using the LLM, wherein each of the plurality of additional LLM outputs is determined based at least in part on the context of the dialogue session and the additional assistant outputs; Causing a given modified assistant output from the set of modified assistant outputs and a given additional modified assistant output from the set of additional modified assistant outputs to be provided for presentation to the user.

2. Determining the assistant output that responds to the assistant query included in the utterance based on processing the stream of the audio data, processing the stream of the audio data using an ASR model to generate a stream of automatic speech recognition (ASR) output; processing the stream of the ASR output using an NLU model to generate a stream of natural language understanding (NLU) data; The method according to claim 1, further comprising making the set of the assistant output be determined based on the stream of the NLU.

3. processing the set of the assistant output and the context of the dialogue session to generate a set of modified assistant output using one or more of the LLM outputs generated using the LLM, processing the set of the assistant output and the context of the dialogue session using the LLM to generate one or more of the LLM outputs; The method according to claim 2, further comprising determining the set of the modified assistant output based on one or more of the LLM outputs.

4. processing the set of the assistant output and the context of the dialogue session using the LLM to generate one or more of the LLM outputs, processing the set of the assistant output and the context of the dialogue session using a first set of LLM parameters among a plurality of separate sets of LLM parameters to determine one or more of the LLM outputs having a first personality among a plurality of separate personalities, The method according to claim 3, wherein the set of the modified assistant output includes one or more first personality assistant outputs reflecting the first personality.

5. processing the set of the assistant output and the context of the dialogue session using the LLM to generate one or more of the LLM outputs, To determine one or more of the LLM outputs having the second personality among the plurality of distinct personalities, comprising the step of processing the set of assistant outputs and the context of the conversation session using a second set of LLM parameters among the plurality of distinct sets of LLM parameters, The set of modified assistant outputs includes one or more second personality assistant outputs that reflect the second personality, The method according to claim 4, wherein the second personality is distinct from the first personality.

6. The one or more first personality assistant outputs included in the set of modified assistant outputs and reflecting the first personality are determined using a first vocabulary associated with the first personality, and the one or more second personality assistant outputs included in the set of modified assistant outputs and reflecting the second personality are determined using a second vocabulary associated with the second personality, and the second personality is distinct from the first personality based on the second vocabulary being distinct from the first vocabulary. The method according to claim 5.

7. The one or more first personality assistant outputs included in the set of modified assistant outputs and reflecting the first personality are associated with a first set of prosodic properties utilized in providing the given modified assistant output for audible presentation to the user, and the one or more second personality assistant outputs included in the set of modified assistant outputs and reflecting the second personality are associated with a second set of prosodic properties utilized in providing the given modified assistant output for audible presentation to the user, and the second personality is distinct from the first personality based on the second set of prosodic properties being distinct from the first set of prosodic properties. The method according to claim 5.

8. To generate the set of assistant outputs modified using one or more of the LLM outputs generated using the LLM, the step of processing the set of assistant outputs and the context of the dialogue session is Based on one or more of the LLM outputs having been previously generated based on previous assistant queries of previous dialogue sessions corresponding to the assistant queries of the dialogue session, and / or based on one or more of the LLM outputs having been previously generated for previous contexts of previous dialogue sessions corresponding to the context of the dialogue session, the step of identifying one or more of the LLM outputs previously generated using the LLM; The method according to any one of claims 2 to 7, comprising the step of causing the set of assistant outputs to be modified using one or more of the LLM outputs to determine the set of modified assistant outputs.

9. The step of identifying one or more of the LLM outputs previously generated using the LLM is Comprising the step of identifying one or more first LLM outputs of the one or more LLM outputs that reflect a first personality among a plurality of distinct personalities, The method according to claim 8, wherein the set of modified assistant outputs includes one or more first personality assistant outputs that reflect the first personality.

10. The step of identifying one or more of the LLM outputs previously generated using the LLM is Comprising the step of identifying one or more second LLM outputs of the one or more LLM outputs that reflect a second personality among a plurality of distinct personalities, The set of modified assistant outputs includes one or more second personality assistant outputs that reflect the second personality, The method according to claim 9, wherein the second personality is different from the first personality.

11. The method according to any one of claims 8 to 10, further comprising the step of determining that the previous assistant query of the previous dialogue session corresponds to the assistant query of the dialogue session based on the ASR output including one or more terms of the previous assistant query of the previous dialogue session.

12. generating an embedding of the assistant query based on one or more terms in the ASR output corresponding to the assistant query; The method according to any one of claims 8 to 11, further comprising the step of determining that the previous assistant query of the previous dialogue session corresponds to the assistant query of the dialogue session based on comparing the embedding of the assistant query with the previously generated embedding of the previous assistant query of the previous dialogue session.

13. The method according to any one of claims 8 to 12, further comprising the step of determining that the previous context of the previous dialogue session corresponds to the context of the dialogue session based on the one or more context signals of the dialogue session corresponding to the one or more context signals of the previous dialogue session.

14. The method according to claim 13, wherein the one or more context signals include one or more of time, day of the week, location of the client device, and ambient noise in the environment of the client device.

15. generating an embedding of the context of the dialogue session based on the context signal of the dialogue session; The method according to any one of claims 8 to 14, further comprising the step of determining that the previous context of the previous dialogue session corresponds to the context of the dialogue session based on comparing the embedding of the one or more context signals with the previously generated embedding of the previous context of the previous dialogue session.

16. processing the set of assistant outputs and the context of the dialogue session to generate the additional assistant query related to the utterance based at least in part on the context of the dialogue session and at least in part on the assistant query. Determining an intention related to the assistant query included in the utterance based on the output of the NLU; Identifying at least one relevant intention regarding the intention related to the assistant query included in the utterance based on the intention related to the assistant query included in the utterance; Generating the additional assistant query related to the utterance based on the at least one relevant intention, the method according to any one of claims 2 to 15.

17. The step of determining the additional assistant output in response to the additional assistant query based on the additional assistant query includes The method according to claim 16, comprising the step of causing the additional assistant query to be sent to one or more first-party systems via an application programming interface (API) to generate the additional assistant output in response to the additional assistant query.

18. The step of determining the additional assistant output in response to the additional assistant query based on the additional assistant query includes Causing the additional assistant query to be sent to one or more third-party systems via one or more networks; and Receiving the additional assistant output in response to the additional assistant query in response to the additional assistant query being sent to one or more of the third-party systems, the method according to claim 16.

19. Processing the additional assistant output and the context of the dialogue session to generate the set of additional modified assistant outputs using the plurality of additional LLM outputs determined using the LLM includes Processing the set of additional assistant outputs and the context of the dialogue session using the LLM to determine the plurality of additional LLM outputs; and Determining the set of additional modified assistant outputs based on the plurality of additional LLM outputs, the method according to any one of claims 16 to 18.

20. Processing the additional assistant output and the context of the dialogue session to generate the set of additional modified assistant outputs using the plurality of additional LLM outputs determined using the LLM, Identifying the plurality of additional LLM outputs previously generated using the LLM, based on the plurality of additional LLM outputs having been previously generated based on previous assistant queries of a previous dialogue session corresponding to the additional assistant queries of the dialogue session and / or based on the plurality of additional LLM outputs having been previously generated for previous contexts of the previous dialogue session corresponding to the context of the dialogue session, The method according to any one of claims 16 to 18, comprising: modifying the set of additional assistant outputs using the plurality of additional LLM outputs to generate the set of additional modified assistant outputs. **Claim 21** The step of providing the given modified assistant output and the given additional modified assistant output for presentation to the user, Combining the given modified assistant output and the given additional modified assistant output, Processing the given modified assistant output and the given additional modified assistant output using a text-to-speech (TTS) model to generate synthetic audio data including a synthetic voice capturing the given modified assistant output and the given additional modified assistant output, The method according to any one of claims 1 to 20, comprising: rendering the synthetic audio data audible for presentation to the user via a speaker of the client device. **Claim 22** The step of providing the given modified assistant output and the given additional modified assistant output for presentation to the user, To generate synthetic audio data including synthetic speech that captures the given modified assistant output and the given additional modified assistant output, using a text-to-speech (TTS) model to process the given modified assistant output and the given additional modified assistant output; rendering the synthetic audio data audibly for presentation to the user via a speaker of the client device. The method according to any one of claims 1 to 21, comprising the step of **Claim 23** The method according to any one of claims 1 to 22, wherein the step of generating a set of modified assistant outputs using one or more of the LLM outputs further comprises processing at least a portion of the assistant queries included in the utterance. **Claim 24** A method implemented by one or more processors, as part of an interaction session between a user of a client device and an automated assistant implemented by the client device, receiving a stream of audio data capturing the user's utterance, the stream of audio data being generated by one or more microphones of the client device, the utterance including an assistant query; determining a set of assistant outputs based on processing the stream of audio data, each assistant output in the set of assistant outputs responding to the assistant query included in the utterance; generating a set of modified assistant outputs using a plurality of LLM outputs generated using a large language model (LLM), each of the plurality of LLM outputs being determined based at least in part on the context of the interaction session and one or more of the assistant outputs included in the set of assistant outputs; generating additional assistant queries related to the utterance based at least in part on the context of the interaction session and at least in part on the assistant query, for which purpose, processing the set of assistant outputs and the context of the interaction session. Based on the additional assistant query, determining an additional assistant output that responds to the additional assistant query; processing the set of modified assistant outputs based on the additional assistant output that responds to the additional assistant query to generate a set of additional modified assistant outputs; selecting a given additional modified assistant output from the set of additional modified assistant outputs for presentation to the user; A method comprising: **Claim 25** A method implemented by one or more processors, as part of an interaction session between a user of a client device and an automated assistant implemented by the client device, receiving a stream of audio data capturing the user's utterance, the stream of audio data being generated by one or more microphones of the client device, the utterance including an assistant query; determining a set of assistant outputs based on processing the stream of audio data, each assistant output in the set of assistant outputs responding to the assistant query included in the utterance; generating a set of modified assistant outputs using a plurality of LLM outputs generated using a large language model (LLM), each of the plurality of LLM outputs being determined based at least in part on the context of the interaction session and one or more of the assistant outputs included in the set of assistant outputs, generating the set of modified assistant outputs using one or more of the LLM outputs comprising: generating a set of first personality responses based on (i) the set of assistant outputs, (ii) the context of the interaction session, and (iii) one or more first LLM outputs of the one or more LLM outputs that reflect a first personality of a plurality of distinct personalities; generating additional assistant queries related to the utterance based at least in part on the context of the dialogue session and at least in part on the assistant query processing the set of assistant outputs and the context of the dialogue session for the purpose of determining additional assistant outputs in response to the additional assistant queries based on the additional assistant queries processing the additional assistant outputs and the context of the dialogue session to generate a set of additional modified assistant outputs using a plurality of additional LLM outputs generated using the LLM, wherein each of the plurality of additional LLM outputs is determined based at least in part on the context of the dialogue session and the additional assistant outputs causing a given modified assistant output from the set of modified assistant outputs and a given additional modified assistant output from the set of additional modified assistant outputs to be provided for presentation to the user A method comprising: Claim 26 A method implemented by one or more processors, comprising: receiving, as part of a dialogue session between a user of a client device and an automated assistant implemented by the client device, a stream of audio data capturing the utterance of the user, the stream of audio data being generated by one or more microphones of the client device and the utterance including an assistant query determining a set of assistant outputs based on processing the stream of audio data, each assistant output in the set of assistant outputs responding to the assistant query included in the utterance determining whether to modify one or more of the assistant outputs included in the set of assistant outputs based on the processing of the utterance in response to determining to modify one or more of the assistant outputs included in the set of assistant outputs Generating a set of assistant outputs modified using a plurality of LLM outputs generated using a large language model (LLM), wherein each of the plurality of LLM outputs is determined based at least in part on the context of the conversation session and one or more of the assistant outputs included in the set of assistant outputs; Generating additional assistant queries related to the utterance based at least in part on the context of the conversation session and at least in part on the assistant query; Processing the set of assistant outputs and the context of the conversation session for the purpose of; Determining additional assistant outputs responsive to the additional assistant queries based on the additional assistant queries; Processing the additional assistant outputs and the context of the conversation session to generate a set of additional modified assistant outputs using a plurality of additional LLM outputs generated using the LLM, wherein each of the plurality of additional LLM outputs is determined based at least in part on the context of the conversation session and the additional assistant outputs; Providing a given modified assistant output from the set of modified assistant outputs and a given additional modified assistant output from the set of additional modified assistant outputs for presentation to the user; A method comprising. Claim 27 The step of determining whether to modify one or more of the assistant outputs included in the set of assistant outputs based on the processing of the utterance; Processing the stream of audio data using an ASR model to generate a stream of automatic speech recognition (ASR) outputs; Processing the stream of ASR outputs using an NLU model to generate a stream of natural language understanding (NLU) data; Identifying the user's intent in providing the utterance based on the stream of NLU data; The method according to claim 26, further comprising determining whether to modify the assistant output based on the user's intent in providing the utterance. Claim 28 The method according to claim 26 or 27, further comprising a step of determining whether to modify one or more of the assistant outputs included in the set of assistant outputs, based on one or more computational costs associated with modifying one or more of the assistant outputs.

29. The method according to claim 28, wherein the one or more computational costs associated with modifying one or more of the assistant outputs include one or more of battery consumption, processor consumption associated with modifying one or more of the assistant outputs, or latency associated with modifying one or more of the assistant outputs.

30. At least one processor, A system comprising a memory storing instructions that, when executed, cause the at least one processor to perform operations corresponding to any one of claims 1 to 29.

31. A non-transitory computer-readable storage medium storing instructions that, when executed, cause at least one processor to perform operations corresponding to any one of claims 1 to 29.

Citation Information

Patent Citations

  • Device and method for answer sentence generation, and program and storage medium thereof

    JP2007102104A

  • Voice interactive device and voice interactive program

    JP2009186989A

  • Response generation device, response generation method, and response generation program

    JP2016045584A

  • Response generation device, response generation method, and response generation program

    WO2020105302A1

Cited By

  • Domain-knowledge guided agent framework for automated system analysis

    US20250335707A1