Generating automated assistant responses using large language models
Patent Information
- Application Number
- CN202180096665.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-11-22
- Filing Date
- 2021-11-30
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2041-11-30
AI Technical Summary
因此,由自动化助理响应于第一人的口述话语提供的响应可能不会与第一人产生共鸣,因为响应可能无法反映多人之间的自然谈话
[0022]应当理解,本文公开的技术可以在客户端设备上本地实现、由经由一个或多个网络连接到客户端设备的服务器远程实现、和/或两者。
Smart Images

Figure CN117136405B_ABST
Abstract
Description
Background Technology
[0001] Humans can engage in human-computer dialogue with interactive software applications referred to herein as “automated assistants” (also known as “chatbots,” “interactive personal assistants,” “intelligent personal assistants,” “personal voice assistants,” “conversation agents,” etc.). Automated assistants typically rely on a series of components to interpret and respond to spoken utterances. For example, an Automatic Speech Recognition (ASR) engine can process audio data corresponding to a user’s spoken utterances to generate ASR outputs, such as ASR hypotheses of the spoken utterances (i.e., sequences of terms and / or other tokens). Furthermore, a Natural Language Understanding (NLU) engine can process the ASR outputs (or touch / type input) to generate NLU outputs, such as the request (e.g., intent) expressed by the user when providing spoken utterances (or touch / type input) and slot values of parameters optionally associated with the intent. Finally, the NLU outputs can be processed by various performance components to generate performance outputs, such as responsive content in response to spoken utterances and / or one or more actions that can be performed in response to spoken utterances.
[0002] Generally, a conversation with an automated assistant begins with a spoken utterance from the user, and the automated assistant responds to the utterance using the aforementioned set of components. The user can continue the conversation by providing additional spoken utterances, and the automated assistant can again respond to these additional utterances using the same set of components. In other words, these conversations are typically turn-based, as the user provides spoken utterances in turns, the automated assistant responds in turns, the user provides additional spoken utterances in additional turns, the automated assistant responds in additional turns, and so on. However, from the user's perspective, these turn-based conversations may feel unnatural because they do not reflect how humans actually talk to each other.
[0003] For example, if the first person provides spoken utterances during a conversation to convey an initial idea to the second person (e.g., “I’m going to the beach today”), the second person can consider those spoken utterances within the context of the conversation to formulate a response to the first person (e.g., “Sounds fun, what are you going to do at the beach?”, “Nice, have you looked at the weather?”, etc.). It is noteworthy that the second person can provide spoken utterances in response to the first person, which allows the first person to continue the conversation in a natural way. In other words, during a conversation, both the first and second person can provide spoken utterances to facilitate natural dialogue, and neither person needs to drive the conversation.
[0004] However, if the second person is replaced by an automated assistant in the above example, the automated assistant may not be able to provide a response that allows the first person to continue the conversation. For example, in response to the first person's spoken statement "I'm going to the beach today," the automated assistant might simply respond with "sounds fun" or "nice" without providing any additional response to facilitate the conversation. This is even though the automated assistant could perform an action and / or provide a response to facilitate the conversation, such as proactively asking the first person what they plan to do at the beach, proactively looking up the weather forecast for a beach the first person frequently visits and including it in the response, or proactively making inferences based on the weather forecast. Therefore, the response provided by the automated assistant in response to the first person's spoken statement may not resonate with the first person because the response may not reflect a natural conversation between multiple people. Furthermore, the first person may have to provide additional spoken statements to explicitly request certain information that the automated assistant could proactively provide (e.g., the beach weather forecast), thus increasing the number of spoken statements directed to the automated assistant and wasting the computing resources of the client device processing these statements. Summary of the Invention
[0005] The embodiments described herein relate to enabling an automated assistant to perform natural conversation with a user during a dialogue session. Some embodiments may receive an audio data stream capturing the user's spoken utterances. The audio data stream may be generated by one or more microphones of a client device, and the spoken utterances may include assistant queries. Some embodiments may further determine an assistant output set based on processing the audio data stream, and process the assistant output set and the context of the dialogue session to generate a modified assistant output set using one or more LLM outputs generated using a large language model (LLM). Each of the one or more LLM outputs may be determined based on at least a portion of the context of the dialogue session and one or more assistant outputs included in the assistant output set. Some embodiments may also enable a given modified assistant output from the modified assistant output set to be provided to the user. Furthermore, each of the one or more LLM outputs may include, for example, a probability distribution of a sequence of one or more words and / or phrases across one or more vocabularies, and one or more words and / or phrases in that sequence may be selected as one or more LLM outputs based on the probability distribution. In addition, the context of a conversation session can be determined based on one or more context signals, such as time of day, day of week, location of the client device, ambient noise detected in the environment of the client device, user profile data, software application data, environmental data about the known environment of the user on the client device, conversation history of the conversation session between the user and the automation assistant, and / or other context signals.
[0006] In some implementations, an assistant output set can be determined based on processing an audio data stream using a streaming Automatic Speech Recognition (ASR) model to generate an ASR output stream, such as one or more identified words or phrases predicted to correspond to spoken utterances, one or more phonemes predicted to correspond to spoken utterances, one or more predictive metrics associated with each of the one or more identified words or phrases and / or one or more predicted phonemes, and / or other ASR outputs. Furthermore, a Natural Language Understanding (NLU) model can be used to process the ASR outputs to generate an NLU output stream, such as one or more predicted intentions of the user when providing spoken utterances and one or more corresponding slot values of one or more parameters associated with each of the one or more predicted intentions. Additionally, the NLU data stream can be processed by one or more first-party (1P) and / or third-party (3P) systems to generate the assistant output set. As used herein, one or more 1P systems include systems developed and / or maintained by the same entity (e.g., a common publisher) that develops and / or maintains the automated assistant described herein, while one or more 3P systems include systems developed and / or maintained by entities different from the entity that develops and / or maintains the automated assistant described herein. It is worth noting that the assistant output set described herein includes assistant outputs typically considered for use in response to spoken utterances. However, by using the claimed techniques, the assistant output set generated in the manner described above can be further processed to generate a modified assistant output set. Specifically, the assistant output set can be modified using one or more LLM outputs, and a given modified assistant output can be selected from the modified assistant output set to be provided to the user in response to the receipt of spoken utterances.
[0007] For example, suppose a user on a client device provides the spoken phrase, "Hey Assistant, I'm thinking about going surfing today." In this example, the automated assistant can process the spoken phrase as described above to generate an assistant output set and a modified assistant output set. In this example, the assistant outputs included in the assistant output set could include, for example, "That sounds like fun!", "Sounds fun!", etc. Furthermore, in this example, the assistant outputs included in the modified assistant output set could include, for example, "That sounds like fun, how long have you been surfing?", "Enjoy it, but if you're going to Example Beach again, be prepared for some light showers," etc. It is worth noting that the assistant outputs included in the assistant output set do not include any assistant outputs that drive the conversation in a manner that further engages the user of the client device in the conversation. However, the assistant outputs included in the modified assistant output set include assistant outputs that drive the conversation by asking context-sensitive questions (e.g., “how long have you been surfing?”), providing context-sensitive information (e.g., “but if you’re going to Example Beach again, be prepared for some light showers”), and / or otherwise resonate with the user of the client device in the context of the conversation.
[0008] In some implementations, a modified set of assistant responses can be generated using one or more LLM outputs generated online. For example, in response to receiving spoken utterances, the automated assistant can generate an assistant output set as described above. Furthermore, also in response to receiving spoken utterances, the automated assistant can use one or more LLMs to process the assistant output set, the context of the dialogue session, and / or assistant queries included in the spoken utterances to generate a modified set of assistant outputs based on one or more LLM outputs generated using one or more LLMs.
[0009] In additional or alternative implementations, a modified set of assistant responses can be generated using one or more LLM outputs generated offline. For example, before receiving spoken utterances, the automated assistant can obtain multiple assistant queries and corresponding contexts of previous conversations for each of the multiple assistant queries from an assistant activity database (which may be a limited set of assistant activities of a user on a client device). Furthermore, the automated assistant can cause, for a given assistant query among the multiple assistant queries, to generate a set of assistant outputs in the manner described above and for that given assistant query. Additionally, the automated assistant can cause to use one or more LLMs to process the set of assistant outputs, the corresponding contexts of the conversations, and / or a given assistant query to generate a modified set of assistant outputs based on one or more LLM outputs generated using one or more LLMs. This process can be repeated for each of the multiple queries and the corresponding contexts of previous conversations obtained by the automated assistant.
[0010] Additionally, the automation assistant can index one or more LLM outputs in memory accessible by the user's client device. In some implementations, the automation assistant can cause one or more LLMs to be indexed in memory based on one or more terms included in multiple assistant queries. In additional or alternative implementations, the automation assistant can generate a corresponding embedding (e.g., a word2vec embedding or another low-dimensional representation) for each of the multiple assistant queries and map each of the corresponding embeddings to an assistant query embedding space to index one or more LLM outputs. In additional or alternative implementations, the automation assistant can cause one or more LLMs to be indexed in memory based on one or more context signals included in a corresponding previous context. In additional or alternative implementations, the automation assistant can generate a corresponding embedding for each of the corresponding contexts and map each of the corresponding embeddings to a context embedding space to index one or more LLM outputs. In additional or alternative implementations, the automation assistant can cause one or more LLMs to be indexed in memory based on one or more terms or phrases of the assistant outputs included in the set of assistant outputs for each of the multiple assistant queries. In an additional or alternative implementation, the automation assistant may generate a corresponding embedding (e.g., a word2vec embedding, or another low-dimensional representation) for each assistant output included in the assistant output set, and map each corresponding embedding to the assistant output embedding space to index one or more LLM outputs.
[0011] Therefore, when spoken utterances are subsequently received at the user's client device, the automation assistant can identify one or more LLM outputs previously generated based on a current assistant query corresponding to one or more assistant queries included in multiple queries, a current context corresponding to one or more corresponding previous contexts, and / or one or more current assistant outputs corresponding to one or more previous assistant outputs. For example, in an implementation where one or more LLM outputs are indexed based on the corresponding embeddings of previous assistant queries, the automation assistant can cause the generation of the embedding of the current assistant query and its mapping to an assistant query embedding space. Furthermore, the automation assistant can determine that the current assistant query corresponds to a previous assistant query based on a distance between the embedding of the current assistant query and the corresponding embedding of the previous assistant query in the query embedding space satisfying a threshold. The automation assistant can retrieve one or more LLM outputs generated based on processing previous assistant queries from memory and utilize these one or more LLM outputs to generate a modified set of assistant outputs. Additionally, for example, in an implementation where one or more LLMs are indexed based on one or more terms included in multiple assistant queries, the automation assistant can determine, for example, the edit distance between the current assistant query and multiple previous assistants to identify the previous assistant queries corresponding to the current assistant query. Similarly, the automation assistant can obtain one or more LLM outputs generated from processing previous assistant queries from memory, and use those one or more LLM outputs to generate a modified set of assistant outputs.
[0012] In some implementations, and in addition to one or more LLM outputs, additional assistant queries can be generated based on the context of processing the assistant query and / or conversation session. For example, when processing the context of the assistant query and / or conversation session, the automation assistant can determine the intent associated with a given assistant query based on NLU data flow. Furthermore, the automation assistant can identify at least one related intent associated with the intent associated with the given assistant query (e.g., based on a mapping of that intent to at least one related intent in a database or memory accessible to the client device and / or based on processing the intent associated with the given assistant query using one or more machine learning (ML) models or heuristically defined rules). Additionally, the automation assistant can generate additional assistant queries based on at least one related intent. For example, suppose the assistant query instructs the user to go to the beach (e.g., “Hey assistant, I'm going to the beach today”). In this example, the additional assistant query could correspond to, for example, “what's the weather at Example Beach?” (e.g., to proactively determine weather information for a beach called Example Beach that the user typically visits). It is worth noting that additional assistant queries may not be provided to users of the client device.
[0013] Conversely, in these implementations, additional assistant outputs can be determined based on processing additional assistant queries. For example, the automated assistant can transmit a structured request to one or more 1P and / or 3P systems to obtain weather information as additional assistant output. Further assuming the weather information indicates that rain is expected at Example Beach. In some versions of these implementations, the automated assistant can further cause additional assistant outputs to be processed using one or more LLM outputs and / or one or more additional LLM outputs to generate an additional modified set of assistant outputs. Thus, in the initial example provided above, a given modified assistant output from the initial modified assistant output set provided to the user in response to receiving the spoken phrase “Hey Assistant, I'm thinking about going surfing today” might be “Enjoy it,” and a given additional modified assistant output from the additional modified assistant output set might be “but if you're going to Example Beach again, be prepared for some light showers.” In other words, the automated assistant…
[0014] In various implementations, when generating a modified set of assistant outputs, each of the one or more LLM outputs utilized can be generated using corresponding parameter sets from multiple disparate parameter sets. Each parameter in the multiple disparate parameter sets can be associated with a disparate personality of the automated assistant. In some versions of those implementations, a single LLM can be used to generate one or more corresponding LLM outputs using corresponding parameter sets for each disparate personality, while in other versions of those implementations, multiple LLMs can be used to generate one or more corresponding LLM outputs using corresponding parameter sets for each disparate personality. Thus, when a given modified assistant output from the modified set of assistant outputs is provided to the user, it can reflect various dynamic contextual personalities via prosodic attributes of different personalities (e.g., intonation, cadence, pitch, pauses, speech rate, stress, rhythm, etc. of these different personalities).
[0015] It is worth noting that the personalized responses described in this paper not only reflect the prosodic attributes of different personalities, but also the dissimilar vocabulary and / or the dissimilar speaking styles of different personalities (e.g., lengthy speaking styles, concise speaking styles, etc.). For example, a first set of parameters can be used to generate a given modified assistant output to be presented to the user, which reflects the primary personality of the automated assistant in terms of a first set of vocabulary used by the automated assistant and / or a first set of prosodic attributes used to provide the modified assistant output for audible presentation to the user. Alternatively, a second set of parameters can be used to generate a modified assistant output to be presented to the user, which reflects the secondary personality of the automated assistant in terms of a second set of vocabulary used by the automated assistant and / or a second set of prosodic attributes used to provide the modified assistant output for audible presentation to the user.
[0016] Therefore, the automated assistant can dynamically adjust the personality used when providing modified assistant output to the user based on both the vocabulary and prosodic attributes used by the automated assistant when rendering the modified assistant output for audible presentation. Notably, the automated assistant can dynamically adjust these personalities based on the context of the conversational session—including previous utterances received from the user and previous assistant output provided by the automated assistant, and / or any other contextual signals described herein. As a result, the modified assistant output provided by the automated assistant can better resonate with the user on the client device. Furthermore, it should be noted that the personality used in a given conversational session can be dynamically adjusted as the context of that conversational session is updated.
[0017] In some implementations, the automated assistant can rank assistant outputs included in the assistant output set (i.e., those not generated using one or more LLM outputs) and modified assistant outputs (i.e., those generated using one or more LLM outputs) according to one or more ranking criteria. Therefore, when selecting a given assistant output to be presented to a user, the automated assistant can choose from both the assistant output set and the modified assistant output set. One or more ranking criteria may include, for example: one or more prediction metrics (e.g., ASR metrics generated when generating the ASR output stream, NLU degrees generated when generating the NLU output stream, fulfillment metrics generated when generating the assistant output set) – which indicate how responsive each assistant output included in the assistant output set and the modified assistant output set is predicted to be to an assistant query included in the spoken utterance; one or more intents included in the NLU output stream; and / or other ranking criteria. For example, if the user's intent on the client device indicates that the user wants a factual answer (e.g., based on spoken utterances providing an assistant query including "why is the sky blue?"), the automated assistant may facilitate one or more assistant outputs included in a set of one or more assistant outputs, because the user may want a direct answer to the assistant query. However, if the user's intent on the client device indicates that the user provides more open-ended input (e.g., based on spoken utterances providing an assistant query including "what time is it?"), the automated assistant may facilitate one or more assistant outputs included in a modified set of assistant outputs, because the user may prefer more conversational aspects.
[0018] In some implementations, and before generating a modified set of assistant outputs, the automation assistant may determine whether or not to generate a modified set of assistant outputs. In some versions of those implementations, the automation assistant may determine whether or not to generate a modified set of assistant outputs based on one or more predicted intentions of the user when providing utterances, such as those indicated by NLU data streams. For example, in an implementation where the automation assistant determines that the utterance requests the automation assistant to perform a search (e.g., the assistant queries "Why is the sky blue?"), the automation assistant may determine not to generate a modified set of assistant outputs because the user is seeking a factual answer. In additional or alternative versions of those implementations, the automation assistant may determine whether or not to generate a modified set of assistant outputs based on one or more computational costs associated with modifying one or more assistant outputs. One or more computational costs associated with modifying one or more assistant outputs may include, for example, one or more of the following: battery consumption, processor consumption, or latency associated with modifying one or more assistant outputs. For example, if the client device is in a low-power mode, the automation assistant may determine not to generate a modified set of assistant outputs to reduce the client device's battery consumption.
[0019] By using the techniques described herein, one or more technical advantages can be achieved. As a non-limiting example, the techniques described herein enable automated assistants to engage in natural conversation with users during a dialogue session. For example, the automated assistant can generate modified assistant outputs using one or more LLM outputs that are inherently more conversational. Thus, the automated assistant can proactively provide contextual information related to the dialogue session that is not directly requested by the user (e.g., by generating additional assistant queries as described herein and by providing additional assistant outputs determined based on these additional assistant queries), thereby making the modified assistant outputs resonate with the user. Furthermore, modified assistant outputs with various personalities can be generated based on vocabulary that is context-adjusted throughout the dialogue session and on prosodic properties used to audibly render the modified assistant outputs, further resonating with the user. This results in various technical advantages in saving computational resources at the client device and can enable dialogue sessions to end more quickly and efficiently and / or reduce the number of dialogue sessions. For example, the amount of user input received at the client device can be reduced because the number of times the user must request information related to the dialogue session context can be reduced, as this can be proactively provided to the user by the automated assistant. Furthermore, for example, in implementations that generate one or more LLM outputs offline and then use them online, runtime waiting time can be reduced.
[0020] As used herein, a “conversational session” can include logically self-contained exchanges between a user and an automated assistant (and in some cases, other human participants). An automated assistant can distinguish between multiple conversational sessions with a user based on various signals, such as the passage of time between sessions, changes in user context (e.g., location, before / during / after a scheduled meeting), detection of one or more interventional interactions between the user and the client device other than the conversation between the user and the automated assistant (e.g., the user switching applications for a period of time, the user leaving and then returning to a standalone voice-activated product), locking / hibernating the client device between sessions, changes in the client device used to interface with the automated assistant, and so on. Notably, during a given conversational session, the user can interact with the automated assistant using various input modalities, including but not limited to dictation, typing, and / or touch input.
[0021] The above description is provided as an overview of some embodiments disclosed herein. Those embodiments and other embodiments are described in more detail herein.
[0022] It should be understood that the techniques disclosed herein can be implemented locally on the client device, remotely by a server connected to the client device via one or more networks, and / or both. Attached Figure Description
[0023] Figure 1 A block diagram of an example environment is depicted, which demonstrates various aspects of this disclosure and in which the implementations disclosed herein can be carried out.
[0024] Figure 2 Example processing flows for generating assistant output using a large language model according to various implementations are described.
[0025] Figure 3 A flowchart illustrating example methods for generating assistant output offline using a large language model, according to various implementations, for subsequent online use.
[0026] Figure 4 A flowchart illustrating example methods for generating assistant output based on a large language model and an assistant query, according to various implementations, is provided.
[0027] Figure 5 The flowchart illustrates example methods for generating assistant output based on personalized assistant responses using a large language model, according to various implementations.
[0028] Figure 6Non-limiting examples of dialogue sessions between a user and an automated assistant according to various implementations are depicted, in which the automated assistant utilizes a large language model to generate assistant output.
[0029] Figure 7 Example architectures of computing devices according to various implementation methods are depicted. Detailed Implementation
[0030] Turn now Figure 1 This document depicts a block diagram of an example environment 100 that illustrates various aspects of this disclosure and in which the embodiments disclosed herein can be implemented. Example environment 100 includes a client device 110 and a natural conversation system 120. In some embodiments, the natural conversation system 120 may be implemented locally at the client device 110. In additional or alternative embodiments, the natural conversation system 120 may be implemented as follows: Figure 1 The client device 110 depicted is remotely implemented (e.g., at a remote server). In these embodiments, the client device 110 and the natural conversation system 120 may be communicatively coupled to each other via one or more networks 199, such as one or more wired or wireless local area networks (“LANs”, including Wi-Fi LANs, mesh networks, Bluetooth, near field communication, etc.) or wide area networks (“WANs”, including the Internet).
[0031] Client device 110 may be one or more of the following: desktop computer, laptop computer, tablet computer, mobile phone, vehicle computing device (e.g., in-vehicle communication system, in-vehicle entertainment system, in-vehicle navigation system), stand-alone interactive speaker (optionally with a display), smart appliance such as a smart TV, and / or wearable device for a user including a computing device (e.g., a watch for a user with a computing device, glasses for a user with a computing device, virtual or augmented reality computing device). Additional and / or alternative client devices may be provided.
[0032] Client device 110 can execute automation assistant client 114. An instance of automation assistant client 114 can be an application separate from the operating system of client device 110 (e.g., installed "on top" of the operating system) – or alternatively, it can be implemented directly by the operating system of client device 110. Automation assistant client 114 can interact with natural conversation system 120, which is implemented locally or remotely at client device 110 and invoked via one or more networks 199, such as... Figure 1As depicted in [the text]. An automation assistant client 114 (and optionally through its interaction with other remote systems (e.g., servers)) can form a logical instance that appears from the user's perspective to be an automation assistant 115, through which the user can engage in human-computer dialogue. An instance of automation assistant 115 is... Figure 1 The diagram is depicted and surrounded by dashed lines representing the Automated Assistant Client 114 (including Client Device 110) and the Natural Conversation System 120. Therefore, it should be understood that a user interacting with the Automated Assistant Client 114 running on Client Device 110 can actually interact with a logical instance of his or her own Automated Assistant 115 (or a logical instance of Automated Assistant 115 shared among family members or other user groups). For the sake of brevity and simplicity, Automated Assistant 115 as used herein will refer to the Automated Assistant Client 114, which runs locally on Client Device 110 and / or remotely on one or more remote servers that may implement the Natural Conversation System 120.
[0033] In various embodiments, client device 110 may include a user input engine 111 configured to detect user input provided by a user of client device 110 using one or more user interface input devices. For example, client device 110 may be equipped with one or more microphones that capture audio data, such as audio data corresponding to a user's spoken words or other sounds in the environment of client device 110. Additionally or alternatively, client device 110 may be equipped with one or more visual components configured to capture visual data corresponding to images and / or movements (e.g., gestures) detected in the field of view of one or more visual components. Additionally or alternatively, client device 110 may be equipped with one or more touch-sensitive components (e.g., keyboard and mouse, stylus, touchscreen, touch panel, one or more hardware buttons, etc.) configured to capture signals corresponding to touch input directed at client device 110.
[0034] In various implementations, client device 110 may include rendering engine 112 configured to provide content for audible and / or visual presentation to a user of client device 110 using one or more user interface output devices. For example, client device 110 may be equipped with one or more speakers, enabling content to be provided to a user for audible presentation via client device 110. Additionally or alternatively, client device 110 may be equipped with a display or projector, enabling content to be provided to a user for visual presentation via client device 110.
[0035] In various embodiments, client device 110 may include one or more presence sensors 113 configured to provide signals indicating a detected presence, particularly a human presence, upon approval from the corresponding user. In some of those embodiments, the automation assistant 115 may identify client device 110 (or another computing device associated with the user of client device 110) at least in part based on the presence of the user at client device 110 (or at another computing device associated with the user of client device 110) to satisfy a spoken utterance. The utterance may be satisfied by rendering response content at client device 110 and / or the other computing device associated with the user of client device 110 (e.g., via rendering engine 112), by controlling client device 110 and / or the other computing device associated with the user of client device 110, and / or by performing any other actions to satisfy the utterance. As described herein, the automation assistant 115 may utilize data determined by the presence sensor 113 to determine the client device 110 (or other computing device) based on where the user is or has recently been in the vicinity, and provide corresponding commands only to the client device 110 (or those other computing devices). In some additional or alternative embodiments, the automation assistant 115 may utilize data determined by the presence sensor 113 to determine whether any user (any user or a specific user) is currently near the client device 110 (or other computing device), and may optionally suppress the provision of data to and / or from the client device 110 (or other computing device) based on the user's proximity to the client device 110 (or other computing device).
[0036] The presence sensor 113 can take various forms. For example, the client device 110 can utilize one or more user interface input components described above with respect to the user input engine 111 to detect the presence of a user. Additionally or alternatively, the client device 110 can be equipped with other types of light-based presence sensors 113, such as a passive infrared (“PIR”) sensor that measures infrared (“IR”) light radiated from an object within its field of view.
[0037] Additionally or alternatively, in some embodiments, the presence sensor 113 may be configured to detect other phenomena associated with human presence or device presence. For example, in some embodiments, the client device 110 may be equipped with the presence sensor 113, which detects various types of wireless signals (e.g., waves such as radio, ultrasound, electromagnetic, etc.) emitted by other computing devices (e.g., mobile devices, wearable computing devices, etc.) and / or other computing devices, such as those carried / operated by the user. For example, the client device 110 may be configured to emit waves imperceptible to humans, such as ultrasound or infrared waves, which can be detected by other computing devices (e.g., via an ultrasound / infrared receiver such as a microphone with ultrasound capability).
[0038] Additionally or alternatively, client device 110 may emit other types of waves imperceptible to humans, such as radio waves (e.g., Wi-Fi, Bluetooth, cellular, etc.) that can be detected by other computing devices (e.g., mobile devices, wearable computing devices, etc.) carried / operated by the user and used to determine the user's specific location. In some embodiments, GPS and / or Wi-Fi triangulation may be used, for example, to detect a person's location based on GPS and / or Wi-Fi signals to / from client device 110. In other embodiments, other wireless signal characteristics such as time of flight, signal strength, etc., may be used by client device 110, alone or in combination, to determine the location of a specific person based on signals emitted by other computing devices carried / operated by the user. Additionally or alternatively, in some embodiments, client device 110 may perform speaker recognition (SID) to identify the user based on the user's voice and / or perform facial recognition (FID) to identify the user based on visual data captured of the user's face.
[0039] In some implementations, the speaker movement can then be determined, for example, by the presence sensor 113 of the client device 110 (and optionally by a GPS sensor, Soli chip, and / or accelerometer of the client device 110). In some implementations, based on this detected movement, the user's location can be predicted, and when any content is rendered at the client device 110 and / or other computing device based at least in part on the proximity of the client device 110 and / or other computing device to the user's location, that location can be assumed to be the user's location. In some implementations, it can be simply assumed that the user is in the location where he or she last interacted with the automation assistant 115, especially if not too much time has passed since the last interaction.
[0040] Furthermore, client device 110 and / or natural conversation system 120 may include one or more memories for storing data and / or software applications, one or more processors for accessing data and executing software applications, and / or other components that facilitate communication via one or more of the networks 199. In some embodiments, one or more of the software applications may be locally installed at client device 110, while in other embodiments, one or more of the software applications may be remotely hosted (e.g., via one or more servers) and accessible by client device 110 via one or more of the networks 199.
[0041] In some implementations, the operations performed by the automation assistant 115 can be implemented locally at the client device 110 via the automation assistant client 114. For example... Figure 1 As shown, the automated assistant client 114 may include an Automatic Speech Recognition (ASR) engine 130A1, a Natural Language Understanding (NLU) engine 140A1, a Large Language Model (LLM) engine 150A1, and a Text-to-Speech (TTS) engine 160A1. In some implementations, such as when... Figure 1 The operations performed by the automation assistant 115 when the natural conversation system 120 is remotely implemented on the client device 110, as described herein, can be distributed across multiple computer systems. In these embodiments, the automation assistant 115 may additionally or alternatively utilize the ASR engine 130A2, NLU engine 140A2, LLM engine 150A2, and TTS engine 160A2 of the natural conversation system 120.
[0042] Each of these engines can be configured to perform one or more functions. For example, ASR engines 130A1 and / or 130A2 can use streaming ASR models (e.g., recurrent neural network (RNN) models, transformer models, and / or any other type of ML model capable of performing ASR) stored in the machine learning (ML) model database 115A to process the audio data stream captured from spoken utterances and generated by the microphone of client device 110 to generate an ASR output stream. Notably, the streaming ASR model can be used to generate the ASR output stream while generating the audio data stream. Furthermore, NLU engines 140A1 and / or 140A2 can use NLU models (e.g., long short-term memory (LSTM), gated recurrent unit (GRU), and / or any other type of RNN or other ML model capable of performing NLU) and / or syntax-based rules stored in the ML model database 115A to process the ASR output stream to generate an NLU output stream. Additionally, the automation assistant 115 can cause the NLU output to be processed to generate a fulfillment data stream. For example, the automation assistant 115 can send one or more structured requests to one or more first-party (1P) systems 191 and / or to one or more third-party (3P) systems 192 via one or more networks 199 (or one or more application programming interfaces (APIs)), and receive performance data from one or more 1P systems 191 and / or 3P systems 192 to generate a performance data stream. The one or more structured requests may include, for example, NLU data included in the performance data stream. The performance data stream may correspond to, for example, a set of assistant outputs predicted to be in response to assistant queries included in spoken utterances captured in an audio data stream processed by ASR engines 130A1 and / or 130A2.
[0043] Furthermore, LLM engines 150A1 and / or 150A2 can process sets of assistant outputs predicted in response to assistant queries included in spoken utterances captured in audio data streams processed by ASR engines 130A1 and / or 130A2. As described herein (e.g., see references...). Figures 2 to 6 In some implementations, LLM engines 150A1 and / or 150A2 can use one or more LLM outputs to modify the assistant output set to generate a modified assistant output set. In some versions of those implementations (e.g., and as per [reference to...]), Figure 3 As described, the automation assistant 115 can cause one or more LLM outputs to be generated offline (e.g., not in response to spoken utterances received during a conversational session) and subsequently used online (e.g., when spoken utterances are received during a conversational session) to generate a modified set of assistant outputs. In additional or alternative versions of those implementations (e.g., and as per [the description]... Figure 4 and Figure 5 As described, the automated assistant 115 can generate one or more LLM outputs online (e.g., when uttered utterances are received during a conversational session). In these embodiments, one or more LLM outputs can be generated based on processing an assistant output set (e.g., fulfilling a data stream) using one or more LLMs (e.g., one or more transformer models, such as Meena, RNN, and / or any other LLM) stored in a model database 115A, the context of the conversational session in which the uttered utterances are received (e.g., based on one or more contextual signals stored in a context database 110A), the identified text corresponding to an assistant query included in the uttered utterances, and / or other information that the automated assistant 115 can utilize when generating one or more LLM outputs. Each of the one or more LLM outputs can include, for example, a probability distribution over a sequence of one or more words and / or phrases across one or more vocabularies, and one or more words and / or phrases in that sequence can be selected as one or more LLM outputs based on the probability distribution. In various embodiments, one or more of the LLM outputs can be stored in an LLM output database 150A for subsequent use to modify one or more assistant outputs included in the assistant output set.
[0044] Furthermore, in some implementations, TTS engines 160A1 and / or 160A2 may use TTS models stored in ML model database 115A to process text data (e.g., text defined by automation assistant 115) to generate synthetic speech audio data including computer-generated synthetic speech. The text data may correspond to, for example, one or more assistant outputs from a set of assistant outputs included in the fulfillment data stream, one or more modified assistant outputs from a set of modified assistant outputs, and / or any other text data described herein. It is noteworthy that the ML models stored in ML model database 115A may be on-device ML models stored locally at client device 110, or shared ML models accessible to both client device 110 and / or remote systems when natural conversation system 120 is not implemented locally at client device 110. In additional or alternative implementations, audio data corresponding to one or more assistant outputs from the set of assistant outputs included in the fulfillment data stream, one or more modified assistant outputs from the modified set of assistant outputs, and / or any other text data described herein may be stored in memory or in one or more databases accessible by the client device 110, such that the automation assistant does not need to use the TTS engine 160A1 and / or 160A2 to generate any synthesized speech audio data when making the audio data available for audible presentation to the user.
[0045] In various implementations, the ASR output stream may include, for example: an ASR hypothesis stream (e.g., term hypotheses and / or transcription hypotheses) predicted to correspond to spoken utterances captured in the audio data stream; one or more corresponding prediction values (e.g., probabilities, log-likelihoods, and / or other values) for each ASR hypothesis; multiple phonemes predicted to correspond to spoken utterances captured in the audio data stream; and / or other ASR outputs. In some versions of those implementations, ASR engines 130A1 and / or 130A2 may select one or more ASR hypotheses as the recognized text corresponding to the spoken utterances (e.g., based on the corresponding prediction values).
[0046] In various implementations, the NLU output stream may include, for example, an annotated identified text stream, which includes one or more annotations for one or more (e.g., all) terms of the identified text. For example, NLU engines 140A1 and / or 140A2 may include a part-of-speech tagger (not depicted) configured to annotate terms with their grammatical roles. Additionally or alternatively, NLU engines 140A1 and / or 140A2 may include entity taggers (not depicted) configured to annotate entity references in one or more segments of the identified text, such as references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, places (real and fictional), etc. In some implementations, data about entities may be stored in one or more databases, such as in a knowledge graph (not depicted). In some implementations, the knowledge graph may include nodes representing known entities (and in some cases, entity attributes), and edges connecting nodes and representing relationships between entities. Entity taggers can annotate entity references at a high-granularity level (e.g., to enable the identification of all references to entity classes such as people) and / or a low-granularity level (e.g., to enable the identification of all references to a specific entity such as a particular person). Entity taggers can rely on the content of natural language input to resolve specific entities and / or can optionally communicate with knowledge graphs or other entity databases to resolve specific entities.
[0047] Additionally or alternatively, NLU engines 140A1 and / or 140A2 may include a coreference resolver (not depicted) configured to group or “cluster” references to the same entity based on one or more contextual clues. For example, based on a mention of “theatretickets” in a client device notification rendered immediately before receiving the input “buy them”, the coreference resolver could be used to resolve the term “them” to “buy theatre tickets” in the natural language input “buy them”. In some implementations, one or more components of NLU engines 140A1 and / or 140A2 may rely on annotations from one or more other components of NLU engines 140A1 and / or 140A2. For example, in some implementations, an entity annotator may rely on annotations from the coreference resolver to annotate all references to a particular entity. Additionally, for example, in some implementations, the coreference resolver may rely on annotations from the entity tagger to cluster references to the same entity.
[0048] Although Figure 1 This description pertains to a single client device with a single user; however, it should be understood that this is for illustrative purposes and does not imply limitation. For example, one or more additional client devices belonging to the user may also implement the techniques described herein. For instance, client device 110, one or more additional client devices, and / or any other computing devices belonging to the user can form an ecosystem of devices that can employ the techniques described herein. These additional client devices and / or computing devices may communicate with client device 110 (e.g., via network 199). As another example, a given client device may be used by multiple users in a shared setup (e.g., a group of users, a home).
[0049] As described herein, the automation assistant 115 can determine whether to use one or more LLM outputs to modify the assistant response set and / or determine one or more modified assistant output sets based on one or more LLM outputs. The automation assistant 115 can utilize the natural conversation system 120 to make these determinations. In various embodiments and as described herein... Figure 1 The natural conversation system 120 described herein may additionally or alternatively include an offline output modification engine 170, an online output modification engine 180, and / or a ranking engine 190. The offline output modification engine 170 may include, for example, an assistant activity engine 171 and an indexing engine 172. Furthermore, the online output modification engine 180 may include, for example, an assistant query engine 181 and an assistant personality engine 182. (See reference...) Figures 2 to 5 These various engines of the Natural Conversation System 120 will be described in more detail.
[0050] Turn now Figure 2 This describes an example processing flow 200 for generating assistant output using LLM. Figure 1 Audio data streams 201 generated by one or more microphones of client device 110 can be processed by ASR engines 130A1 and / or 130A2 to generate a stream of ASR output 203. Furthermore, ASR output 203 can be processed by NLU engines 140A1 and / or 140A2 to generate a stream of NLU output 204. In some embodiments, NLU engines 140A1 and / or 140A2 can process the context 202 of a conversational session between the user of client device 110 and an automation assistant 115 that is at least partially executed at the user's client device 110. In some versions of those embodiments, the context 202 of the conversational session can be determined based on one or more contextual signals generated by client device 110 (e.g., time of day, day of week, location of client device 110, ambient noise detected in the environment of client device 110, and / or other contextual signals generated by client device 110). In additional or alternative versions of those implementations, the context 202 of a dialogue session can be determined based on one or more contextual signals stored in a context database 110A accessible at the client device 110 (e.g., user profile data, software application data, environmental data about the known environment of the user of the client device 110, dialogue history of the ongoing dialogue session between the user and the automation assistant 115 and / or historical dialogue history of one or more previous dialogue sessions between the user and the automation assistant 115, and / or other contextual data stored in the context database 110A). Furthermore, the stream of NLU output 204 can be processed by one or more 1P systems 191 and / or 3P systems to generate a performance data stream including a set of one or more assistant outputs 205, wherein each assistant output in the set of one or more assistant outputs 205 is predicted in response to spoken utterance captured in the audio data stream 201.
[0051] Typically, in a turn-based conversational session that does not utilize LLM, ranking engine 190 may process one or more sets of assistant outputs 205 to rank each of the one or more assistant outputs included in the one or more sets of assistant outputs 205 according to one or more ranking criteria, and automation assistant 115 may select one or more given assistant outputs 207 from the one or more sets of assistant outputs 205 in response to receiving spoken utterances to provide for presentation to the user of client device 110. In some embodiments, the selected one or more given assistant outputs 207 may be processed by TTS engines 160A1 and / or 160A2 to generate synthesized speech audio data, which includes synthesized speech corresponding to the selected one or more given assistant outputs 207, and rendering engine 112 may cause the synthesized speech audio data to be audibly rendered by the speaker of client device 110 for audible presentation to the user of client device 110. In additional or alternative embodiments, rendering engine 112 may cause text data corresponding to the selected one or more given assistant outputs 207 to be visually rendered by the display of client device 110 for visual presentation to the user of client device 110.
[0052] However, when using the claimed technology, the automation assistant 115 can also cause one or more sets of assistant outputs 205 to be processed by LLM engines 150A1 and / or 150A2 to generate one or more modified sets of assistant outputs 206. In some implementations, one or more LLM outputs can be generated offline using an offline output modification engine 170 (e.g., before receiving audio data streams 201 generated by one or more microphones of client device 110), and one or more LLM outputs can be stored in an LLM output database 150A. (See also: Regarding...) Figure 3 As described, one or more LLM outputs can be indexed in the LLM output database 150A prior to the corresponding assistant query and / or the corresponding context of the corresponding dialogue session in which the corresponding assistant query is received. Furthermore, LLM engines 150A1 and / or 150A2 can determine that an assistant query included in spoken utterances captured in audio data stream 201 matches a corresponding assistant query, and / or that the context 202 of the dialogue session in which the assistant query is received matches the corresponding context of the corresponding dialogue session in which the corresponding assistant query is received. LLM engines 150A1 and / or 150A2 can obtain one or more LLM outputs indexed by the corresponding assistant query and / or the corresponding context matching the assistant query to modify one or more sets of assistant outputs 205. Additionally, one or more sets of assistant outputs 205 can be modified based on one or more LLM outputs, thereby generating one or more modified sets of assistant outputs 206.
[0053] In additional or alternative implementations, the online output modification engine 180 can be used to generate one or more LLM outputs online (e.g., in response to receiving an audio data stream 201 generated by one or more microphones of the client device 110). See reference... Figure 4 and Figure 5 The described method can generate one or more LLM outputs based on one or more sets of assistant outputs 205 processed using one or more LLMs stored in the ML model database 115A, the identified text corresponding to assistant queries included in spoken utterances captured in the audio data stream 201 (e.g., included in the stream of ASR output 203), and / or the context 202 of the dialogue session between the user of the client device 110 and the automation assistant 115. Furthermore, one or more sets of assistant outputs 205 can be modified based on one or more LLM outputs to produce one or more modified sets of assistant outputs 206. In other words, in these embodiments, LLM engines 150A1 and / or 150A2 can directly generate one or more modified sets of assistant outputs 206 using one or more LLMs stored in the ML model database 115A.
[0054] In these implementations, and compared to the typical turn-based dialogue session that does not utilize LLM described above, the ranking engine 190 can process one or more sets of assistant outputs 205 and one or more modified sets of assistant outputs 206 to rank each of the assistant outputs included in both the one or more sets of assistant outputs 205 and the one or more modified sets of assistant outputs 206 according to one or more ranking criteria. Therefore, when selecting one or more given assistant outputs 207, the automated assistant 207 can select from the one or more sets of assistant outputs 205 and the one or more modified sets of assistant outputs 206. It is worth noting that the assistant outputs included in the one or more modified sets of assistant outputs 206 can be generated based on the one or more sets of assistant outputs 205 and convey the same or similar information, also along with the context 202 of the dialogue (e.g., regarding...). Figure 4 The described additional information, along with relevant and / or more natural, fluent and / or more consistent with the characteristics of the automated assistant, conveys the same or similar message, so that one or more given assistant outputs 207 better resonate with the user of the client device 110.
[0055] One or more ranking criteria may include, for example: one or more predictive metrics (e.g., ASR metrics generated by ASR engines 130A1 and / or 130A2 when generating the stream of ASR output 203, NLU metrics generated by NLU engines 140A1 and / or 140A2 when generating the stream of NLU output 204, and performance metrics generated by one or more of 1P systems 191 and / or 3P systems 192) – which indicate how responsive each assistant output in one or more sets of assistant outputs 205 and one or more modified sets of assistant outputs 206 is to assistant queries included in spoken utterances captured in audio data stream 201; one or more intents included in the stream of NLU outputs 204; metrics derived from a classifier that processes each assistant output in one or more sets of assistant outputs 205 and one or more modified sets of assistant outputs 206 to determine how natural, fluent, and / or consistent with the characteristics of an automated assistant is when provided to a user; and / or other ranking criteria. For example, if the user of client device 110's intent indicates that the user wants a factual answer (e.g., based on providing spoken utterances including the assistant query "Why is the sky blue?"), ranking engine 190 may favor one or more assistant outputs included in one or more sets of assistant outputs 205, because the user may want a simple answer to the assistant query. However, if the user of client device 110's intent indicates that the user provides more open-ended input (e.g., based on providing spoken utterances including the assistant query "What time is it?"), ranking engine 190 may favor one or more assistant outputs included in one or more sets of modified assistant outputs 206, because the user may prefer more conversational aspects.
[0056] Although this article describes voice-based dialogue sessions Figure 1 and Figure 2 However, it should be understood that this is for illustrative purposes and does not imply limitation. Rather, it should be understood that the techniques described herein can be utilized regardless of the user's input modality. For example, in some implementations where the user provides typed input and / or touch input as an assistant query, the automation assistant 115 can use NLU engines 140A1 and / or 140A2 to process the typed input to generate a stream of NLU outputs 204 (e.g., skipping the processing of audio data stream 201), and LLM engines 150A1 and / or 150A2 can utilize text input corresponding to the assistant query (e.g., derived from typed input and / or touch input) to generate one or more modified sets of assistant outputs 206 in the same or similar manner as described above.
[0057] Turn now Figure 3This describes a flowchart illustrating an example method 300 that utilizes a large language model to generate assistant output offline for subsequent online use. For convenience, the implementation details are from [source missing]. Figure 2 The operation of method 300 is described in terms of the system of operation of processing flow 200. The system of method 300 includes a computing device (e.g., Figure 1 Client device 110 Figure 6 Client devices 610 and / or Figure 7 The computing device 710, one or more servers and / or other computing devices, and one or more processors, memories and / or other components. Furthermore, although the operations of method 300 are shown in a specific order, this does not imply limitation. One or more operations may be reordered, omitted and / or added.
[0058] At box 352, the system obtains multiple assistant queries pointing to the automated assistant and the corresponding context of the previous dialogue session for each of the multiple assistant queries. For example, the system can enable... Figure 1 and Figure 2 The offline output modification engine of the assistant activity engine 171 obtains multiple assistant queries and corresponding contexts from previous dialogue sessions, where, for example, Figure 1 The assistant activity database 170A depicted receives multiple assistant queries. In some implementations, the multiple assistant queries and the corresponding context of the previous conversation session in which the multiple assistant queries were received may be limited to the user of the client device (e.g., Figure 1 Those associated with the users of the client device 110. In other embodiments, the multiple assistant queries and the corresponding context of the previous conversational session in which the multiple assistant queries were received may be limited to those associated with the multiple users of the respective client devices (e.g., it may include or may not include). Figure 1 (User of client device 110).
[0059] At box 354, the system uses one or more LLMs to process a given assistant query among multiple assistant queries to generate one or more corresponding LLM outputs, wherein each of the one or more corresponding LLM outputs is predicted to be in response to the given assistant query. Each of the one or more corresponding LLM outputs may include, for example, a probability distribution over a sequence of one or more words and / or phrases across one or more vocabularies, and the one or more words and / or phrases in the sequence may be selected as one or more corresponding LLM outputs based on the probability distribution. In various implementations, when generating one or more corresponding LLM outputs for a given assistant query, the system may use one or more LLMs together with the assistant query to process the corresponding context of the previous dialogue session in which the given assistant query was received, and / or the set of assistant outputs predicted to be in response to the given assistant query (e.g., based on using references...). Figure 2 The described ASR engine 130A1 and / or 130A2, NLU engine 140A1 and / or 140A2, and 1P system 191 and / or 3P system 192 are described as a set of assistant outputs generated by processing audio data corresponding to a given assistant query. In some embodiments, the system can process recognized text corresponding to a given assistant query, while in additional or alternative embodiments, the system can process audio data capturing spoken utterances including a given assistant query. In some embodiments, the system can make the LLM engine 150A1 available to the user (e.g., Figure 1 In some implementations, the system allows the LLM engine 150A2 to process a given assistant query using one or more LLMs local to the user's client device (e.g., at a remote server). As described herein, the outputs of one or more corresponding LLMs can reflect more natural conversational output than typical assistant outputs that an automated assistant can provide. This allows the automated assistant to drive the conversation more smoothly, making assistant outputs modified based on one or more corresponding LLM outputs more likely to resonate with the user who perceives the modified assistant output.
[0060] In some implementations, in addition to one or more corresponding LLM outputs, one or more LLM models can be used to generate additional assistant queries based on the corresponding context of processing a given assistant query and / or a corresponding previous dialogue session in which the given assistant query was received. For example, when processing a given assistant query and / or the corresponding context of a corresponding previous dialogue session in which the given assistant query was received, one or more LLMs can determine the intent associated with the given assistant query (e.g., based on the use of...). Figure 2The NLU output 204 is generated by the NLU engines 140A1 and / or 140A2 in the NLU engine. Furthermore, one or more LLMs can identify at least one related intent associated with the intent associated with the given assistant query (e.g., based on a mapping of intent to at least one related intent in a database or memory accessible to the client device 110 and / or based on rules defined using one or more machine learning (ML) models or heuristics). Additionally, one or more LLMs can generate additional assistant queries based on at least one related intent. For example, suppose the assistant query indicates that the user has not yet eaten dinner (e.g., a given assistant query “I’m feeling pretty hungry” received at the user’s actual residence in the evening, as indicated by the corresponding context of a previous conversational session associated with the user’s intent indicating that he / she wants to eat). In this example, additional helper queries could correspond to, for example, “what types of cuisine has the user indicated he / she prefers?” (e.g., reflecting the intent to find relevant cuisine types associated with the user’s intent to indicate that he / she wants to eat), “what restaurants nearby are open?” (e.g., reflecting the intent to find relevant restaurants associated with the user’s intent to indicate that he / she wants to eat), and / or other additional helper queries.
[0061] In these implementations, additional assistant output can be determined based on processing additional assistant queries. In the example above where the additional assistant query is “what types of cuisine has the user indicated he / she prefers?”, user profile data from one or more 1P systems 191 stored locally at client device 110 can be used to determine that the user indicates he / she prefers Mediterranean and Indian cuisine. Based on the user profile data indicating that the user prefers Mediterranean and Indian cuisine, one or more corresponding LLM outputs can be modified to ask the user if Mediterranean and / or Indian cuisine sounds appetizing (e.g., “how does Mediterranean cuisine or Indiancuisine sound for dinner?”).
[0062] In the example above where the additional assistant query is “what restaurants nearby are open?”, restaurant data from one or more of 1P system 191 and / or 3P system 192 can be used to determine which restaurants are open near the user’s primary residence (and optionally limited to restaurants serving Mediterranean and Indian cuisine based on the user profile data). Based on the results, one or more corresponding LLM outputs can be modified to provide the user with a list of one or more restaurants open near the user’s primary residence (e.g., “Example Mediterranean Restaurant is open until 9:00 PM and Example Indian Restaurant is open until 10:00 PM”). It is worth noting that the additional assistant queries initially generated using the LLM (e.g., “what types of cuisine has the user indicated he / she prefers?” and “what types of cuisine has the user indicated he / she prefers?” in the example above) may not be included in one or more corresponding LLM outputs and therefore may not be presented to the user. Conversely, additional assistant outputs determined based on additional assistant queries (e.g., “how does Mediterraneancuisine or Indian cuisine sound for dinner” and “Example Mediterranean Restaurant is open until 9:00 PM and Example Indian Restaurant is open until 10:00 PM”) can be included in one or more corresponding LLM outputs and thus can be provided to the user.
[0063] In additional or alternative implementations, each of one or more corresponding LLM outputs (and optionally additional assistant outputs determined based on additional assistant queries) can be generated using corresponding parameter sets from multiple dissimilar parameter sets of one or more LLMs. Each parameter in the multiple dissimilar parameter sets can be associated with a dissimilar personality of the automation assistant. In some versions of those implementations, a single LLM can be used to generate one or more corresponding LLM outputs using corresponding parameter sets for each dissimilar personality, while in other versions of those implementations, multiple LLMs can be used to generate one or more corresponding LLM outputs using corresponding parameter sets for each dissimilar personality. For example, a single LLM can be used to: generate a first LLM output using a first set of parameters reflecting a first personality (e.g., the chef personality in the example above, where a given assistant query corresponds to "I'm feeling pretty hungry"); generate a second LLM output using a second set of parameters reflecting a second personality (e.g., the butler personality in the example above, where a given assistant query corresponds to "I'm feeling pretty hungry"); and so on for multiple other dissimilar personalities. Furthermore, for example, a first LLM can be used to generate a first LLM output using a first set of parameters that reflects a first personality (e.g., the chef personality in the example above, where a given assistant query corresponds to "I'm feeling pretty hungry"); a second LLM can be used to generate a second LLM output using a second set of parameters that reflects a second personality (e.g., the butler personality in the example above, where a given assistant query corresponds to "I'm feeling pretty hungry"); and so on for multiple other dissimilar personalities. Thus, when providing corresponding LLM outputs to a user, they can reflect various dynamic contextual personalities via prosodic attributes of the different personalities (e.g., intonation, cadence, pitch, pauses, speech rate, stress, rhythm, etc. of these different personalities). Additionally or alternatively, a user can define one or more personalities to be used by the automated assistant in a continuous manner (e.g., always using the butler personality) and / or contextually (e.g., using the butler personality in the morning and evening, but using a different personality in the afternoon) (e.g., via settings of the automated assistant application associated with the automated assistant described herein).
[0064] It is worth noting that the personalized responses described in this paper not only reflect the prosodic attributes of different personalities, but also the vocabulary of different personalities and / or different speaking styles (e.g., verbose speaking style, concise speaking style, friendly personality, sarcastic personality, etc.). For example, the chef personality mentioned above may have a specific chef vocabulary, such that the probability distribution on one or more word and / or phrase sequences of one or more corresponding LLM outputs generated using a parameter set specific to the chef personality can promote the sequence of words and / or phrases used by the chef to outperform other word and / or phrase sequences of other personalities (e.g., scientist personality, librarian personality, etc.). Therefore, when one or more corresponding LLM outputs are provided to the user, they can reflect various dynamic contextual personalities not only in terms of the prosodic attributes of different personalities, but also in terms of the accurate and realistic vocabulary of different personalities, allowing one or more corresponding LLM outputs to better resonate with the user in different contextual scenarios. Furthermore, it should be understood that the vocabulary and / or speaking style of different personalities can be defined at different levels of granularity. Continuing with the example above, the chef personality could have specific Mediterranean chef terms when querying Mediterranean cuisine based on an additional assistant query associated with Mediterranean cuisine, specific Indian chef terms when querying Indian cuisine based on an additional assistant query associated with Indian cuisine, and so on.
[0065] At box 356, the system, based on a given assistant query and / or the corresponding context of a previous dialogue session for that given assistant query, accesses memory (e.g., at the client device) in the memory that is accessible at the client device. Figure 1 The system indexes one or more corresponding LLM outputs in the LLM output database (150A). For example, the system can enable... Figure 1 and Figure 2The offline output modification engine's indexing engine 172 indexes one or more corresponding LLM outputs in the LLM output database 150A. In some embodiments, the indexing engine 172 may index one or more corresponding LLM outputs based on one or more terms included in a given assistant query and / or one or more context signals included in the corresponding context of a previous dialogue session in which the given assistant query was received. In additional or alternative embodiments, the indexing engine 172 may cause the generation of embeddings of a given assistant query (e.g., word2vec embeddings or any other lower-dimensional representation) and / or embeddings of one or more context signals in the corresponding context of a previous dialogue session in which the given assistant query was received, and one or more of these embeddings may be mapped to an embedding space (e.g., a lower-dimensional space). In these embodiments, one or more corresponding LLM outputs indexed in the LLM output database 150A may then be used by the online output modification engine 180 to modify the assistant output set (e.g., as described below regarding...). Figure 3 (as described in blocks 362, 364, and 366 of method 300). In various embodiments, the system may additionally or alternatively generate one or more assistant outputs (e.g., as per [reference to...]). Figure 2 The described one or more assistant outputs (205) are predicted in response to a given assistant query (but not generated using LLM engines 150A1 and / or 150A2), and one or more corresponding LLM outputs may additionally or alternatively be indexed by one or more assistant outputs.
[0066] In some implementations, and as indicated in box 358, the system may optionally receive user input to review and / or modify one or more corresponding LLM outputs. For example, a human reviewer may analyze one or more corresponding LLM outputs generated using one or more LLM models and modify one or more corresponding LLM outputs by changing one or more terms and / or phrases included in the outputs. Additionally, for example, a human reviewer may reindex, discard, and / or otherwise modify the indexes of one or more corresponding LLM outputs. Therefore, in these implementations, one or more corresponding LLM outputs generated using one or more LLM models can be curated by a human reviewer to ensure the quality of the outputs. Furthermore, any undiscarded, reindexed, and / or curated LLM outputs can be used to modify or retrain the LLM offline.
[0067] At box 360, the system determines whether there are additional assistant queries included in the multiple assistant queries obtained at box 352 that have not yet been processed by one or more LLMs. If, at the iteration of box 360, the system determines that there are additional assistant queries included in the multiple assistant queries obtained at box 352 that have not yet been processed by one or more LLMs, the system returns to box 354 and performs additional iterations of boxes 354 and 356, but with respect to the additional assistants rather than the given assistant queries. These operations can be repeated for each assistant query included in the multiple assistant queries obtained at box 352. In other words, the system can index one or more corresponding LLM outputs for each assistant query and / or the corresponding context of the previous dialogue session in which one or more corresponding LLM outputs are received, before utilizing one or more corresponding LLM outputs.
[0068] If, at the iteration in box 360, the system determines that there are no additional assistant queries included in the multiple assistant queries obtained in box 352 that have not yet been processed using one or more LLMs, the system can proceed to box 362. At box 362, the system can monitor an audio data stream generated by one or more microphones on the client device to determine whether the audio data stream captures spoken utterances from the user of the client device directed to the automation assistant. For example, the system can monitor one or more specific words or phrases included in the audio data stream (e.g., monitoring one or more specific words or phrases that invoke the automation assistant using a hot word detection model). Additionally, optionally, the system can monitor speech directed to the client device, for example, in addition to one or more other signals (e.g., one or more gestures captured by the client device's visual sensors, eye gaze directed at the client device, etc.). If, at the iteration in box 362, the system determines that the audio data stream does not capture spoken utterances from the user of the client device directed to the automation assistant, the system can continue monitoring the audio data stream at box 362. If, at the iteration in box 362, the system determines that the audio data stream captures spoken utterances from the user of the client device directed to the automation assistant, the system can proceed to box 364.
[0069] At box 364, the system determines, based on processing the audio data stream, that the spoken utterance includes the current assistant query corresponding to one of multiple assistant queries and / or that the spoken utterance is received in the current context of the current dialogue session corresponding to the corresponding context of the previous dialogue session for one of the multiple assistant queries. For example, the system may use ASR engines 130A1 and / or 130A2 to process the audio data stream (e.g., Figure 2 The audio data stream 201) is used to generate an ASR output stream (e.g., Figure 2The ASR output stream 203). Furthermore, the system can use NLU engines 140A1 and / or 140A2 to process the ASR output stream to generate an NLU output stream (e.g., Figure 2 The NLU output stream 204). Furthermore, based on the ASR output stream and / or NLU output stream, the system can identify the current assistant query. In some implementations, the system can also determine one or more assistant outputs (e.g., by causing one or more processing ASR output streams and / or NLU output streams in the 1P system 191 and / or 3P system). Figure 2 A set of one or more assistant outputs (205).
[0070] exist Figure 3 In some implementations of method 300, the system may utilize online output modification engine 180 to determine that one or more terms or phrases of the current assistant query correspond to one or more terms or phrases of one of a plurality of assistant queries, for which one or more corresponding LLM outputs are indexed (e.g., using any known technique to determine whether a term or phrase corresponds to one or the other, such as exact matching, soft matching, edit distance, phonetic similarity, embedding, etc.). In response to determining that one or more terms or phrases of the current assistant query correspond to one or more terms or phrases of one of a plurality of assistant queries, the system may (e.g., from LLM output database 150A) obtain one or more corresponding LLM outputs indexed at the iteration of box 356 and associated with one of the plurality of assistant queries. For example, if the current query includes the phrase “I'm hungry,” the system may obtain one or more corresponding LLM outputs generated in the above example describing the given assistant query. For example, the system may determine that both the current query and the given assistant query include the phrase “I'm hungry” based on comparing the edit distance between terms of the current query and terms of the given assistant query. Additionally, for example, the system can generate an embedding of the current assistant query and map the current query's embedding to the embedding space described above with respect to box 356. Furthermore, the system can determine that the current assistant query corresponds to a given assistant query based on the distance between the generated embedding of the current query in the embedding space and the previously generated embedding of the given assistant query satisfying a distance threshold.
[0071] exist Figure 3In an additional or alternative implementation of method 300, the system may utilize online output modification engine 180 to determine that one or more context signals detected when the current assistant query is received correspond to one or more corresponding context signals when one of a plurality of assistant queries is received (e.g., received on the same day of the week, at the same time of day, at the same location, with the same ambient noise present in the client device's environment, received in a specific sequence of spoken utterances during a conversational session, etc.). In response to determining that one or more context signals associated with the current assistant query correspond to one or more context signals of one of a plurality of assistant queries, the system may (e.g., from LLM output database 150A) obtain one or more corresponding LLM outputs indexed at the iteration of box 356 and associated with one of the plurality of assistant queries. For example, if the current query is received at night at the user's primary residence, the system may obtain one or more corresponding LLM outputs generated in the above example describing a given assistant query. For example, the system may determine that both the described current query and the given assistant query are associated with a time context signal of "night" and a location context signal of "primary residence". Additionally, for example, the system can generate embeddings of one or more context signals associated with the current assistant query and map these embeddings to the embedding space described above with respect to box 356. Furthermore, the system can determine that one or more context signals associated with the current assistant query correspond to one or more context signals associated with the given assistant query based on a distance threshold satisfied between the generated embeddings of one or more context signals associated with the current query and previously generated embeddings of one or more context signals associated with the given assistant query.
[0072] It is worth noting that the system can utilize one or both of the current assistant query and the context of the dialogue session in which the current assistant query is received (e.g., one or more detected context signals) to determine one or more corresponding LLM outputs that will be used to generate one or more current assistant outputs to be provided to the user in response to the current assistant query. In various implementations, the system can additionally or alternatively utilize one or more assistant outputs generated for the current assistant query to determine one or more corresponding LLM outputs to be used in generating one or more current assistant outputs to be provided to the user in response to the current assistant query. For example, the system can additionally or alternatively generate one or more assistant outputs (e.g., as per [reference to [reference to a specific function]). Figure 2The described one or more assistant outputs 205, which are predicted in response to the current assistant query (but not generated using LLM engines 150A1 and / or 150A2), and the various techniques described above are used to determine that the one or more assistant outputs predicted in response to the current assistant query correspond to one or more previously generated assistant outputs for one of a plurality of assistant queries.
[0073] At box 366, the system enables the automation assistant to generate one or more current assistant outputs to be provided to the user of the client device using one or more corresponding LLM outputs. For example, the system can rank one or more assistant outputs according to one or more ranking criteria (e.g., regarding...). Figure 2 The described one or more assistant outputs 205) and one or more corresponding LLM outputs (e.g., as about Figure 2 The system ranks one or more modified assistant outputs (206) as described. Furthermore, the system can select one or more current assistant outputs from one or more assistant outputs and one or more corresponding LLM outputs. Additionally, the system can cause one or more current assistant outputs to be visually and / or audibly rendered for presentation to the user of the client device.
[0074] Although the method involves generating one or more corresponding LLM outputs offline (e.g., generating corresponding LLM outputs by utilizing offline output modification engine 170 to use one or more LLMs, and indexing one or more corresponding LLM outputs in LLM output database 150A), and subsequently utilizing one or more corresponding LLM outputs online (e.g., utilizing online output modification engine 180 to determine utilization from LLM output database 150A based on the current assistant query), the method is described as follows: Figure 3 The implementation of method 300 is described, but it should be understood that it is for illustrative purposes only and does not imply limitation. For example, see the following regarding... Figure 4 and Figure 5 As described, in the preferred embodiment, the online output modification engine 180 may additionally or alternatively utilize the LLM engines 150A1 and / or 150A2 online.
[0075] Turn now Figure 4 The flowchart illustrates an example method 400 that uses a large language model to generate assistant output based on an assistant query. For convenience, the implementation is referenced from... Figure 2 The operation of method 400 is described using the system of the processing flow 200. The system of method 400 includes a computing device (e.g., Figure 1 Client device 110 Figure 6 Client devices 610 and / or Figure 7The computing device 710, one or more servers and / or other computing devices, includes one or more processors, memories and / or other components. Furthermore, although the operations of method 400 are shown in a specific order, this does not imply limitation. One or more operations may be reordered, omitted and / or added.
[0076] At box 452, the system receives an audio data stream capturing a user's spoken utterance, which includes an assistant query directed to an automated assistant, and this spoken utterance is received during a conversational session between the user and the automated assistant. In some implementations, the system may process the audio data stream only in response to determining that one or more conditions are met to identify its captured assistant query. For example, the system may monitor one or more specific words or phrases included in the audio data stream (e.g., monitoring one or more specific words or phrases that invoke the automated assistant using a hot word detection model). Furthermore, for example, optionally, in addition to one or more other signals (e.g., one or more gestures captured by the client device's visual sensors, eye gaze directed at the client device, etc.), the system may also monitor speech directed at the client device.
[0077] At box 454, the system determines an assistant output set based on processing the audio data stream, each of the assistant outputs included in the set being in response to an assistant query included in the spoken utterance. For example, the system may use ASR engines 130A1 and / or 130A2 to process the audio data stream (e.g., audio data stream 201) to generate an ASR output stream (e.g., ASR output 203). Furthermore, the system may use NLU engines 140A1 and / or 140A2 to process the ASR output stream (e.g., audio data stream 201) to generate an NLU output stream (e.g., a stream of NLU output 204). Additionally, the system may cause one or more 1P systems 191 and / or 3P systems 192 to process the NLU output stream (e.g., a stream of NLU output 204) to generate an assistant output set 205 (e.g., an assistant output set). It is worth noting that the set of assistant outputs can correspond to one or more candidate assistant outputs such that the automated assistant can consider these one or more candidate assistant outputs in response to spoken utterances without the techniques described herein (i.e., techniques that do not utilize LLM Engine 150A1 and / or 150A2 in the modified assistant outputs as described herein).
[0078] At box 456, the system processes the assistant output set and the context of the dialogue session to: (1) generate a modified assistant output set using one or more LLM outputs, each of which is determined at least based on the context of the dialogue session and / or one or more assistant outputs included in the assistant output set; and (2) generate additional assistant queries related to the spoken utterance based on at least a portion of the context of the dialogue session and at least a portion of the assistant queries included in the spoken utterance. In various embodiments, each LLM output may be further determined based on assistant queries included in the spoken utterance captured in the audio data stream. In some embodiments, when generating the modified assistant output set, one or more LLM outputs may have already been generated offline (e.g., before the spoken utterance is received and the offline output modification engine 170 is used, as referenced above). Figure 3 In these embodiments, the system can determine that an assistant query included in the spoken utterance captured in the audio data stream corresponds to a previous assistant query for which one or more LLM outputs have already been generated, the context of the dialogue session corresponds to the previous context of the previous dialogue session in which the previous assistant query was received, and / or one or more assistant outputs included in the set of assistant outputs determined at box 454 correspond to one or more previous assistant outputs determined based on the previous query. Furthermore, the system can obtain information based on... Figure 3 Method 300 describes one or more LLM outputs that are indexed for a previous assistant query, a previous context, and / or one or more previous assistant outputs that correspond to the assistant query, context, and / or one or more assistant outputs included in the assistant output set (e.g., using an online output modification engine 180), and utilizes one or more LLM outputs as the modified assistant output set.
[0079] In additional or alternative implementations, when generating a modified assistant output set, the system may process at least the context of the dialogue session and / or one or more assistant outputs included in the assistant output set online (e.g., in response to receiving spoken utterances and using online output modification engine 180) to generate one or more LLM outputs. For example, the system may cause LLM engines 150A1 and / or 150A2 to use one or more LLMs to process the context of the dialogue session, assistant queries, and / or one or more assistant outputs included in the assistant output set to generate a modified assistant output set. Reference can be made to the above regarding generating one or more LLM outputs offline. Figure 3Method 300, box 354, describes the same or similar approach but in response to receiving spoken utterances at a client device to generate one or more LLM outputs online. For example, a system can use one or more LLMs to process one or more assistant outputs included in an assistant output set to generate one or more personalized responses for each of the one or more assistant outputs included in the assistant output set, as described above regarding... Figure 3 The method 300 is described in box 354. In other words, each assistant output included in the set of assistant outputs may have a limited vocabulary and a consistent personality in terms of the prosodic attributes associated with each assistant output. However, in processing each assistant output included in the set of assistant outputs to generate a modified set of assistant outputs, each modified assistant output may have a much larger vocabulary and greater diversity in terms of the prosodic attributes associated with each modified assistant output, depending on the use of one or more LLMs to generate one or more modified assistant outputs. As a result, each modified assistant output may correspond to a context-sensitive assistant output that resonates better with users engaging in conversational dialogue with the automated assistant.
[0080] Similarly, in some implementations, when generating the additional assistant query, the additional assistant may have already been generated offline (e.g., before receiving the spoken utterance and modifying the offline output engine 170, as mentioned above). Figure 3 In these embodiments, and similar to those described above regarding obtaining one or more LLM outputs previously generated offline, the system can obtain additional assistant queries indexed based on previous assistant queries, previous contexts, and / or one or more previous assistant outputs respectively corresponding to assistant queries, contexts, and / or included in the assistant output set as described with respect to method 300, and utilize the previously generated additional assistant queries as additional assistant queries.
[0081] Furthermore, similarly, in additional or alternative implementations, when generating additional assistant queries, the system can process at least the context of the dialogue session and / or one or more assistant outputs included in the assistant output set online (e.g., in response to receiving spoken utterances and using online output modification engine 180) to generate additional assistant queries. For example, the system can cause LLM engines 150A1 and / or 150A2 to use one or more LLMs to process the context of the dialogue session, assistant queries, and / or one or more assistant outputs included in the assistant output set to generate additional assistant queries. This can be referenced above regarding generating additional assistant queries offline. Figure 3The method described in block 354 of method 300 is in the same or similar manner but generates additional assistant queries online in response to receiving spoken utterances at a client device. In some implementations, the one or more LLMs described herein may have multiple disparate layers dedicated to performing certain functions. For example, one or more first layers of the one or more LLMs may be used to generate the personalized responses described herein, and one or more second layers of the one or more LLMs may be used to generate the additional assistant queries described herein. In additional or alternative implementations, the one or more LLMs may communicate with one or more additional layers not included in the one or more LLMs when generating the additional assistant queries described herein. For example, one or more layers of the one or more LLMs may be used to generate the personalized responses described herein, and one or more additional layers of another ML model communicating with the one or more LLMs may be used to generate the additional assistant queries described herein. Reference will be made below. Figure 6 Provide more detailed, non-restricted examples of personalized replies and additional assistant queries.
[0082] At box 458, the system determines the additional assistant output in response to the additional assistant query based on the additional assistant query. In some implementations, the system may cause the additional assistant query to be generated by one or more of 1P system 191 and / or 3P system 192 regarding... Figure 2 The processing assistant query is handled in the same or similar manner to generate additional assistant outputs. In some implementations, the additional assistant output may be a single additional assistant output, while in other implementations, the additional assistant output may be included in a set of additional assistant outputs (e.g., similar to...). Figure 2 (The additional assistant output set 205). In additional or alternative implementations, additional assistant queries may be directly mapped to additional assistant outputs based on, for example, the user profile data of the user providing the spoken utterance and / or any other data accessible to the automated assistant.
[0083] At box 460, the system processes the modified set of assistant outputs based on the additional assistant outputs in response to the additional assistant query to generate an additional modified set of assistant outputs. In some implementations, the system may prepend or append the additional assistant outputs for one or more modified assistant outputs to each of the one or more assistant outputs included in the modified set of assistant outputs generated at box 456. In additional or alternative implementations and as indicated at box 460A, the system may process the additional assistant outputs and the context of the dialogue session to generate an additional modified set of assistant outputs using one or more LLM outputs used at box 456 and / or one or more additional LLM outputs generated based on at least a portion of the context of the dialogue session and at least a portion of the additional assistant outputs, in addition to the one or more LLM outputs used at box 456. Additional modified assistant output sets can be generated in the same or similar manner as described above regarding the generation of modified assistant output sets, but based on additional assistant outputs instead of the assistant output set (e.g., using one or more LLM outputs generated offline and / or online using LLM engines 150A1 and / or 150A2).
[0084] At box 462, the system causes a given modified assistant output from the modified assistant output set and / or a given additional modified assistant output from the additional modified assistant output set to be presented to the user. In some implementations, the system may cause the ranking engine 190 to rank each of the one or more modified assistant outputs included in the modified assistant output set (and optionally, each of the one or more assistant outputs included in the assistant output set) according to one or more ranking criteria, and select a given modified assistant output from the modified assistant output set (or select a given assistant output from the assistant output set). Furthermore, the system may also cause the ranking engine 190 to rank each of the one or more additional modified assistant outputs included in the additional modified assistant output set (and optionally, additional assistant outputs) according to one or more ranking criteria, and select a given additional modified assistant output from the additional modified assistant output set (or select an additional assistant output as a given additional assistant output). In these implementations, the system can combine a given modified assistant output and a given additional assistant output, and make the given modified assistant output and the given additional assistant output available for visual and / or audible presentation to a user of a client device engaging in a conversation with the automated assistant.
[0085] Turn now Figure 5The document describes a flowchart of an example method 500 that utilizes a large language model to generate assistant output based on personalized assistant responses. For convenience, the implementation details are provided below. Figure 2 The operation of method 500 is described in terms of the system of operation of processing flow 200. The system of method 500 includes a computing device (e.g., Figure 1 Client device 110 Figure 6 Client devices 610 and / or Figure 7 The computing device 710, one or more servers and / or other computing devices, includes one or more processors, memories and / or other components. Furthermore, although the operations of method 500 are shown in a specific order, this does not imply limitation. One or more operations may be reordered, omitted and / or added.
[0086] At box 552, the system receives an audio data stream capturing a user's spoken utterance, which includes an assistant query directed to the automated assistant, and this spoken utterance is received during a conversational session between the user and the automated assistant. At box 554, the system determines a set of assistant outputs based on the processed audio data stream, including each assistant output in this set in response to an assistant query included in the spoken utterance. This can be discussed in relation to... Figure 4 Methods 400, boxes 452, and 454 describe the same or similar ways of execution. Figure 5 Method 500 refers to the operations in boxes 552 and 554.
[0087] At box 556, the system determines whether to modify one or more assistant outputs included in the assistant output set. The system may determine whether to modify one or more assistant outputs based on, for example, the user's intent to provide spoken utterances (e.g., included in the stream of NLU output 204), one or more assistant outputs included in the assistant output set (e.g., the assistant output set 205), one or more computational costs associated with modifying one or more assistant outputs included in the assistant output set (e.g., battery consumption, processor consumption, latency, etc.), the duration of interaction with the automated assistant, and / or other considerations. For example, if the user's intent indicates that the user providing the spoken utterances expects a quick and / or factual response (e.g., "Why is the sky blue?", "What's the weather?", "What time is it?", etc.), in some cases, the system may determine not to modify one or more assistant outputs to reduce latency and computational resource consumption when providing content in response to the spoken utterances. Additionally, for example, if the user's client device is in power-saving mode, the system may determine not to modify one or more assistant outputs to conserve battery power. Additionally, for example, if a user has been engaged in a conversation for a duration exceeding a threshold duration (e.g., 30 seconds, 1 minute, etc.), the system may determine not to modify one or more assistant outputs in an attempt to end the conversation in a faster and more efficient manner.
[0088] If, at the iteration in box 556, the system determines that one or more assistant outputs included in the assistant output set should not be modified, the system can proceed to box 558. At box 558, the system causes a given assistant output from the assistant output set to be provided to the user. For example, the system may cause ranking engine 190 to rank each assistant output included in the assistant output set according to one or more ranking criteria, and select a given assistant output to be provided to the user for visual and / or audible presentation based on the ranking.
[0089] If, at the iteration in box 556, the system determines that it needs to modify one or more assistant outputs included in the assistant output set, the system can proceed to box 560. At box 560, the system processes the assistant output set and the context of the dialogue session to generate a modified assistant output set using one or more LLM outputs, each of which is determined based on the context of the dialogue session and / or the one or more assistant outputs included in the assistant output set, and each of the one or more LLM outputs reflects a corresponding personality of the automated assistant from among several disparate personalities. (See also: Regarding...) Figure 3As described in block 354 of method 300, one or more LLM outputs can be generated using various disparate parameters to reflect the different personalities of the automated assistant. In various implementations, each LLM output can be further determined based on assistant queries included in spoken utterances captured in the audio data stream. In some implementations, the system can process the assistant output set and the context of the dialogue session (and optionally assistant queries) to generate a modified assistant output set using one or more of the LLM outputs previously generated offline as described herein, while in additional or alternative implementations, the system can process the assistant output set and the context of the dialogue session (and optionally assistant queries) to generate a modified assistant output set online as described herein. As mentioned above regarding... Figure 4 Method 400 indicates the following regarding Figure 6 Provide more detailed, non-restrictive examples of personalized responses and additional assistant queries.
[0090] At box 562, the system causes a given assistant output from the assistant output set to be presented to the user. In some implementations, the system may cause the ranking engine 190 to rank each of one or more modified assistant outputs included in the modified assistant output set (and optionally, each of one or more assistant outputs included in the assistant output set) according to one or more ranking criteria, and to select a given modified assistant output from the modified assistant output set (or select a given assistant output from the assistant output set). Furthermore, the system may cause a given modified assistant output to be presented to the user visually and / or audibly.
[0091] Although there is no description of generating any additional helper queries. Figure 5 However, it should be understood that this is for illustrative purposes and does not imply limitation. On the contrary, it should be understood that... Figure 4 Method 400 generates additional assistant queries based on the determination that context-sensitive additional assistant outputs exist, which may be provided to facilitate conversational dialogue and to provide a more natural conversational experience for the user. This ensures that any assistant outputs presented to the user better reflect human-to-human conversations and that the conversation between the user and the automated assistant resonates more with the user. Furthermore, while there is no description of whether to modify the set of assistant outputs... Figure 3 and Figure 4 However, it should be understood that this is for illustrative purposes and does not imply limitation. On the contrary, it should be understood that... Figure 3 , Figure 4 and Figure 5 In any of the example methods, a determination is made as to whether to use one or more LLMs to trigger modifications to the assistant response set.
[0092] Turn now Figure 6 This describes a non-limiting example of a dialogue session between a user and an automated assistant, where the automated assistant utilizes one or more LLMs to generate assistant outputs. As described herein, in some implementations, the automated assistant may utilize one or more LLM outputs previously generated offline to generate a modified set of assistant outputs (e.g., as described above regarding...). Figure 3 (as described in method 300). For example, the automation assistant may determine that a previous assistant query for which one or more LLM outputs have been generated corresponds to an assistant query included in a spoken utterance, a previous context of a previous conversational session corresponds to the context of a conversational session between the user and the automation assistant in which the spoken utterance was received, and / or one or more previous assistant outputs correspond to one or more assistant outputs included in a set of assistant outputs for an assistant query included in a spoken utterance. Furthermore, the automation assistant may obtain one or more LLM outputs indexed based on the previous assistant query, the previous context, and / or one or more previous assistant outputs included in the set of previous assistant outputs (e.g., in an LLM output database 150A), and utilize one or more LLM outputs as a modified set of assistant outputs. In an additional or alternative implementation, the automation assistant may use one or more LLMs to process the assistant query, the context of the conversational session, and / or one or more assistant outputs included in the set of assistant outputs to generate one or more LLM outputs to be used online as a modified set of assistant outputs (e.g., as described above regarding...). Figure 4 Method 400 and Figure 5 Method 500 described. Therefore, providing Figure 6 The following are non-limiting examples to illustrate how the use of LLMs based on the techniques described herein can lead to improved natural conversations between users and automated assistants.
[0093] Client device 610 (e.g., Figure 1An instance of client device 610 may include various user interface components, including, for example: a microphone for generating audio data based on spoken words and / or other audible input; a speaker 680 for audibly rendering synthesized speech and / or other audible output; and / or a display 680 for visually rendering visual output. Furthermore, the display 680 of client device 610 may include various system interface elements 681, 682, and 683 (e.g., hardware and / or software interface elements) that can be interacted with by a user of client device 610 to cause client device 610 to perform one or more actions. The display 680 of the client device 610 enables the user to interact with the content rendered on the display 680 via touch input (e.g., by directing user input to the display 680 or portions thereof (e.g., to a text input box (not depicted), to a keyboard (not depicted), or to other portions of the display 680)) and / or via verbal input (e.g., by selecting a microphone interface element 684 at the client device 610—or simply by speaking without having to select the microphone interface element 684 (i.e., the automation assistant can monitor one or more words or phrases, gestures, gaze, mouth movements, lip movements, and / or other conditions activating verbal input)). Although Figure 6 The client device 610 depicted is a mobile phone, but it should be understood that this is for illustrative purposes only and is not intended to be limiting. For example, the client device 610 can be a stand-alone speaker with a display, a stand-alone speaker without a display, a home automation device, an in-vehicle system, a laptop computer, a desktop computer, and / or any other device capable of performing automated assistants to conduct human-computer dialogue sessions with the user of the client device 610.
[0094] For example, suppose a user of client device 610 provides the spoken words 652, "Hey Assistant, what time is it?". In this example, the automation assistant can cause the ASR engine 130A1 and / or 130A2 to process the audio data of the captured spoken words 652 to generate an ASR output stream. Furthermore, the automation assistant can cause the NLU engine 140A1 and / or 140A2 to process the ASR output stream to generate an NLU output stream. Additionally, the automation assistant can cause one or more of the 1P system 191 and / or 3P system 192 to process the NLU output stream to generate a set of one or more assistant outputs. This set of assistant outputs may include, for example, "8:30 AM," "Good morning, it's 8:30 AM," and / or any other output conveying the current time to the user of client device 610.
[0095] exist Figure 6In the example, it is further assumed that the automation assistant determines to modify one or more assistant outputs included in the assistant output set to generate a modified assistant output set. For example, the automation assistant may determine that the assistant query included in spoken utterance 652 requests the automation assistant to provide the current time to be presented to the user. The automation assistant may determine that an instance of a previous assistant query that requested the automation assistant to provide the current time to be presented to the user, for which one or more LLM outputs were previously generated, corresponds to the one included in the previous assistant query. Figure 6 The assistant query in spoken utterance 652, the previous context of the previous conversation corresponds to Figure 6 The context of the dialogue session between the user and the automated assistant (e.g., the user requests the automated assistant to provide the current time in the morning (and optionally makes the request by initiating a dialogue session), the client device 610 is located in a specific location, and / or other contextual signals), and / or one or more previous assistant outputs corresponding to the input included in the dialogue session. Figure 6 The automated assistant may include one or more assistant outputs in the assistant output set of the assistant query in spoken utterance 652. Furthermore, the automated assistant may obtain one or more LLM outputs indexed based on previous assistant queries, previous context, and / or one or more previous assistant outputs included in the previous assistant output set (e.g., in LLM output database 150A), and utilize one or more LLM outputs as a modified assistant output set. Additionally, for example, the automated assistant may cause one or more LLMs to be used to process assistant queries, the context of the dialogue session, and / or one or more assistant outputs included in the assistant output set, thereby generating one or more LLM outputs online for use as a modified assistant output set.
[0096] exist Figure 6In the example, it is further assumed that the automation assistant determines to provide a modified assistant output 654, “Good morning [User]! It's 8:30 AM. Any fun plans today?”, for presentation to the user, where the modified assistant output 654 is determined based on one or more LLM outputs. The modified assistant output 654 provided to the user is personalized or customized for the context of the user and the conversational session on client device 610, as the modified assistant output 654 greets the user with a context-appropriate greeting (e.g., “Good Morning”) and addresses the user of client device 610 by name (e.g., “[User]”). It is noteworthy that one or more LLM outputs may include one or more corresponding placeholders (e.g., as indicated by “[User]!” in the modified assistant output 654), which may be populated with user profile data accessible to the automation assistant. While Figure 6 Examples include placeholders for the names of the users of client device 610, but it should be understood that this is for illustrative purposes and does not imply limitation. For example, one or more placeholders may be filled with any data accessible to the automation assistant, such as smart networked device identifiers (e.g., smart lights, smart TVs, smart appliances, smart speakers, smart locks, etc.), known locations associated with the user of client device 610 (e.g., city, state, county, region, province, country, physical address of workplace, user's business, or the user's primary residence of client device 610, etc.), entity references (e.g., referencing people, places, things, etc.), software applications accessible at the user's client device 610, and / or any other data accessible to the automation assistant.
[0097] Furthermore, the modified assistant output 654 functions in response to assistant queries included in spoken utterance 652 (e.g., “It’s 8:30 AM”). However, the modified assistant output 654 not only functions to be personalized or customized for the user and to respond to assistant queries, but it also helps drive the conversation between the user and the automated assistant by further engaging the user in a dialogue (e.g., “Any funplans today?”). Without using the technique described herein regarding modifying the initially generated assistant output set based on processed spoken utterance 652 using one or more LLM outputs, the automated assistant could simply reply “It’s 8:30 AM” without offering any greeting to the user of client device 610 (e.g., “Good morning”), without addressing the user of client device 610 by name (e.g., “[User]”), and without further engaging the user of client device 610 in a conversational dialogue (e.g., “Any fun plans today?”). Therefore, the modified assistant output 654 resonates more effectively with the user of client device 610 compared to any assistant output included in the originally generated assistant output set that does not utilize one or more LLM outputs.
[0098] exist Figure 6 In the example, it is further assumed that a user of client device 610 provides the spoken utterance 656, "Yes, I'm thinking about going to the beach." In this example, the automation assistant can cause the processing of audio data captured from the spoken utterance 656 to generate an assistant output set generated without using one or more LLM outputs. Furthermore, the automation assistant can cause the processing of assistant queries included in the spoken utterance 656, the assistant output set, and / or the context of the dialogue session to generate an assistant output set modified using one or more LLM outputs (e.g., offline and / or online), and optionally generate additional assistant queries based on the assistant queries.
[0099] In this example, the assistant outputs included in the set of assistant outputs (generated without using one or more LLM outputs) can be limited because the assistant query included in utterance 656 does not request the automated assistant to perform any action. For example, the assistant outputs included in the set of assistant outputs may include “Sounds fun!”, “Surf’s up!”, “That sounds like fun!”, and / or other assistant outputs in response to utterance 656, but do not further enable the user of client device 610 to engage in a conversational session. In other words, the assistant outputs included in the set of assistant outputs may be limited in terms of lexical variation because the assistant outputs are not generated using one or more LLM outputs as described herein. Nevertheless, the automated assistant can utilize the assistant outputs included in the set of assistant outputs to determine how to modify one or more assistant outputs using one or more LLM outputs.
[0100] Furthermore, and as regarding Figure 3 and Figure 4 As described, the automated assistant can generate additional assistant queries based on the assistant query and using one or more LLMs or different ML models that communicate with one or more LLMs. For example, in Figure 6 In the example, spoken utterance 656 provided by a user of client device 610 instructs the user to go to the beach. Based on the identified intent associated with spoken utterance 656, indicating the user's intention to go to the beach, the automation assistant can determine a relevant intent associated with finding the weather at a beach frequently visited by the user of client device 610 (e.g., an example beach named "Half Moon Bay"). Based on the identified relevant intent, the automation assistant can generate an additional assistant query with the location parameter "Half Moon Bay" and ask "What's the weather?", submitting the query to one or more 1P systems 191 and / or 3P systems 192 to obtain additional assistant output, which includes "weather" for "Half Moon Bay," indicating, for example, that it will rain at "Half Moon Bay" and that the temperature will be cold and overcast throughout the day. In some implementations, the automation assistant may cause the processing of the additional assistant output and / or the context of the dialogue session to generate an additional modified set of assistant outputs determined using one or more LLM outputs and / or one or more additional LLM outputs.
[0101] exist Figure 6In the example, the automated assistant can cause the assistant outputs included in the assistant output set and the modified assistant output set to be ranked according to one or more ranking criteria, and can select one or more assistant outputs based on the ranking (e.g., selecting the given assistant output "Sounds fun!"). Furthermore, the automated assistant can cause the assistant outputs included in the additional modified assistant output set and the additional assistant outputs to be ranked according to one or more ranking criteria, and can select one or more assistant outputs based on the ranking (e.g., selecting the given additional assistant output "But if you're going to Half Moon Bay again, expect rain and chilly temps."). Furthermore, the automated assistant can combine a selected given assistant output with a selected given additional assistant output to produce a modified assistant output 658 such as “Sounds fun! But if you're going to Half Moon Bay again, expect rain and chilly temps,” and this modified assistant output 658 can be provided to the user of client device 610 for visual and / or audible presentation. Thus, in this example, the automated assistant can process spoken utterance 656 and provide additional contextual information associated with the spoken utterance (e.g., the weather at a beach the user of client device 610 might visit) to further enable the user of client device 610 to engage in a conversational session. Without the techniques described herein, although the automated assistant is capable of determining and providing weather information, the user of client device 610 might need to actively request weather information from the automated assistant, thereby increasing the amount of user input and wasting computational resources at client device 610 processing the increased amount of user input, and also increasing the cognitive load on the user of client device 610.
[0102] exist Figure 6In the example, it is further assumed that the user of client device 610 provides the spoken utterance 660, “Oh no…thanks for the heads up, can you remind me to check the weather again in two hours?” In this example, the automation assistant can cause processing to capture the audio data of spoken utterance 660 to generate an assistant output set generated without using one or more LLM outputs. Furthermore, the automation assistant can cause processing to include the assistant query, assistant output set, and / or the context of the conversation session included in spoken utterance 660 to generate an assistant output set modified using one or more LLM outputs (e.g., in offline and / or online manner). Based on processing spoken utterance 660, the automation assistant can determine to set a reminder at 10:30 AM (e.g., two hours after the start of the conversation session) to remind the user of client device 610 to check the weather at “Half Moon Bay,” or proactively provide the user of client device 610 with the weather at “Half Moon Bay” at 10:30 AM. Furthermore, the automated assistant can make the modified assistant output 662 “Sure thing, I set the reminder and hope the weather clears up for you” from (i.e., generated using one or more LLM outputs) a modified set of assistant outputs, presented visually and / or audibly to the user of client device 610. It is worth noting that... Figure 6 The modified assistant output 662 in the example is context-sensitive relative to the dialogue session because it instructs the user of the desire for the weather to clear up. Conversely, assistant outputs included in the assistant output set (i.e., those generated without using one or more LLM outputs) may simply provide instructions for setting a reminder, regardless of the context of the dialogue session.
[0103] Although it describes the use of one or more LLM outputs to generate specific modifications based on the specific utterances and context of the dialogue session, and the selection of the specific modifications to be presented to the user by the assistant output. Figure 6However, it should be understood that this is for illustrative purposes and not intended to be limiting. Rather, it should be understood that the techniques described herein can be used in any conversational session between any user and a corresponding instance of the automation assistant. Furthermore, while a transcription corresponding to a conversational session between a user and the automation assistant is depicted at the display 680 of the client device 610, it should also be understood that this is for illustrative purposes and not intended to be limiting. For example, it should be understood that the conversational session can be executed on any device capable of executing the automation assistant, regardless of whether the client device includes a display.
[0104] Furthermore, it should be understood that, in Figure 6 The assistant output provided to the user during a conversational session can include various personalized responses described herein. For example, a first set of parameters can be used to generate a modified assistant output 654, reflecting the automation assistant's primary personality in terms of a first set of words to be used by the automation assistant and / or a first set of prosodic attributes to be used to provide the modified assistant output 654 for audible presentation to the user. Furthermore, a second set of parameters can be used to generate a modified assistant output 658, reflecting the automation assistant's secondary personality in terms of a second set of words to be used by the automation assistant and / or a second set of prosodic attributes to be used to provide the modified assistant output 658 for audible presentation to the user. In this example, the primary personality could reflect, for example, the personality of a butler or maid providing a morning greeting in response to a user requesting the current time and asking if he / she has any plans for the day. Furthermore, the secondary personality could reflect, for example, the personality of a weather forecaster, surfer, or lifeguard (or a combination thereof), suggesting that going to the beach sounds interesting, but the weather might not be suitable for a day at the beach. For example, the Surfer personality could be used to provide the modified Assistant output 658's "Sounds fun!" part as the user instructs him / her to go to the beach, and the Weather Forecaster personality could be used to provide "But if you're going to Half Moon Bay again, expect rain and chilly temps" as the automated assistant is providing weather information to the user.
[0105] Therefore, the automated assistant can dynamically adjust the personality of the modified assistant output presented to the user based on both the vocabulary used by the automated assistant and the prosodic attributes used in rendering the modified assistant output for audible presentation to the user. Notably, the automated assistant can dynamically adjust these personalities when providing modified assistant output based on the context of the conversational session—including previously uttered utterances received from the user and previous assistant output provided by the automated assistant, and / or any other contextual signals described herein. As a result, the modified assistant output provided by the automated assistant can better resonate with the user of the client device.
[0106] Furthermore, although this article is about users providing spoken utterances throughout the conversation to describe... Figure 6 However, it should be understood that this is for illustrative purposes and does not imply limitation. For example, a user may additionally or alternatively provide typed input and / or touch input throughout the conversation. In these implementations, the automated assistant may process typed input (e.g., using NLU engines 140A1 and / or 140A2) to generate an NLU output stream (e.g., and optionally skip any processing using ASR engines 130A1 and / or 130A2), and may process NLU data streams and text input corresponding to assistant queries derived from typed input and / or touch input (e.g., using LLM engines 150A1 and / or 150A2) to generate one or more sets of modified assistant outputs in the same or similar manner as described above.
[0107] Turn now Figure 7 This diagram depicts a block diagram of an example computing device 710 that can be optionally used to perform one or more aspects of the techniques described herein. In some implementations, one or more of a client device, a cloud-based automation assistant component, and / or other components may include one or more components of the example computing device 710.
[0108] Computing device 710 typically includes at least one processor 714 that communicates with multiple peripheral devices via a bus subsystem 712. These peripheral devices may include a storage subsystem 724 (including, for example, a memory subsystem 725 and a file storage subsystem 726), a user interface output device 720, a user interface input device 722, and a network interface subsystem 716. The input and output devices allow users to interact with computing device 710. The network interface subsystem 716 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.
[0109] User interface input device 722 may include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen incorporated into a display, audio input devices such as a voice recognition system, a microphone, and / or other types of input devices. Generally, the term "input device" is used to encompass all possible types of devices and methods for inputting information into computing device 710 or a communication network.
[0110] User interface output device 720 may include a display subsystem, printer, fax machine, or non-visual display, such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide non-visual displays, such as via an audio output device. Generally, the term "output device" is used to encompass all possible types of devices and methods for outputting information from computing device 710 to a user or another machine or computing device.
[0111] Storage subsystem 724 stores the programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 724 may include selected aspects for performing the methods disclosed herein and for implementing... Figure 1 and Figure 2 The logic of the various components described in the text.
[0112] These software modules are typically executed by processor 714 alone or in combination with other processors. The memory 725 used in storage subsystem 724 may include multiple memories, including main random access memory (RAM) 730 for storing instructions and data during program execution and read-only memory (ROM) 732 for storing fixed instructions. File storage subsystem 726 can provide persistent storage for program and data files and may include hard disk drives, floppy disk drives, and associated removable media, CD-ROM drives, optical disk drives, or removable media cartridges. Modules implementing the functionality of certain embodiments may be stored by file storage subsystem 726 in storage subsystem 724 or in other machines accessible to processor 714.
[0113] The bus subsystem 712 provides a mechanism for enabling the various components and subsystems of the computing device 710 to communicate with each other as intended. Although the bus subsystem 712 is schematically shown as a single bus, alternative implementations of the bus subsystem 712 may use multiple buses.
[0114] The computing device 710 can be of different types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the constantly evolving nature of computers and networks, [the following applies]. Figure 7 The description of the computing device 710 depicted herein is intended only as a specific example for illustrating some embodiments. Many other configurations of the computing device 710 may have... Figure 7 The computing device depicted in the text has more or fewer components.
[0115] In cases where the systems described herein collect or otherwise monitor personal information about users, or may utilize personal and / or monitored information, users may be given the opportunity to control whether a program or feature collects user information (e.g., information about a user's social networks, social behaviors or activities, occupation, user preferences, or the user's current geographic location), or to control whether and / or how content that may be more relevant to the user is received from a content server. Furthermore, certain data may be processed in one or more ways before storage or use to remove personally identifiable information. For example, a user's identity may be processed to the point that the user's personally identifiable information cannot be determined, or a user's geographic location may be generalized (e.g., down to the city, zip code, or state level) where geographic location information is available, making it impossible to determine the user's specific geographic location. Therefore, users can control how information about themselves is collected and / or used.
[0116] In some implementations, a method implemented by one or more processors is provided, the method comprising: as part of a dialogue session between a user of a client device and an automated assistant implemented by the client device; receiving an audio data stream capturing spoken utterances of the user, the audio data stream being generated by one or more microphones of the client device, and the spoken utterances including an assistant query; determining an assistant output set based on processing the audio data stream, each assistant output in the assistant output set responding to an assistant query included in the spoken utterances; processing the assistant output set and the context of the dialogue session to: generate a modified assistant output set using one or more LLM outputs generated leveraging a large language model (LLM), each of the one or more LLM outputs being based on at least a portion of the context of the dialogue session and including... One or more assistant outputs in the assistant output set are determined, and additional assistant queries related to spoken utterances are generated based on at least a portion of the context of the dialogue session and at least a portion of the assistant query; additional assistant outputs in response to the additional assistant queries are determined based on the additional assistant queries; the additional assistant outputs and the context of the dialogue session are processed to generate an additional modified set of assistant outputs using LLM outputs generated using LLM or one or more of the additional LLM outputs, each of the additional LLM outputs being determined based on at least a portion of the context of the dialogue session and the additional assistant outputs; and a given modified assistant output from the modified set of assistant outputs and a given additional modified assistant output from the additional modified set of assistant outputs are provided to the user.
[0117] These and other embodiments of the technology disclosed herein may optionally include one or more of the following features.
[0118] In some implementations, determining the assistant output in response to an assistant query included in spoken utterance based on processing an audio data stream may include: processing the audio data stream using an automatic speech recognition (ASR) model to generate an ASR output stream; processing the ASR output stream using a natural language understanding (NLU) model to generate an NLU data stream; and causing the set of assistant outputs to be determined based on the NLU stream.
[0119] In some versions of those implementations, processing the assistant output set and the context of the dialogue session to generate a modified assistant output set using one or more of the LLM outputs generated using LLM may include: processing the assistant output set and the context of the dialogue session using LLM to generate one or more of the LLM outputs; and determining the modified assistant output set based on one or more of the LLM outputs. In some further versions of those implementations, processing the assistant output set and the context of the dialogue session using LLM to generate one or more of the LLM outputs may include: processing the assistant output set and the context of the dialogue session using a first LLM parameter set from a plurality of disparate LLM parameter sets to determine one or more of the LLM outputs having a first property among a plurality of disparate properties. The modified assistant output set may include one or more first-property assistant outputs reflecting the first property. In yet another version of those implementations, processing the assistant output set and the context of the dialogue session using LLM to generate one or more of the LLM outputs may include: processing the assistant output set and the context of the dialogue session using a second LLM parameter set from a plurality of disparate LLM parameter sets to determine one or more of the LLM outputs having a second property among a plurality of disparate properties. The modified assistant output set may include one or more secondary-nature assistant outputs reflecting a second nature, and the second nature may be unique from the first nature. In a further version of those embodiments, a first word associated with the first nature may be used to determine one or more primary-nature assistant outputs included in the modified assistant output set and reflecting the first nature, and a second word associated with the second nature may be used to determine one or more secondary-nature assistant outputs included in the modified assistant output set and reflecting the second nature, wherein the second nature differs from the first nature based on the second word being different from the first word. In a further additional or alternative version of those embodiments, one or more primary-nature assistant outputs included in the modified assistant output set and reflecting the first nature may be associated with a first prosodic attribute set used when providing a given modified assistant output for audible presentation to a user, and one or more secondary-nature assistant outputs included in the modified assistant output set and reflecting the second nature may be associated with a second prosodic attribute set used when providing a given modified assistant output for audible presentation to a user, and the second nature may differ from the first nature based on the second prosodic attribute set being different from the first prosodic attribute set.
[0120] In some versions of those implementations, processing the assistant output set and the context of the dialogue session to generate a modified assistant output set using one or more of the LLM outputs generated using the LLM model may include: identifying one or more of the LLM outputs previously generated using the LLM model, based on the fact that one or more of the LLM outputs were previously generated based on a previous assistant query of a previous dialogue session corresponding to the assistant query of the dialogue session, and / or based on the fact that one or more of the LLM outputs were previously generated based on a previous context of a previous dialogue session corresponding to the context of the dialogue session; and causing the assistant output set to be modified using one or more of the LLM outputs to determine the modified assistant output set. In some further versions of those implementations, identifying one or more of the LLM outputs previously generated using the LLM model may include: identifying one or more first LLM outputs among the one or more LLM outputs that reflect the firstness of a plurality of dissimilar personalities. The modified assistant output set includes one or more first-nature assistant outputs reflecting the firstness. In a further version of those implementations, identifying one or more of the LLM outputs previously generated using the LLM model may include: identifying one or more second LLM outputs among the one or more LLM outputs that reflect the secondness of a plurality of dissimilar personalities. The modified assistant output set may include one or more second-nature assistant outputs reflecting a second nature, and the second nature may differ from the first nature. In further additional or alternative versions of those embodiments, the method may further include: determining that a previous assistant query of a previous conversation corresponds to an assistant query of the current conversation based on the ASR output including one or more terms of an assistant query corresponding to one or more terms of a previous assistant query of a previous conversation. In further additional or alternative versions of those embodiments, the method may further include: generating an embedding of the assistant query based on one or more terms in the ASR output corresponding to the assistant query; and determining that a previous assistant query of a previous conversation corresponds to an assistant query of the current conversation based on comparing the embedding of the assistant query with a previously generated embedding of a previous assistant query of a previous conversation. In further additional or alternative versions of those embodiments, the method may further include: determining that a previous context of a previous conversation corresponds to a context of the current conversation based on one or more context signals of the current conversation corresponding to one or more context signals of the previous conversation. In further still versions of those embodiments, the one or more context signals may include one or more of the following: time of day, day of week, location of the client device, and ambient noise in the environment of the client device.In further additional or alternative versions of these implementations, the method may also include: generating an embedding of the context of the dialogue session based on context signals of the dialogue session; and determining that the previous context of the previous dialogue session corresponds to the context of the dialogue session by comparing the embeddings of one or more context signals with a previously generated embedding of a previous context of a previous dialogue session.
[0121] In some versions of those implementations, processing the assistant output set and the context of the dialogue session to generate additional assistant queries related to the spoken utterance based on at least a portion of the context of the dialogue session and at least a portion of the assistant query may include: determining an intent associated with the assistant query included in the spoken utterance based on the NLU output; identifying at least one related intent related to the intent associated with the assistant query included in the spoken utterance based on the intent associated with the assistant query included in the spoken utterance; and generating additional assistant queries related to the spoken utterance based on the at least one related intent. In some further versions of those implementations, determining additional assistant output in response to the additional assistant query based on the additional assistant query may include: causing the additional assistant query to be transmitted via an application programming interface (API) to one or more first-party systems to generate additional assistant output in response to the additional assistant query. In some additional or alternative further versions of those implementations, determining additional assistant output in response to the additional assistant query based on the additional assistant query may include: causing the additional assistant query to be transmitted via one or more networks to one or more third-party systems; and receiving additional assistant output in response to the additional assistant query being transmitted to one or more third-party systems. In some additional or alternative versions of those implementations, processing additional assistant outputs and the context of the dialogue session using one or more of the LLM outputs determined by LLM or one or more of the additional LLM outputs to generate an additional modified set of assistant outputs may include: processing the additional assistant output set and the context of the dialogue session using LLM to determine one or more of the additional LLM outputs; and determining the additional modified set of assistant outputs based on one or more of the additional LLM outputs. In some additional or alternative versions of those implementations, processing additional assistant outputs and the context of the dialogue session using one or more of the LLM outputs determined by LLM or one or more of the additional LLM outputs to generate an additional modified set of assistant outputs may include: identifying one or more of the additional LLM outputs previously generated using the LLM model based on the fact that one or more of the additional LLM outputs were previously generated based on a previous assistant query of a previous dialogue session corresponding to the additional assistant query of the dialogue session, and / or based on the fact that one or more of the additional LLM outputs were previously generated for a previous context of a previous dialogue session corresponding to the context of the dialogue session; and causing the additional assistant output set to be modified using one or more of the additional LLM outputs to generate an additional modified set of assistant outputs.
[0122] In some implementations, the method may further include: ranking a superset of assistant outputs based on one or more ranking criteria, the superset of assistant outputs including at least a set of assistant outputs and a modified set of assistant outputs; and selecting a given modified assistant output from the modified set of assistant outputs based on the ranking. In some versions of those implementations, the method may further include: ranking an additional superset of assistant outputs based on one or more ranking criteria, the superset of assistant outputs including at least an additional set of assistant outputs and an additional set of modified assistant outputs; and selecting a given additional modified assistant output from the additional set of modified assistant outputs based on the ranking. In some other versions of those implementations, providing a given modified assistant output and a given additional modified assistant output to be presented to the user may include: combining the given modified assistant output and the given additional modified assistant output; processing the given modified assistant output and the given additional modified assistant output using a text-to-speech (TTS) model to generate synthesized speech audio data, the synthesized speech audio data including synthesized speech captured from the given modified assistant output and the given additional modified assistant output; and making the synthesized speech audio data audibly rendered to be presented to the user via a speaker of a client device.
[0123] In some implementations, the method may further include: ranking a superset of assistant outputs based on one or more ranking criteria, the superset of assistant outputs including a set of assistant outputs, a modified set of assistant outputs, an additional set of assistant outputs, and an additional set of modified assistant outputs; and selecting a given modified assistant output from the modified set of assistant outputs and a given additional modified assistant output from the additional set of modified assistant outputs based on the ranking. In some other versions of those implementations, making the given modified assistant output and the given additional modified assistant output available to the user may include: processing the given modified assistant output and the given additional modified assistant output using a text-to-speech (TTS) model to generate synthesized speech audio data, the synthesized speech audio data including synthesized speech captured from the given modified assistant output and the given additional modified assistant output; and making the synthesized speech audio data audibly rendered to be presented to the user via a speaker of a client device.
[0124] In some implementations, using one or more of the LLM outputs to generate a modified set of assistant outputs may also be based on processing at least a portion of the assistant queries included in spoken utterances.
[0125] In some implementations, a method implemented by one or more processors is provided, the method comprising: as part of a dialogue session between a user of a client device and an automated assistant implemented by the client device; receiving an audio data stream capturing spoken utterances of the user, the audio data stream being generated by one or more microphones of the client device, and the spoken utterances including an assistant query; determining an assistant output set based on processing the audio data stream, each assistant output in the assistant output set responding to an assistant query included in the spoken utterances; processing the assistant output set and the context of the dialogue session to: generate a modified assistant output set using one or more LLM outputs generated using a large language model (LLM), each of the one or more LLM outputs being determined based on at least a portion of the context of the dialogue session and one or more assistant outputs included in the assistant output set, and generating additional assistant queries related to the spoken utterances based on at least a portion of the context of the dialogue session and at least a portion of the assistant queries; determining additional assistant outputs responding to the additional assistant queries based on the additional assistant queries; processing the modified assistant output set based on the additional assistant outputs responding to the additional assistant queries to generate an additional modified assistant output set; and causing a given additional modified assistant output from the additional modified assistant output set to be provided to the user.
[0126] In some implementations, a method implemented by one or more processors is provided, the method comprising: as part of a dialogue session between a user of a client device and an automated assistant implemented by the client device; receiving an audio data stream capturing spoken utterances of the user, the audio data stream being generated by one or more microphones of the client device, and the spoken utterances including an assistant query; determining an assistant output set based on processing the audio data stream, each assistant output in the assistant output set responding to an assistant query included in the spoken utterances; processing the assistant output set and the context of the dialogue session using one or more LLM outputs generated using a large language model (LLM) to generate a modified assistant output set, each of the one or more LLM outputs being determined based on at least a portion of the context of the dialogue session and one or more assistant outputs included in the assistant output set. Generating a modified assistant output set using one or more LLM outputs includes: generating a first-person response set based on (i) the assistant output set, (ii) the context of the dialogue session, and (iii) one or more first LLM outputs reflecting a first-personality among a plurality of disparate personalities. The method further includes causing a given modified assistant output from the modified assistant output set to be provided to the user.
[0127] In some implementations, a method implemented by one or more processors is provided, the method comprising: as part of a dialogue session between a user of a client device and an automated assistant implemented by the client device; receiving an audio data stream capturing spoken utterances of the user, the audio data stream being generated by one or more microphones of the client device, and the spoken utterances including an assistant query; determining an assistant output set based on processing the audio data stream, each assistant output in the assistant output set responding to an assistant query included in the spoken utterances; processing the assistant output set and the context of the dialogue session using one or more LLM outputs generated using a large language model (LLM) to generate a modified assistant output set, each of the one or more LLM outputs being determined based on at least a portion of the context of the dialogue session and one or more assistant outputs included in the assistant output set; and causing a given modified assistant output from the modified assistant output set to be provided to the user.
[0128] In some implementations, a method implemented by one or more processors is provided, the method comprising: as part of a dialogue session between a user of a client device and an automated assistant implemented by the client device; receiving an audio data stream capturing spoken utterances of the user, the audio data stream being generated by one or more microphones of the client device, and the spoken utterances including an assistant query; determining an assistant output set based on processing the audio data stream, each assistant output in the assistant output set responding to an assistant query included in the spoken utterances; determining, based on processing the spoken utterances, whether to modify one or more of the assistant outputs included in the assistant output set; in response to determining to modify one or more of the assistant outputs included in the assistant output set: processing the assistant output set and the context of the dialogue session using one or more LLM outputs generated using a large language model (LLM) to generate a modified assistant output set, each of the one or more LLM outputs being determined based on at least a portion of the context of the dialogue session and one or more of the assistant outputs included in the assistant output set; and causing a given modified assistant output from the modified assistant output set to be provided to the user.
[0129] These and other embodiments of the technology disclosed herein may optionally include one or more of the following features.
[0130] In some implementations, determining whether to modify one or more of the assistant outputs included in the assistant output set based on the processing of spoken utterance may include: processing an audio data stream using an automatic speech recognition (ASR) model to generate an ASR output stream; processing the ASR output stream using a natural language understanding (NLU) model to generate an NLU data stream; identifying the user's intent in providing spoken utterance based on the NLU data stream; and determining whether to modify the assistant output based on the user's intent in providing spoken utterance.
[0131] In some implementations, determining whether to modify one or more of the assistant outputs included in the assistant output set may also be based on one or more computational costs associated with modifying one or more of the assistant outputs. In some versions of those implementations, the one or more computational costs associated with modifying one or more of the assistant outputs may include one or more of the following: battery consumption, processor consumption, or latency associated with modifying one or more of the assistant outputs.
[0132] In some implementations, a method implemented by one or more processors is provided, the method comprising: obtaining a plurality of assistant queries directed to an automation assistant and a corresponding context of a previous dialogue session for each of the plurality of assistant queries; for each of the plurality of assistant queries: processing a given assistant query among the plurality of assistant queries using one or more large language models (LLMs) to generate a corresponding LLM output in response to the given assistant query; and indexing the corresponding LLM output in memory accessible at a client device based on the given assistant query and / or the corresponding context of a previous dialogue session for the given assistant query; and in memory accessible at the client device. After indexing the corresponding LLM output in the memory, and as part of the current dialogue session between the user of the client device and the automated assistant implemented by the client device: receiving an audio data stream capturing the user's spoken words, the audio data stream being generated by one or more microphones of the client device; based on processing the audio data stream, determining that the spoken words include the current assistant query corresponding to a given assistant query and / or the spoken words being received in the current context of the current dialogue session corresponding to the previous dialogue session corresponding to the given assistant query; and causing the automated assistant to generate assistant output to be presented to the user in response to the spoken words using the corresponding LLM output.
[0133] These and other embodiments of the technology disclosed herein may optionally include one or more of the following features.
[0134] In some implementations, multiple assistant queries targeting the automated assistant may have been previously submitted by a user via a client device. In other implementations, multiple assistant queries targeting the automated assistant may have been previously submitted by multiple additional users, in addition to the user on the client device, via their respective client devices.
[0135] In some implementations, indexing the corresponding LLM output in memory accessible at the client device may be based on the embedding of a given assistant query generated when processing a given assistant query. In some implementations, indexing the corresponding LLM output in memory accessible at the client device may be based on one or more terms or phrases included in a given assistant query generated when processing a given assistant query. In some implementations, indexing the corresponding LLM output in memory accessible at the client device may be based on the embedding of the corresponding context of a previous conversation session for a given assistant query. In some implementations, indexing the corresponding LLM output in memory accessible at the client device may be based on one or more context signals included in the corresponding context of a previous conversation session for a given assistant query.
[0136] Additionally, some embodiments include one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in associated memory, and wherein the instructions are configured to cause any of the aforementioned methods to be performed. Some embodiments also include one or more non-transitory computer-readable storage media storing computer instructions executable by one or more processors to perform any of the aforementioned methods. Some embodiments also include a computer program product comprising instructions executable by one or more processors to perform any of the aforementioned methods.
Claims
1. A method implemented by one or more processors, the method comprising: As part of a conversation between the user of the client device and the automated assistant implemented by the client device: Receive an audio data stream that captures the user's spoken words, the audio data stream being generated by one or more microphones of the client device, and the spoken words including an assistant query; An assistant output set is determined based on processing the audio data stream, each assistant output in the assistant output set being in response to an assistant query included in the spoken utterance; Process the assistant output set and the context of the dialogue session to: A modified set of assistant outputs is generated using one or more LLM outputs generated using a large language model LLM, each of the one or more LLM outputs being determined based on at least a portion of the context of the dialogue session and one or more assistant outputs included in the set of assistant outputs. Additional assistant queries related to the spoken utterance are generated based on at least a portion of the context of the dialogue session and at least a portion of the assistant query. The additional assistant output in response to the additional assistant query is determined based on the additional assistant query; The additional assistant outputs and the context of the dialogue session are processed to generate an additional modified set of assistant outputs using the LLM outputs generated using the LLM or one or more of the additional LLM outputs, each of the additional LLM outputs being determined based on at least a portion of the context of the dialogue session and the additional assistant outputs; as well as This allows a given modified assistant output from the modified assistant output set and a given additional modified assistant output from the additional modified assistant output set to be provided to the user.
2. The method according to claim 1, wherein, The assistant output determined based on processing the audio data stream in response to the assistant query included in the spoken utterance includes: The audio data stream is processed using an Automatic Speech Recognition (ASR) model to generate an ASR output stream; The ASR output stream is processed using a natural language understanding NLU model to generate an NLU data stream; and This enables the determination of the assistant output set based on the NLU stream.
3. The method according to claim 2, wherein, Processing the assistant output set and the context of the dialogue session to generate the modified assistant output set using one or more LLM outputs generated using the LLM includes: The LLM is used to process the assistant output set and the context of the dialogue session to generate one or more LLM outputs from the LLM outputs; and The modified set of assistant outputs is determined based on one or more of the LLM outputs.
4. The method according to claim 3, wherein, Using the LLM to process the assistant output set and the context of the dialogue session to generate one or more LLM outputs from the LLM outputs includes: The assistant output set and the context of the dialogue session are processed using a first LLM parameter set from multiple disparate LLM parameter sets to determine one or more LLM outputs with a first characteristic among multiple disparate characteristics. The modified assistant output set includes one or more primary personality assistant outputs that reflect the first personality.
5. The method according to claim 4, wherein, Using the LLM to process the assistant output set and the context of the dialogue session to generate one or more LLM outputs from the LLM outputs includes: The assistant output set and the context of the dialogue session are processed using a second set of multiple disparate LLM parameter sets to determine one or more LLM outputs having a second personality among multiple disparate personalities. The modified assistant output set includes one or more second-personality assistant outputs reflecting the second personality, and The second personality is different from the first personality.
6. The method according to claim 5, wherein, The one or more first-personality assistant outputs included in the modified assistant output set and reflecting the first personality are determined using a first word associated with the first personality, wherein the one or more second-personality assistant outputs included in the modified assistant output set and reflecting the second personality are determined using a second word associated with the second personality, and wherein the second personality is different from the first personality based on the second word being different from the first word.
7. The method according to claim 5, wherein, The one or more first-personality assistant outputs included in the modified assistant output set and reflecting the first personality are associated with a first prosodic attribute set used in providing the given modified assistant output for audible presentation to the user, wherein the one or more second-personality assistant outputs included in the modified assistant output set and reflecting the second personality are associated with a second prosodic attribute set used in providing the given modified assistant output for audible presentation to the user, and wherein the second personality is different from the first personality based on the second prosodic attribute set being different from the first prosodic attribute set.
8. The method according to claim 2, wherein, Processing the assistant output set and the context of the dialogue session to generate the modified assistant output set using one or more LLM outputs generated using the LLM includes: Based on the fact that one or more of the LLM outputs were generated prior to a previous assistant query of a previous conversation corresponding to the assistant query of the conversation session, and / or based on the fact that one or more of the LLM outputs were generated prior to a previous context of the previous conversation corresponding to the context of the conversation session, identify one or more LLM outputs generated prior to the LLM outputs; and This allows the assistant output set to be modified using one or more of the LLM outputs to determine the modified assistant output set.
9. The method according to claim 8, wherein, Identifying one or more LLM outputs generated prior to the use of the LLM includes: Identify one or more first LLM outputs that reflect the first of multiple dissimilar personalities among the one or more LLM outputs. The modified assistant output set includes one or more primary personality assistant outputs that reflect the first personality.
10. The method according to claim 9, wherein, Identifying one or more LLM outputs generated prior to the use of the LLM includes: Identify one or more second LLM outputs that reflect a second nature among multiple dissimilar personalities in the one or more LLM outputs. The modified assistant output set includes one or more second-personality assistant outputs reflecting the second personality, and The second personality is different from the first personality.
11. The method of claim 8, further comprising: Based on the ASR output including one or more terms of an assistant query that correspond to one or more terms of the previous assistant query in the previous conversation, it is determined that the previous assistant query in the previous conversation corresponds to the assistant query in the conversation.
12. The method according to claim 8, further comprising: Based on one or more terms in the ASR output that correspond to the assistant query, generate the embedding of the assistant query; as well as Based on comparing the embedding of the assistant query with the previously generated embedding of the previous assistant query in the previous conversation, it is determined that the previous assistant query in the previous conversation corresponds to the assistant query in the current conversation.
13. The method of claim 8, further comprising: Based on one or more context signals of the current dialogue session corresponding to one or more context signals of the previous dialogue session, it is determined that the previous context of the previous dialogue session corresponds to the context of the current dialogue session.
14. The method according to claim 13, wherein, The one or more context signals include one or more of the following: time of day, day of week, location of the client device, and ambient noise in the environment of the client device.
15. The method of claim 8, further comprising: The context of the dialogue session is embedded based on the context signals of the dialogue session. as well as Based on comparing the embedding of the one or more context signals with the previously generated embedding of the previous context of the previous dialogue session, it is determined that the previous context of the previous dialogue session corresponds to the context of the dialogue session.
16. The method according to claim 2, wherein, Processing the assistant output set and the context of the dialogue session to generate the additional assistant query related to the spoken utterance based on at least a portion of the context of the dialogue session and at least a portion of the assistant query includes: The intent associated with the assistant query included in the spoken utterance is determined based on the NLU output; Based on the intent associated with the assistant query included in the spoken utterance, identify at least one related intent associated with the intent, the intent being associated with the assistant query included in the spoken utterance; and The additional assistant query related to the spoken utterance is generated based on the at least one relevant intent.
17. The method according to claim 16, wherein, Determining the additional assistant output in response to the additional assistant query based on the additional assistant query includes: This causes the additional assistant query to be transmitted to one or more first-party systems via an application programming interface (API) to generate the additional assistant output in response to the additional assistant query.
18. The method according to claim 16, wherein, Determining the additional assistant output in response to the additional assistant query based on the additional assistant query includes: This causes the additional assistant query to be transmitted to one or more third-party systems via one or more networks; and In response to the additional assistant query being transmitted to one or more of the third-party systems, receive the additional assistant output in response to the additional assistant query.
19. The method of claim 16, wherein, Processing the additional assistant outputs and the context of the dialogue session using one or more LLM outputs determined by the LLM or one or more additional LLM outputs from the additional LLM outputs to generate the additional modified set of assistant outputs includes: The LLM is used to process the set of additional assistant outputs and the context of the dialogue session to determine one or more additional LLM outputs from the additional LLM outputs; and The additional set of auxiliary outputs for modification is determined based on one or more of the additional LLM outputs.
20. The method of claim 16, wherein, Processing the additional assistant outputs and the context of the dialogue session using one or more of the LLM outputs determined by the LLM, or one or more of the additional LLM outputs, to generate the additional modified set of assistant outputs includes: Based on the fact that one or more additional LLM outputs generated prior to the LLM are generated based on a previous assistant query of a previous conversation corresponding to the additional assistant query of the conversation session, and / or based on the fact that one or more additional LLM outputs are generated prior to the previous context of the previous conversation corresponding to the context of the conversation session, identify one or more additional LLM outputs; and This allows the additional assistant output set to be modified using one or more of the additional LLM outputs to generate the additional modified assistant output set.
21. The method according to claim 1, further comprising: The superset of assistant outputs is ranked based on one or more ranking criteria, wherein the superset of assistant outputs includes at least the assistant output set and the modified assistant output set; as well as Based on the ranking, the given modified assistant output is selected from the set of modified assistant outputs.
22. The method of claim 21, further comprising: The superset of the additional assistant outputs is ranked based on one or more of the ranking criteria, wherein the superset of the assistant outputs includes at least the additional assistant outputs and the additional modified set of assistant outputs. as well as Based on the ranking, the given additional modified assistant output is selected from the set of additional modified assistant outputs.
23. The method according to claim 22, wherein, The provision of the given modified assistant output and the given additional modified assistant output to the user includes: Combine the given modified assistant output with the given additional modified assistant output; The given modified assistant output and the given additional modified assistant output are processed using a text-to-speech (TTS) model to generate synthesized speech audio data, the synthesized speech audio data comprising synthesized speech capturing the given modified assistant output and the given additional modified assistant output; and This allows the synthesized speech audio data to be audibly rendered and presented to the user via the speaker of the client device.
24. The method according to claim 1, further comprising: The superset of assistant outputs is ranked based on one or more ranking criteria, wherein the superset of assistant outputs includes the assistant output set, the modified assistant output set, the additional assistant outputs, and the additional modified assistant output set. as well as Based on the ranking, the given modified assistant output is selected from the modified assistant output set, and the given additional modified assistant output is selected from the additional modified assistant output set.
25. The method according to claim 24, wherein, The provision of the given modified assistant output and the given additional modified assistant output to the user includes: The given modified assistant output and the given additional modified assistant output are processed using a text-to-speech (TTS) model to generate synthesized speech audio data, the synthesized speech audio data comprising synthesized speech capturing the given modified assistant output and the given additional modified assistant output; and This allows the synthesized speech audio data to be audibly rendered and presented to the user via the speaker of the client device.
26. The method according to any one of claims 1 to 25, wherein, The modified assistant output set is generated using one or more of the LLM outputs, also based on processing at least a portion of the assistant query included in the spoken utterance.
27. A method implemented by one or more processors, the method comprising: As part of a conversation between the user of the client device and the automated assistant implemented by the client device: Receive an audio data stream that captures the user's spoken words, the audio data stream being generated by one or more microphones of the client device, and the spoken words including an assistant query; An assistant output set is determined based on processing the audio data stream, each assistant output in the assistant output set being in response to an assistant query included in the spoken utterance; Process the assistant output set and the context of the dialogue session to: A modified set of assistant outputs is generated using one or more LLM outputs generated using a large language model LLM, each of the one or more LLM outputs being determined based on at least a portion of the context of the dialogue session and one or more assistant outputs included in the set of assistant outputs. Additional assistant queries related to the spoken utterance are generated based on at least a portion of the context of the dialogue session and at least a portion of the assistant query. The additional assistant output in response to the additional assistant query is determined based on the additional assistant query; Based on the additional assistant output in response to the additional assistant query, the modified assistant output set is processed to generate an additional modified assistant output set; as well as This causes a given additional modified assistant output from the set of additional modified assistant outputs to be provided to the user.
28. A system comprising: At least one processor; as well as A memory for storing instructions, which, when executed, cause the at least one processor to perform an operation corresponding to any one of claims 1 to 27.
29. A non-transitory computer-readable storage medium storing instructions, which, when executed, cause at least one processor to perform an operation corresponding to any one of claims 1 to 27.
Citation Information
Patent Citations
Adaptive interface in a voice-based networked system
US20190318724A1
Speaker awareness using speaker dependent speech model(s)
WO2021112840A1