Dynamic adaptation of a given assistant output based on a given persona assigned to the automated assistant
By dynamically adapting to the output of the automatic assistant, and generating and controlling the visual queue according to the roles assigned by the user, the problem of lack of richness of the visual content of the existing automatic assistant is solved, and more rich and context-related visual interactions are achieved.
Patent Information
- Application Number
- JP2024561899
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-21
- Filing Date
- 2022-10-18
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-10-18
AI Technical Summary
When interacting with users, the visual content provided by existing automatic assistants lacks richness and cannot effectively convey nonverbal elements in human communication, such as facial expressions, body language and animation.
By dynamically adapting to the output of the automatic assistant, visual queues, including text content and visual animation, are generated and controlled according to the roles assigned by the user, to enhance the visual richness of the interaction.
It realizes that when interacting with users, automatic assistants provide richer and more context-related visual content, improve user experience and enhance the vividness of interaction.
Smart Images

Figure 2025514775000001_ABST
Abstract
Description
[Background technology]
[0001] Humans can engage in human-computer interaction using interactive software applications referred to herein as “automated assistants” (also referred to as “chatbots,” “interactive personal assistants,” “intelligent personal assistants,” “personal voice assistants,” “conversational agents,” etc.). Automated assistants typically rely on a pipeline of components in interpreting and responding to speech. For example, an automatic speech recognition (ASR) engine may process audio data corresponding to a user's utterance to generate an ASR output, such as an ASR hypothesis for the utterance (i.e., a sequence of term(s) and / or other token(s)). Additionally, a natural language understanding (NLU) engine may process the ASR output (or touch / typed input) to generate an NLU output, such as a request (e.g., intent) expressed by the user in giving the utterance (or touch / typed input) and, optionally, slot value(s) for parameter(s) associated with the intent. Finally, the NLU output may be processed by various fulfillment components to generate fulfillment outputs, such as responsive content responsive to the utterance and / or one or more actions that may be performed in response to the utterance.
[0002] Generally, these automated assistants respond to speech using a pipeline of the aforementioned components. For example, these automated assistants may cause auditory content to be provided using various text-to-speech (TTS) technologies for audible presentation to the user, such as responses to queries, confirmation that one or more actions have been performed on the user's behalf, etc. Furthermore, these automated assistants may additionally or alternatively cause visual content to be provided for visual presentation to the user, such as displaying information cards containing information requested by the user, various interfaces for various applications accessible at the client device on which these automated assistants are implemented, etc. While improvements have been made with respect to the auditory content provided by these automated assistants for presentation to the user during these interaction sessions (e.g., to more accurately reflect natural conversations between humans), improvements in the visual content provided by these automated assistants for visual presentation during these interaction sessions have remained relatively unchanged.
[0003] For example, assume that a given user directs the utterance "Good morning" to a given automated assistant implemented at a given client device of the given user. In this example, further assume that this utterance causes the given automated assistant to greet the given user, provide a weather forecast summary to the given user, and provide news headlines to the given user, such as through an auditory presentation of a synthesized voice of "Good morning John, the weather today is 65 and sunny, here is your news." In particular, the given automated assistant may cause the display of the given client device to display text content corresponding to the system voice, information cards related to the weather forecast, information cards related to new headlines, and the like. However, in this example, the visual content provided for presentation by the given automated assistant to the given user merely confirms the information provided in the auditory content. Thus, the visual content provided for presentation to a given user lacks the visual richness that humans may express when communicating, such as facial expressions, body language, animation, and the like.
[0004] One solution to this lack of visual richness of human representations may include providing visualized representations (e.g., avatars) for these automated assistants. However, current visualized representations of these automated assistants are relatively limited in terms of the visual richness of the information they can convey to users. Continuing with the above example, a given visualized representation of a given automated assistant may persist on the display of a given client device, or may employ a hard-coded rule of "context[greeting]=action[wave]" such that a given visualized representation of a given automated assistant waves to a given user when they say "good morning". However, the range of visual richness that humans may express when communicating may be virtually limitless. As a result, there is a need in the art for techniques aimed at improving the visual content provided by these automated assistants for presentation to users during these interaction sessions. Summary of the Invention
[0005] The implementations described herein aim to enable an automated assistant to dynamically adapt a given assistant output based on a given persona assigned to the automated assistant from among a plurality of heterogeneous personas. The given assistant output may include, for example, a stream of corresponding textual content and a stream of corresponding visual cues. The stream of corresponding textual content may include textual content responsive to a corresponding utterance and synthesized for auditory presentation to the user who provided the corresponding utterance. Furthermore, the stream of corresponding visual cues may include instructions for controlling a display of a client device (e.g., on which an instance of the automated assistant is implemented) in response to the corresponding utterance, and / or instructions for controlling a visual representation of the instance of the automated assistant. For example, the stream of corresponding visual cues may include display animations that cause the display of the client device to be dynamically adapted, animated body movement gestures performed by the visual representation of the automated assistant, and / or any other instructions that otherwise control the display and / or the visual representation of the instance of the automated assistant. A given persona may be embodied, for example, by a given vocabulary that is specific to the given persona and utilized in generating a corresponding stream of textual content, a given set of prosodic features that are specific to the given persona and utilized in synthesizing the corresponding stream of textual content for aural presentation to a user, and / or a given set of visual cues that include some visual cues that are specific to the given persona (e.g., animated body movement gestures commonly associated with visualized representations of instances of automated assistants) and some visual cues that are common among multiple ones of the multiple heterogeneous personas (e.g., waving, certain facial expressions, etc.).The visual representation of the automated assistant may be, for example, an animated avatar or entity that represents an instance of the automated assistant and may be based, for example, on a real human, a fictional character, animated object(s), and / or animal(s), and / or other visual representation.
[0006] An embodiment may receive a stream of audio data capturing a user's speech and generate a given assistant output responsive to the speech based on processing the stream of audio data. The given assistant output may include a stream of text content and a stream of visual cues for controlling a display of a client device in response to the speech and / or for controlling a visual representation of an instance of an automated assistant that is visually rendered for presentation to a user via a display of the client device. In accordance with an embodiment, synthetic speech audio data capturing synthetic speech corresponding to the stream of text content may be audibly rendered for presentation to a user via a speaker(s) of the client device, and the stream of visual cues may be utilized to control a display of the client device and / or control a visual representation of an instance of an automated assistant. In particular, the given assistant output responsive to the speech may be specific to a given persona that a user has assigned to an instance of a client device.
[0007] In some versions of those embodiments, the embodiments may process a stream of audio data capturing the speech using an automatic speech recognition (ASR) model to generate a stream of ASR output. Additionally, the embodiments may process the stream of ASR output using a natural language understanding (NLU) model to generate a stream of NLU output. Additionally, the embodiments may determine a given assistant output responsive to the speech based on at least the stream of NLU output. Notably, the given assistant output in these embodiments is generated using a typical automated assistant pipeline and may not include a stream of visual cues or may only include a very rudimentary stream of visual cues, such as static graphics or information cards provided for visual presentation to the user. Thus, the embodiments may modify the given assistant output to generate a modified given assistant output that adapts the given assistant output to a given persona assigned to the automated assistant. For example, an embodiment may process a given Assistant output using a large-scale language model (LLM) (e.g., one or more Transformer models, such as Meena, RNN, and / or any other LLM) that is specific to a given persona and / or utilizes given persona data specific to a given persona to generate a modified given Assistant output. Also, for example, an embodiment may determine previously generated LLM outputs that were previously generated based on the same or similar utterances provided by a user to generate a modified given Assistant output.
[0008] As used herein, previously generated LLM outputs utilized in an offline manner (e.g., generated prior to receiving an utterance) and / or LLM outputs generated in an online manner (e.g., generated in response to receiving an utterance) may include one or more probability distributions. For example, in determining the stream of textual content described herein, these LLM outputs may include corresponding probability distributions of sequences of one or more words and / or phrases across one or more vocabularies. The stream of textual content may be selected from one or more words and / or phrases based on their probability of inclusion in the probability distribution. In some implementations, the one or more vocabularies may include vocabulary specific to a given persona assigned to the automated assistant. In additional or alternative implementations, the one or more vocabularies may include general vocabulary, but the selection of textual content for inclusion in the stream of textual content may be biased toward one or more words and / or phrases specific to a given persona.
[0009] Also, for example, in determining the stream of visual cues described herein, these LLM outputs may include a corresponding probability distribution of a sequence of tokens representing one or more animated body movement gestures that can be performed by a visual representation of an instance of an automated assistant and / or one or more display animations that can be implemented by a display of a client device. The stream of visual cues may be selected from one or more animated body movement gestures and / or one or more display animations based on the probabilities included in the probability distribution and for the sequence of tokens. In some implementations, the one or more animated body movement gestures and / or the one or more display animations may include animated body movement gestures specific to a given persona assigned to the automated assistant. In additional or alternative implementations, the one or more animated body movement gestures may include generic animated body movement gestures, but the selection of animated body movement gestures for inclusion in the stream of text cues may be biased toward one or more animated body movement gestures specific to a given persona.
[0010] In additional or alternative versions of these embodiments, the embodiment may use the LLM to process a stream of audio data capturing the speech, a stream of ASR output of the speech, a stream of NLU output of the speech, and / or the context (if any) of the dialogue session in which the speech was received to generate a given assistant output. In these embodiments, the embodiment may not subsequently modify a given assistant output, since the given assistant output may be generated uniquely for a given persona through the use of the LLM. For example, an instance of the LLM may be pre-trained to generate a given assistant output for a given persona, such that each of multiple heterogeneous personas may be associated with a corresponding instance of the LLM. Also, for example, an LLM may be generic to multiple personas that can be assigned to an instance of an automated assistant, but the LLM may further process given persona data that is specific to a given persona, such as a given persona token, a given persona embedding, a given persona vector, and / or other data that can be utilized to tailor a given assistant response generated using the generic LLM to multiple personas, in generating a given assistant output.
[0011] In various implementations, the implementations may synchronize the auditory rendering of synthetic speech corresponding to the stream of textual context for presentation to the user with the use of the stream of visual cues in controlling the display of the client device and / or controlling the visual representation of the instance of the automated assistant. In some versions of those implementations, such as when a given assistant output is later modified, the stream of modified textual content may be annotated with one or more corresponding visual cue timestamps indicating when the stream of visual cues is used to control the display. For example, the one or more corresponding visual cue timestamps may include a corresponding start visual cue timestamp indicating when to start using one or more of the visual cues included in the stream of visual cues, a corresponding pause visual cue timestamp indicating when to pause using one or more of the visual cues included in the stream of visual cues, a corresponding resume visual cue timestamp indicating when to resume using one or more of the visual cues included in the stream of visual cues, a corresponding end visual cue timestamp indicating when to stop using one or more of the visual cues included in the stream of visual cues, and / or other visual cues.
[0012] By using the techniques described herein, one or more technical advantages can be achieved. As one non-limiting example, the techniques described herein enable an automated assistant to provide not only a more robust and contextually relevant stream of textual content that is synthesized for presentation to a user, but also a robust and contextual stream of visual cues that are utilized to control the display of a client device and / or to control the visual representation of an instance of the automated assistant. As a result, an interaction session between a user and an automated assistant may be better empathized with the user through the utilization of the LLM described herein. As a result, the amount of instances in which a user repeats an utterance and / or an interaction session fails may be reduced, thereby reducing the amount of computational and / or network resources consumed in the user repeating an utterance and / or the interaction session failing.
[0013] An "interaction session," as used herein, may include a logically self-contained exchange between a user and an automated assistant (and possibly other human participants). The automated assistant may distinguish between multiple interaction sessions with a user based on various signals, such as the passage of time between sessions, changes in user context (e.g., location, before / during / after a scheduled meeting, etc.) between sessions, detection of one or more intervening interactions between the user and the client device other than the interaction between the user and the automated assistant (e.g., a user briefly switches applications, a user moves away from a standalone voice-activated product and then returns), locking / sleeping the client device between sessions, changing the client device used to interface with the automated assistant, etc. In particular, during a given interaction session, a user may interact with the automated assistant using various input modalities, including, but not limited to, spoken input, typed input, and / or touch input.
[0014] The above description is provided by way of example only as a summary of some of the embodiments disclosed herein. These and other embodiments are described in further detail herein.
[0015] It should be understood that the techniques disclosed herein may be implemented locally on a client device, remotely by a server(s) connected to the client device via one or more networks, and / or both. [Brief description of the drawings]
[0016] [Figure 1] FIG. 1 illustrates various aspects of the present disclosure and shows a block diagram of an exemplary environment in which the embodiments disclosed herein may be implemented. [Diagram 2] 1 depicts a flowchart illustrating an example method for dynamically adapting a given assistant output based on a given persona assigned to an automated assistant from among multiple disparate personas, according to various implementations. [Diagram 3] 1 shows a flowchart illustrating another example method for dynamically adapting a given assistant output based on a given persona assigned to an automated assistant from among multiple disparate personas, according to various implementations. [Figure 4] 1 shows a flowchart illustrating an example method for training a large-scale language model for use in dynamically adapting a given assistant output based on a given persona assigned to an automated assistant from among multiple disparate personas, according to various embodiments. [Diagram 5] 1 shows a flowchart illustrating an example method for generating persona training instances for use in training a large-scale language model that is used to dynamically adapt a given assistant output based on a given persona assigned to an automated assistant from among multiple heterogeneous personas, according to various embodiments. [Figure 6A]1 illustrates a non-limiting example of dynamically adapting the display of a client device on which an automated assistant is implemented based on a given persona assigned to the automated assistant from among multiple disparate personas, according to various embodiments. [Figure 6B] 1 illustrates a non-limiting example of dynamically adapting the display of a client device on which an automated assistant is implemented based on a given persona assigned to the automated assistant from among multiple disparate personas, according to various embodiments. [Figure 7A] 1 illustrates a non-limiting example of dynamically adapting a visual representation of an automated assistant based on a given persona assigned to the automated assistant from among multiple disparate personas, according to various embodiments. [Figure 7B] 1 illustrates a non-limiting example of dynamically adapting a visual representation of an automated assistant based on a given persona assigned to the automated assistant from among multiple disparate personas, according to various embodiments. [Figure 8] 1 illustrates an exemplary architecture of a computing device in accordance with various embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0017] Referring now to FIG. 1, a block diagram of an example environment 100 is shown that illustrates various aspects of the disclosure and in which the embodiments disclosed herein may be implemented. The example environment 100 includes a client device 110 and a persona system 120. In some implementations, the persona system 120 may be implemented locally at the client device 110. In additional or alternative implementations, the persona system 120 may be implemented remotely from the client device 110 (e.g., at a remote server(s)), as shown in FIG. 1. In these implementations, the client device 110 and the persona system 120 may be communicatively coupled to one another via one or more networks 199, such as one or more wired or wireless local area networks ("LANs" including Wi-Fi LANs, mesh networks, Bluetooth, near field communications, etc.) or wide area networks ("WANs" including the Internet).
[0018] The client device 110 may be, for example, one or more of a desktop computer, a laptop computer, a tablet, a mobile phone, a vehicle computing device (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker (optionally having a display), a smart appliance such as a smart television, and / or a user wearable device that includes a computing device (e.g., a user's watch with a computing device, a user's glasses with a computing device, a virtual reality or augmented reality computing device). Additional and / or alternative client devices may be provided.
[0019] The client device 110 may run an automated assistant client 114. An instance of the automated assistant client 114 may be an application separate from the operating system of the client device 110 (e.g., installed "on top" of the operating system), or alternatively, may be directly implemented by the operating system of the client device 110. The automated assistant client 114 may interact with a persona system 120, which may be implemented locally at the client device 110 or remotely as shown in FIG. 1 and invoked over one or more of the networks 199. The automated assistant client 114 (and optionally through interactions with other remote systems (e.g., server(s))) may form what appears from the user's perspective to be a logical instance of an automated assistant 115 that may engage in human-computer interaction during an interaction session with the user. An instance of the automated assistant 115 is shown in FIG. 1 and is surrounded by a dashed line including the automated assistant client 114 and the natural conversation system 120 of the client device 110. Thus, it should be understood that a user engaging with an automated assistant client 114 running on a client device 110 may actually be engaging with the user's own logical instance of the automated assistant 115 (or a logical instance of the automated assistant 115 shared among a household or other group of users). For brevity and simplicity, the automated assistant 115 as used herein refers to an automated assistant client 114 running locally on the client device 110 and / or remotely on one or more remote servers that may implement the persona system 120.
[0020] In various implementations, client device 110 may include a user input engine 111 configured to detect user input provided by a user of client device 110 using one or more user interface input devices. For example, client device 110 may include one or more microphones that generate a stream of audio data, such as a stream of audio data that captures a user's speech and / or other sounds in the environment of client device 110. Additionally or alternatively, client device 110 may include one or more vision components configured to generate a stream of vision data that captures images, video, and / or certain movements (e.g., gestures) detected within the field of view of one or more of the vision components. Additionally or alternatively, client device 110 may include one or more touch-sensing components (e.g., a keyboard and mouse, a stylus, a touch screen, a touch panel, one or more hardware buttons, etc.) that are configured to generate a signal(s) corresponding to touch input directed at client device 110 (e.g., in implementations in which client device 110 includes a touch screen display).
[0021] In various implementations, client device 110 may include a rendering engine 112 configured to provide content for aural and / or visual presentation to a user of client device 110 using one or more user interface output devices. For example, client device 110 may include one or more speakers that enable it to provide aural content for aural presentation to a user via client device 110. Additionally or alternatively, client device 110 may include a display or projector that enables it to provide visual content for visual presentation to a user via client device 110.
[0022] In various implementations, the client device 110 may include one or more presence sensors 113 configured to provide a signal indicative of a detected presence, particularly a human presence, with the approval of the corresponding user(s). In some of those implementations, the automated assistant 115 may identify the client device 110 (or other computing device associated with the user of the client device 110) and fulfill an utterance based at least in part on the presence of the user at the client device 110 (or other computing device associated with the user of the client device 110). The utterance may be fulfilled by rendering a given assistant output at the client device 110 and / or other computing device(s) associated with the user of the client device 110 (e.g., via the rendering engine 112), by causing the client device 110 and / or other computing device(s) associated with the user of the client device 110 to be controlled, and / or by causing the client device 110 and / or other computing device(s) associated with the user of the client device 110 to perform any other action to fulfill the utterance. As described herein, the automated assistant 115 may utilize data determined based on the presence sensor 113 in determining which client device 110 (or other computing device(s)) the user is near or has recently been near, and provide corresponding commands only to the client device 110 (or those other computing device(s)).In some additional or alternative implementations, the automated assistant 115 may utilize data determined based on the presence sensor 113 in determining whether any user(s) (any user or a particular user) are currently in proximity to the client device 110 (or other computing device(s)), and may optionally refrain from providing data to and / or from the client device 110 (or other computing device(s)) based on the user(s) in proximity to the client device 110 (or other computing device(s)).
[0023] The presence sensor 113 may take a variety of forms. For example, the client device 110 may detect the presence of a user utilizing one or more of the user interface input components described above with respect to the user input engine 111. Additionally or alternatively, the client device 110 may include other types of light-based presence sensors 113, such as a passive infrared ("PIR") sensor that measures infrared ("IR") light emanating from objects within its field of view.
[0024] Additionally or alternatively, in some implementations, the presence sensor 113 may be configured to detect other phenomena related to human presence or device presence. For example, in some embodiments, the client device 110 may include a presence sensor 113 that detects various types of wireless signals (e.g., radio, ultrasonic, electromagnetic, etc. waves) emitted by, for example, other computing devices (e.g., mobile devices, wearable computing devices, etc.) carried / operated by a user and / or by the other computing devices. For example, the client device 110 may be configured to emit waves that are imperceptible to humans, such as ultrasonic or infrared waves that can be detected (e.g., via an ultrasonic / infrared receiver, such as an ultrasound-enabled microphone) by the other computing device(s).
[0025] Additionally or alternatively, client device 110 may emit other types of human-imperceptible waves, such as radio waves (e.g., Wi-Fi, Bluetooth, cellular, etc.), that may be detected by other computing device(s) carried / operated by the user (e.g., mobile devices, wearable computing devices, etc.) and used to determine a particular location of the user. In some implementations, GPS and / or Wi-Fi triangulation may be used to detect a person's location, for example, based on GPS and / or Wi-Fi signals to / from client device 110. In other implementations, other wireless signal characteristics, such as time of flight, signal strength, etc., may be used by client device 110, either alone or collectively, to determine a particular user's location based on signals emitted by other computing device(s) carried / operated by the user. Additionally or alternatively, in some implementations, client device 110 may perform speaker identification (SID) to recognize a user from the user's voice, and / or face identification (FID) to recognize a user from visual data captured of the user's face.
[0026] In some implementations, the speaker's movements may then be determined, for example, by the presence sensor 113 of the client device 110 (and optionally a GPS sensor, Soli chip, and / or accelerometer of the client device 110). In some implementations, based on such detected movements, the user's location may be predicted, which may be presumed to be the user's location when any content is rendered on the client device 110 and / or other computing device(s), based at least in part on the proximity of the client device 110 and / or other computing device(s) to the user's location. In some implementations, the user may simply be presumed to be at the location where the user last engaged with the automated assistant 115, especially if not much time has passed since the last engagement.
[0027] Additionally, client device 110 and / or persona system 120 may include one or more memories for storing data and / or software applications, one or more processors for accessing data and executing software applications, and / or other components that facilitate communication over one or more of networks 199. In some implementations, one or more of the software applications may be installed locally on client device 110, while in other implementations, one or more of the software applications may be hosted remotely (e.g., by one or more servers) and accessible by client device 110 over one or more of networks 199.
[0028] In some implementations, the operations performed by the automated assistant 115 may be implemented locally at the client device 110 via the automated assistant client 114. As shown in FIG. 1, the automated assistant client 114 may include an automatic speech recognition (ASR) engine 130A1, a natural language understanding (NLU) engine 140A1, a fulfillment (LLM) engine 150A1, and a text-to-speech (TTS) engine 160A1. In some implementations, the operations performed by the automated assistant 115 may be distributed across multiple computer systems, such as when the persona system 120 is implemented remotely from the client device 110, as shown in FIG. 1. In these implementations, the automated assistant 115 may additionally or alternatively utilize the ASR engine 130A2, the NLU engine 140A2, the fulfillment engine 150A2, and the TTS engine 160A2 of the persona system 120.
[0029] Each of these engines may be configured to perform one or more functions. For example, ASR engines 130A1 and / or 130A2 may use streaming ASR model(s) stored in machine learning (ML) model(s) database 115A (e.g., recurrent neural network (RNN) models, Transformer models, and / or any other type of ML model capable of performing ASR) to process streams of audio data captured by speech and generated by microphone(s) of client device 110 to generate corresponding streams of ASR output. In particular, the streaming ASR models may be utilized to generate corresponding streams of ASR output as streams of audio data are generated. Further, NLU engine 140A1 and / or 140A2 may process the corresponding stream of ASR outputs using NLU model(s) and / or grammar-based rule(s) stored in ML model(s) database 115A (e.g., long short-term memory (LSTM), gated recurrent unit (GRU), and / or any other type of RNN or other ML model capable of performing NLU) to generate a corresponding stream of NLU outputs. Further, fulfillment engine 150A1 and / or 150A2 may cause the corresponding stream of NLU outputs to be processed to generate a corresponding stream of fulfillment data. For example, the automated assistant 115 may generate and send one or more corresponding structured requests to one or more first party (1P) systems 191 via one or more of the networks 199 (or one or more application programming interfaces (APIs)) and / or to one or more third party (3P) systems 192 via one or more of the networks, receive corresponding fulfillment data from one or more of the 1P systems 191 and / or 3P systems 192, and generate a corresponding stream of fulfillment data.As used herein, one or more 1P systems 191 refer to any systems developed and / or maintained by the same entity that develops and / or maintains the automated assistant 115, and one or more 3P systems refer to any systems developed and / or maintained by an entity separate from the entity that develops and / or maintains the automated assistant 115.
[0030] The one or more corresponding structured requests may include, for example, NLU data included in the stream of corresponding NLU output. The corresponding stream of fulfillment data may correspond, for example, to a corresponding given assistant output predicted to respond to an utterance captured in the corresponding stream of audio data processed by ASR engine 130A1 and / or 130A2. Finally, TTS engine 160A1 and / or 160A2 may process the stream of corresponding text content (e.g., text formulated by automated assistant 115) using TTS model(s) stored in ML model(s) database 115A to generate synthetic speech audio data including computer-generated synthetic speech. The corresponding stream of text content may correspond, for example, to one or more given assistant outputs, one or more of the modified given assistant outputs, and / or any other text content described herein.
[0031] In particular, the ML model(s) stored in ML model(s) database 115A may be an on-device ML model stored locally on client device 110, a remote ML model executed remotely from the client device (e.g., at remote server(s)), or a shared ML model accessible to both client device 110 and / or a remote system (e.g., remote server(s)). In additional or alternative implementations, corresponding streams of synthetic voice audio data corresponding to one or more given assistant outputs, one or more of the modified given assistant outputs, and / or any other text content described herein may be pre-cached in memory or one or more databases accessible by client device 110, thereby eliminating the need for the automated assistant to use TTS engines 160A1 and / or 160A2 to generate the corresponding synthetic voice audio data.
[0032] In various implementations, the corresponding stream of ASR output may include, for example, a stream of ASR hypotheses (e.g., term hypotheses and / or transcription hypotheses) predicted to correspond to the user's utterance(s) captured in the corresponding stream of audio data, one or more corresponding prediction measures (e.g., probabilities, log-likelihoods, and / or other values) for each of the ASR hypotheses included in the stream of ASR hypotheses, a number of phonemes predicted to correspond to the user's utterance(s) captured in the corresponding stream of audio data, and / or other ASR outputs. In some versions of those implementations, ASR engine 130A1 and / or 130A2 may select one or more of the ASR hypotheses as the corresponding recognized text corresponding to the utterance(s).
[0033] In various implementations, the corresponding stream of NLU output may include, for example, a stream of annotated recognized text including one or more annotations of the recognized text for one or more (e.g., all) of the terms of the recognized text, one or more corresponding predictive measures (e.g., probabilities, log-likelihoods, and / or other values) for the NLU output and / or other NLU outputs included in the stream of NLU output. For example, NLU engines 140A1 and / or 140A2 may include a portion of a speech tagger (not shown) configured to annotate terms with their grammatical roles. Additionally or alternatively, NLU engines 140A1 and / or 140A2 may include an entity tagger (not shown) configured to annotate entity references in one or more segments of the recognized text, such as references to people (e.g., including fictional characters, celebrities, public figures, etc.), organizations, places (real and fictional), etc. In some implementations, data about entities may be stored in one or more databases, such as a knowledge graph (not shown). In some implementations, a knowledge graph may include nodes that represent known entities (and possibly entity attributes) and edges that connect the nodes and represent relationships between the entities. An entity tagger may annotate references to entities at a high level of granularity (e.g., to enable identification of all references to an entity class, such as people) and / or at a lower level of granularity (e.g., to enable identification of all references to a particular entity, such as a particular person). An entity tagger may rely on the content of the natural language input to resolve specific entities and / or may optionally communicate with a knowledge graph or other entity database to resolve specific entities.
[0034] Additionally or alternatively, NLU engines 140A1 and / or 140A2 may include a coreference resolver (not shown) configured to group or “cluster” references to the same entity based on one or more contextual cues. For example, the coreference resolver may be utilized to resolve the term “them” in a natural language input of “buy them” to “buy theater tickets” based on the “theater tickets” being mentioned in a client device notification rendered immediately prior to receiving the “buy them” input. In some implementations, one or more components of NLU engines 140A1 and / or 140A2 may rely on annotations from one or more other components of NLU engines 140A1 and / or 140A2. For example, in some implementations, an entity tagger may rely on annotations from a coreference resolver in annotating all references to a particular entity. Also for example, in some implementations, a coreference resolver may rely on annotations from an entity tagger in clustering references to the same entity.
[0035] Although FIG. 1 is described with respect to a single client device with a single user, it should be understood that this is for illustrative purposes and is not intended to be limiting. For example, one or more additional client devices of the user may also implement the techniques described herein. For example, the client device 110, one or more additional client devices, and / or any other computing devices of the user may form an ecosystem of devices that may employ the techniques described herein. These additional client devices and / or computing devices may communicate with the client device 110 (e.g., via the network(s) 199). As another example, a given client device may be utilized by multiple users in a shared configuration (e.g., a group of users, a household). For example, multiple users located at the same location in a household may each utilize a given client device, and each of the multiple users may have a personal separate automated assistant account for each of the multiple users. In these examples, the automated assistant 115 may utilize one or more user identification techniques described herein to identify a given user from among the multiple users, and accordingly adapt any interaction session by utilizing information specific to the given user (e.g., calendar information for the given user, mobile device information from the given user's mobile device (e.g., incoming electronic communications intended for the given user), etc.), and / or other information to adapt the interaction session to the given user, using a given persona that the given user has assigned to the automated assistant 115 (which may be different, for example, from other personas assigned to the automated assistant 115 by other users of the multiple users).
[0036] As described herein, automated assistant 115 may utilize persona system 120 to generate a given assistant output specific to a given persona of multiple different personas assigned to automated assistant 115. Each given assistant output may include, for example, a corresponding stream of textual content determined based on a vocabulary specific to the given persona and synthesized (e.g., using TTS engines 160A1 and / or 160A2) using a set of prosodic features specific to the given persona, and a corresponding stream of visual cues used to control a display of client device 110 and / or a corresponding visual representation of an instance of automated assistant 115 associated with the given persona (e.g., as described with respect to FIGS. 6A and 6B). A given persona that can be assigned to an automated assistant can be embodied by a vocabulary that is specific to the given persona and that is utilized to generate a corresponding stream of textual content, a set of prosodic features specific to the given persona, a corresponding stream of visual cues, and / or a corresponding visual representation of an instance of the automated assistant 115 associated with the given persona. Thus, for additional personas distinct from the given persona, one or more of these may differ to embody the additional personas in a manner distinct from the given persona.
[0037] In various embodiments, as shown in FIG. 1, persona system 120 may additionally or alternatively include persona training engine 170 and persona inference engine 180. Persona training engine 170 may include, for example, training instance engine 171 and training engine 172. Additionally, persona inference engine 180 may include, for example, user identification engine 181, LLM engine 182, output modification engine 183, ranking engine 184, and synchronization engine 185. These various engines of persona system 120 are described in more detail with respect to FIGS. 2-5. While certain engines are shown in FIG. 1, it should be understood that this is for purposes of illustration and is not intended to be limiting. For example, the various engines shown in FIG. 1 may be combined and / or omitted in various embodiments. As one non-limiting example, one or more of the LLM engine 182, the output modification engine 183, the ranking engine 184, and / or the synchronization engine 185 may be combined in an embodiment where a given LLM specific to a given persona is utilized to generate a given assistant output specific to a given persona. As another non-limiting example, the user identification engine 181 may be omitted in an embodiment where the client device 110 is personal to the user (e.g., the user's mobile device).
[0038] In some implementations, persona training instance engine 171 may generate a given persona training instance based on interaction data (e.g., stored in interaction data database 170A) for multiple disparate personas assignable to automated assistant 115, and store the given persona training instance in one or more databases (e.g., training instance(s) database 170B) accessible to persona training engine 170. The interaction data may include any data that may be utilized in a given persona training instance, such as, for example, a corresponding stream of audio data capturing the corresponding utterance, a corresponding stream of ASR output generated based on processing the corresponding stream of audio data, a corresponding stream of NLU output generated based on processing the corresponding stream of ASR output, a corresponding context of a corresponding interaction session in which the corresponding utterance was received, a corresponding stream of visual data capturing a human or character that provided the corresponding utterance, corresponding developer input annotating the interaction data, and / or any other data that may be utilized in a given persona training instance. Generating a given persona training instance based on this interaction data is described in more detail herein (eg, with respect to FIG. 5).
[0039] In these implementations, training engine 172 may train instances of a given LLM (e.g., stored in ML model(s) database 115A) specific to one or more of the multiple heterogeneous personas. For example, a given persona training instance may be indexed (e.g., in training instance(s) database 170B) by each of the multiple heterogeneous personas such that each of the multiple heterogeneous personas is associated with a corresponding set of given persona training instances. Thus, training engine 172 can retrieve (e.g., from training instance(s) database 170B) a corresponding given persona training instance for a given persona for training an instance of the given LLM specific to the given persona, retrieve (e.g., from training instance(s) database 170B) a corresponding given persona training instance for an additional persona for training additional instances of the given LLM specific to additional personas, etc. Training an instance of a given LLM specific to a given persona is described in more detail herein (e.g., with respect to FIG. 4). Thus, in these implementations, the instance of a given LLM may later be utilized to generate a given assistant response specific to a given persona that a user has assigned to an instance of the automated assistant 115, based on which the instance of the given LLM is trained specifically for the given persona.
[0040] In additional or alternative embodiments, a given LLM that is generic to a plurality of disparate personas (e.g., all or a subset of the disparate personas) may be utilized to generate a given assistant response. In these embodiments, the given LLM may additionally process given persona data (e.g., stored in persona data database 170C) for a given persona assigned to automated assistant 115. In these embodiments, the given persona data may correspond, for example, to a given persona token, a given persona embedding, a given persona vector, and / or other data that may be utilized to instantiate a given persona specific to an instance of automated assistant 115. In these embodiments, the given persona data may be defined by a developer associated with automated assistant 115 and / or the persona system and / or may be learned using various machine learning techniques. Thus, in these embodiments, a given LLM can be utilized to generate a given assistant response specific to a given persona that a user has assigned to an instance of the automated assistant 115 based on given persona data specific to the given persona.
[0041] In some implementations, the user identification engine 181 may be utilized to determine the identity of the user who provided the speech captured in the stream of audio data. The user identification engine 181 may determine the identity of the user who provided the speech captured in the stream of audio data based on processing the stream of audio data (e.g., using speaker identification (SID) model(s)), based on processing a stream of visual data capturing the user who provided the speech (e.g., using face identification (FID) model(s)), based on an automated assistant account associated with an active automated assistant at the client device 110, and / or by using other techniques. Identifying the user who provided the speech captured in the stream of audio data is described in more detail herein (e.g., with respect to FIGs. 2 and 3).
[0042] In some implementations, the LLM engine 182 may utilize a given instance of the LLM (e.g., as described with respect to FIG. 3) to generate a given Assistant output provided for direct presentation to a user. In additional or alternative implementations, the LLM engine 182 and / or the output modification engine 183 may then modify a given Assistant output utilizing a given instance of the LLM (e.g., as described with respect to FIG. 2) to generate a given Assistant output provided for direct presentation to a user. Notably, in these implementations, the LLM engine 182 may actively process data in response to receiving a stream of audio data capturing speech. Thus, these implementations may utilize the LLM in an online manner. In additional or alternative implementations, the LLM engine 182 may utilize a previously generated LLM output (e.g., stored in the LLM output(s) database 180) that was previously generated based on the same or similar speech captured in the stream of audio data (e.g., as described with respect to FIG. 2).
[0043] In some implementations, the ranking engine 184 may rank the multiple assistant outputs generated by the LLM engine 182 and / or the output modification engine 183 and select a given assistant output to be provided for presentation to the user based on the ranking. The ranking engine 184 may rank the multiple assistant outputs based on various ranking criteria. The ranking criteria may include, for example, one or more predictive measures (e.g., ASR predictive measures generated when generating a stream of ASR outputs, NLU predictive measures generated when generating a stream of NLU outputs, etc.) indicating how responsive each of the multiple assistant outputs is predicted to be to the utterance, one or more intents included in the corresponding stream of NLU outputs, and / or other ranking criteria.
[0044] In some implementations, the synchronization engine 185 may be used to synchronize a given assistant output for presentation to a user of the client device 110. For example, a stream of visual cues of a given assistant output used to control the display of the client device 110 and / or the visual representation of the automated assistant 115 may be defined with respect to a stream of textual content of the given assistant output that is synthesized for audible presentation to the user of the client device 110. For example, the synchronization may annotate the stream of textual content with corresponding visual cue timestamps that indicate when the visual cues should be used to control the display of the client device 110 and / or the visual representation of the automated assistant 115. Synchronizing a given assistant output for presentation to a user is described in more detail herein (e.g., with respect to FIGS. 2, 3, 6A, 6B, 7A, and 7B).
[0045] 2, a flow chart is shown illustrating an example method 200 of dynamically adapting a given assistant output based on a given persona assigned to an automated assistant from among a plurality of disparate personas. For convenience, the operations of method 200 are described with respect to a system that performs the operations. The system of method 200 includes one or more processors, memory, and / or other component(s) of computing device(s) (e.g., client device 110 of FIG. 1, FIG. 6A, FIG. 6B, FIG. 7A, and FIG. 7B, persona system 120, and / or computing device 810 of FIG. 8, one or more servers, and / or other computing devices). Furthermore, although the operations of method 200 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, and / or added.
[0046] In block 252, the system receives a stream of audio data capturing speech of a user of a client device, the speech being directed at least in part to an instance of an automated assistant running on the client device. The stream of audio data may be generated, for example, via one or more microphones of the client device. In some implementations, the system may receive the stream of audio data in response to an instance of an automated assistant being explicitly invoked, such as based on detection of a particular word or phrase (e.g., "Assistant", "Hey Assistant", etc.) on the client device, based on activation of a button on the client device (e.g., a hardware button on the client device, a software button on a display of the client device), based on detection of a particular gesture on the client device, and / or using other invocation techniques. In other implementations, the system may receive the stream of audio data without an instance of an automated assistant being explicitly invoked, such as based on the stream of audio data being received while a user is gazing at the client device and / or using other techniques.
[0047] In various implementations, the system may implement one or more user identification techniques to identify a user of the client device that provided the utterance. For example, the system may process the stream of audio data using speaker identification (SID) model(s) (e.g., text-dependent (TD) SID model(s) and / or text-independent (TI) SID model(s)) to identify a user of the client device based on previously generated speaker embeddings. Additionally or alternatively, the system may process a stream of visual data generated by one or more visual sensors of the client device used for face identification (FID) model(s) (e.g., immediately before the stream of audio data is received, while the stream of audio data is being received, and / or immediately after receiving the stream of audio data) to identify a user based on previously generated face embeddings. Additionally or alternatively, the system may identify a user based on a user account active on the client device (e.g., a user account associated with an instance of an automated assistant).
[0048] At block 254, the system generates a given Assistant output responsive to the utterance based on processing the stream of audio data, the given Assistant output including (1) a stream of textual content and (2) a stream of visual cues. For example, the system may process the stream of audio data using an ASR model(s) to generate a stream of ASR output, such as one or more recognized terms corresponding to the utterance captured in the stream of audio data. Further, the system may process the stream of ASR output using an NLU model(s) to generate a stream of NLU output, such as one or more intents and slot values corresponding to one or more parameters associated with one or more of the intents.
[0049] In some implementations, the system may generate one or more structure requests to be sent to one or more 1P systems and / or one or more 3P systems based on at least the stream of NLU outputs. In response to sending one or more of the structure requests to one or more 1P systems and / or one or more 3P systems, the system may receive responsive content from one or more 1P systems and / or one or more 3P systems. The responsive content may be utilized in generating a stream of textual content and a stream of visual cues to be included in a given Assistant output. For example, if an utterance corresponds to "Assistant, what's the weather," a stream of audio data capturing the utterance may be processed to obtain a stream of textual content (e.g., "The weather today is rainy and 45") and a stream of visual cues (e.g., weather-related information cards) that are synchronized for audible presentation to the user. In some versions of those implementations, the system may utilize the stream of textual content and the stream of visual cues generated based on the response as a given Assistant output. In additional or alternative versions of those embodiments, the system further processes the stream of textual content and the stream of visual cues generated based on the responsive content from the one or more 1P systems and / or one or more 3P systems and / or the responsive content from the one or more 1P systems and / or one or more 3P systems using an LLM (e.g., in an online manner as described herein) or a previously generated LLM output that previously generated the same or a similar utterance (e.g., in an offline manner as described herein) to generate a given assistant output (e.g., "The weather today is rainy and 45 degrees.In other words, in these embodiments, the streams of textual content and visual cues used as a given assistant output may correspond to textual content and visual content generated using a typical automated-assistant pipeline that does not include any LLM, or may correspond to textual content and visual content generated using a typical automated-assistant pipeline but augmented using an LLM or a previously generated LLM output.
[0050] In additional or alternative embodiments, the system may process the stream of audio data and / or the stream of NLU output without generating or sending a structuring request to one or more 1P systems and / or one or more 3P systems using the LLM (e.g., in an online manner as described herein) or a previously generated LLM output (e.g., in an offline manner as described herein). In other words, in these embodiments, the stream of text content and the stream of visual cues utilized as a given assistant output may be generated directly based on the stream of audio data and / or the stream of NLU output.
[0051] In block 256, the system modifies the given assistant output based on a given assistant persona assigned by the user to the instance of the automated assistant from among a plurality of disparate personas to generate a modified given assistant output, the modified given assistant output including (1) a modified stream of textual content different from the stream of textual content, and / or (2) a modified stream of visual cues different from the stream of visual cues. In generating the modified given assistant output, the system may also take into account the context (if any) of the dialogue session in which the utterance is provided by the user. The given persona may include, for example, a given vocabulary specific to the given persona and utilized to modify the stream of textual content to generate a stream of modified textual content, a given set of prosodic features specific to the given persona and utilized to subsequently synthesize the stream of modified textual content for aural presentation to the user, and / or a given set of visual cues specific to the given persona and utilized to modify the stream of visual cues to generate a stream of modified visual cues. The context of the interaction session may be determined based on one or more context signals, including, for example, the time of day, the day of the week, the location of the client device, ambient noise detected in the environment of the client device, user profile data, software application data, environmental data related to the known environment of the user of the client device, an interaction history of the interaction session between the user and the automated assistant, and / or other context signals, and may be represented in various manners (e.g., vector representations, semantic token representations, and / or other representations). In particular, each of the other disparate personas may also be associated with a corresponding vocabulary, a corresponding set of prosodic features, and / or a corresponding set of visual cues.
[0052] In some implementations, a given persona may be assigned to an instance of an automated assistant when a user initially configures a client device. For example, when setting up a client device, a user may be prompted to select a given persona to be assigned to an instance of an automated assistant from among a plurality of disparate personas. In additional or alternative implementations, a given persona may be assigned to an instance of an automated assistant by a setting of an automated assistant application associated with the instance of an automated assistant. For example, a user may be able to navigate to a setting of an automated assistant application associated with an automated instance of an automated assistant and select a given persona from among a plurality of different personas. In additional or alternative implementations, a given persona may be assigned to an instance of an automated assistant based on a voice command included in an utterance directed to the instance of the automated assistant. For example, a user may provide the utterance "talk to Walter" (e.g., where "Walter" is a reference to a given persona (e.g., a butler persona)) or "Assistant, pretend you are Blackbeard" (e.g., where "Blackbeard" is a reference to a given persona (e.g., a pirate persona). In these examples, the system may process the audio data capturing the utterance to identify a voice command that assigns the given persona to an instance of an automated assistant using various components described herein (e.g., ASR, NLU, fulfillment, etc.).
[0053] In some implementations, a given persona may be assigned to a client device such that multiple different users associated with the client device interact with an instance of the automated assistant that is assigned the given persona. In additional or alternative implementations, a given persona may be assigned to a user of a client device such that other users of the client device may assign different personas to be utilized by the instance of the automated assistant when interacting with the automated assistant through the client device. In these implementations, the identity of the user (e.g., determined as described above with respect to the operation of block 252) may be utilized to determine which persona is assigned to the instance of the automated assistant for multiple different users. In other words, a given persona may be assigned to a client device such that the given persona is utilized when interacting with each user of the client device, or a given persona may be assigned to a particular user of the client device (e.g., the user who provided the utterance captured in the stream of audio data) such that different personas are utilized when interacting with different users of the client device. In particular, a given persona assigned to an instance of an automated assistant may be associated with a visual representation as described herein, whereby a stream of visual cues may be utilized to control both the visual representation and the display of the client device, while other personas that may be assigned to an automated assistant may not have a visual representation, whereby a stream of visual cues may be utilized only to control the display of the client device.
[0054] In some implementations, as shown in block 256A, the modified given assistant output may be generated using the LLM in an online manner (e.g., in response to receiving an utterance and having the LLM process the given assistant output). In these implementations, the system may process the given assistant output using one or more LLMs, and optionally the context of the dialogue session in which the utterance was provided, to generate the modified assistant output. For example, the system may have the LLM engine 182 from FIG. 1 process the stream of text content of the given assistant output, the stream of visual cues of the given assistant output, the stream of audio (e.g., a stream of ASR output, a stream of NLU output, etc.), and / or the context of the dialogue session in which the utterance was provided by the user, using one or more LLMs to generate a set of given modified assistant outputs. In other words, the given assistant output may have a limited vocabulary and a generic personality with respect to the limited vocabulary and prosodic characteristics utilized to synthesize speech. However, when generating the modified given assistant output, the system introduces the given persona into the automated assistant instance to provide the automated assistant instance with a much larger vocabulary specific to the given persona as a function of using one or more LLMs when generating the modified given assistant output, a much larger variance in prosodic features associated with the given persona when generating the modified given assistant output, and a much more robust and interactive visual cue associated with the given persona when generating the modified given assistant output, as compared to the given assistant output. As a result, the modified given assistant output may better resonate with a user engaged in an interaction session with the automated assistant instance.
[0055] In additional or alternative implementations, as shown in block 256B, the modified given output may be generated using previously generated LLM output from the LLM in an offline manner (e.g., in response to receiving the utterance but without having the LLM process the given assistant output). In these implementations, the system may determine that the utterance captured in the stream of audio data corresponds to a previous utterance from which a previously generated LLM output was previously generated, and optionally, that the context of the interaction session in which the utterance was provided by the user corresponds to a previous context of a previous interaction session in which the previous utterance occurred. Additionally, the system may retrieve the previously generated LLM output, since the previously generated LLM output may be indexed based on the previous utterance and the previous context. Similarly, these techniques result in an instance of an automated assistant having a much larger vocabulary specific to a given persona as a function of using one or more LLMs in generating a modified given assistant output, a much larger variance in prosodic features associated with a given persona in generating a modified given assistant output, and much more robust and interactive visual cues associated with a given persona in generating a modified given assistant output, as compared to a given assistant output. As a result, the modified given assistant output may better resonate with a user engaged in an interaction session with the instance of the automated assistant.
[0056] At block 258, the system causes the synthetic speech audio data capturing synthetic speech corresponding to the stream of textual context or the stream of modified textual context to be aurally rendered for presentation to the user. For example, the system may process the stream of textual content using the TTS model(s) to generate synthetic speech audio data capturing synthetic speech corresponding to the stream of textual content (e.g., when the stream of textual content is not modified) or the stream of modified textual content (e.g., when the stream of textual content is modified). In particular, the system may utilize a given set of prosodic characteristics assigned to a given persona in generating the synthetic speech audio data to reflect a given tone, rhythm, pitch, intonation, and / or other prosodic characteristics associated with the given persona. Additionally, the system may cause the synthetic speech audio data to be aurally rendered for presentation to the user via one or more speakers of the client device.
[0057] In block 260, the system causes the stream of visual cues or the modified stream of visual cues to be utilized to control the display of the client device and / or to control a visual representation of the instance of the automated assistant. In some implementations, the stream of visual cues or the modified stream of visual cues is utilized to control the display of the client device. In these implementations, the stream of visual cues or the modified stream of visual cues may cause the display of the client device to visually render one or more screen animations for presentation to the user (e.g., as described with respect to FIG. 6A and FIG. 6B). In additional or alternative implementations, the stream of visual cues or the modified stream of visual cues is utilized to control a visual representation of the instance of the automated assistant. The visual representation of the instance of the automated assistant may be, for example, an avatar corresponding to an animated person (e.g., real or fictional), a character (e.g., a butler, a pirate, a chef), an object (e.g., an animated assistant dot), an animal, and / or any other visual representation. In these implementations, the stream of visual cues or the stream of modified visual cues may cause the visualized representation of the automated assistant to perform one or more animated body gesture movements.These body gesture movements may include, for example, generic body gesture movements that are generic across all visual representations (e.g., waving, typing on the client device display, closing the client device display, etc.), persona-specific body gesture movements that are specific to a given persona assigned to the automated assistant (e.g., a fictional character's special move), emotions (e.g., joy, sadness, anger, etc.) depicted by the visual representation that may be optionally combined with the generic or persona-specific body gesture movements, facial expressions (e.g., smile, frown, etc.) that may be optionally combined with the generic or persona-specific body gesture movements, and / or other body gesture movements. In these implementations, the stream of visual cues or the stream of modified visual cues may cause the visual representation of the automated assistant instance to visually perform animated body gesture movements (e.g., as described with respect to FIG. 7A and FIG. 7B).
[0058] In various implementations, the system may synchronize the aural rendering of synthetic speech corresponding to the stream of visual cues or the stream of modified textual context with the use of the stream of visual cues or the stream of modified visual cues in controlling the display of the client device and / or to control the visual representation of the instance of the automated assistant for presentation to the user. In particular, in implementations that utilize LLMs or LLM outputs to generate a given assistant output and / or a modified given assistant output, this synchronization may be performed automatically by the LLMs (e.g., via LLM engine 182 of FIG. 1 ) trained in the manner described herein (e.g., with respect to FIGS. 4 and 5 ) and / or the LLM outputs generated by those LLMs. In other implementations, this synchronization may be performed by a dedicated synchronization engine and / or other components of the system that can synchronize the stream of visual cues or the stream of modified textual context with the use of the stream of visual cues or the stream of modified visual cues (e.g., by synchronization engine 185 of FIG. 1 ).
[0059] For example, assume that a user utters the utterance "Hi Assistant," assume that the audio capturing the utterance has been processed in the manner described above with respect to the operations of blocks 254 and 256, and assume that the user has assigned a given persona associated with the visual representation to an instance of the automated assistant. In this example, the system:<text_stream=“Hey there, how are you doing?”>
[0023] Assume further that the system generates a stream of textual content or modified textual content represented by a data structure:<visual_stream=hand lift_”Hey there” / body gesture> and<visual_stream=smile ”how are you doing?” / face gesture> Assume further that the stream of visual cues or modified visual cues is generated by a data structure represented by: In this example, the stream of textual content or modified textual content may be synchronized with the stream of visual cues or modified visual cues, resulting in a synchronized data structure: <start=body gesture_hand lift / “Hey there” / end=body gesture_hand lift / start=face gesture_smile / “how are you doing?” / end=face gesture_smile> In particular, in this example, the stream of visual cues or modified visual cues is annotated with visual cue timestamps that indicate when the stream of visual cues or modified visual cues should be utilized in controlling the visual representation of an instance of an automated assistant with respect to the stream of textual content or modified textual content and / or with respect to synthetic speech audio data including synthetic speech corresponding to the stream of textual content or modified textual content.
[0060] For example, the synchronized data structure includes corresponding start visual cue timestamps indicating when to begin utilizing a given visual cue included in the stream of visual cues in controlling the visual representation of an instance of the automated assistant, as indicated by "start=body gesture_hand lift" being defined with respect to a first portion of the textual content (e.g., starting simultaneously or within a threshold time that "Hey there" is audibly presented for presentation to the user), and "start=face gesture_smile" being defined with respect to a second portion of the textual content (e.g., starting simultaneously or within a threshold time that "how are you doing?" is audibly presented for presentation to the user). Additionally, the synchronized data structure includes a corresponding stop visual cue timestamp indicating when to stop utilizing a given visual cue included in the stream of visual cues in controlling the visual representation of the instance of the automated assistant, as indicated by "end=body gesture_hand lift" being defined with respect to the first portion of the textual content (e.g., ending simultaneously or within a threshold time that "Hey there" is audibly presented for presentation to the user) and "start=face gesture_smile" being defined with respect to the second portion of the textual content (e.g., starting simultaneously or within a threshold time that "how are you doing?" is audibly presented for presentation to the user).
[0061] In this case, these timestamps can be automatically annotated to the synchronization data structure (e.g., via the LLM engine 182 of FIG. 1) by an LLM trained in the manner described herein (e.g., with respect to FIGS. 4 and 5). The synchronization can be automatically performed (e.g., via the LLM engine 182 of FIG. 1) by an LLM trained in the manner described herein (e.g., with respect to FIGS. 4 and 5). It should be understood that while the above instances are described with respect to visual cues for controlling the visual representation of an instance of an automated assistant, this is for illustrative purposes and is not intended to be limiting, and that similar visual cues and one or more corresponding visual cue timestamps can additionally or alternatively be utilized to control the display of a client device (e.g., as described with respect to FIGS. 6A and 6B). In another example, since tokens corresponding to the text content or modified stream and tokens corresponding to the stream of visual cues are available during TTS generation, these timestamps can be automatically annotated to the synchronization data structure when the TTS model(s) are used to generate synthetic speech audio data including synthetic speech corresponding to the stream of text content or modified text content. This ensures that the stream of visual cues is synchronized with the rendering of the synthetic speech, as well as with the original stream of text content included in the synthetic speech or modified text content.
[0062] In block 262, the system determines whether an additional stream of audio data has been received. If the system determines in the iteration of block 262 that an additional stream of audio data has not been received, the system continues to monitor the additional stream of audio data in block 262. If the system determines in the iteration of block 262 that an additional stream of audio data has been received, the system returns to block 254 to generate a given additional assistant output based on processing the additional stream of audio data. In particular, the additional stream of audio data may capture additional utterances from a user or additional users of the client device (e.g., determined using various techniques described above with respect to the operation of block 252). Thus, the processing of the additional audio data stream may be adapted in the same or similar manner as described above (e.g., assuming that the additional utterances are provided by additional users, who have assigned additional personas to the instance of the automated assistant).
[0063] In various implementations, if, in an iteration of block 262, the system determines that no additional stream of audio data has been received for a threshold period (e.g., 10 seconds, 30 seconds, 3 minutes, 5 minutes, etc.), the system may generate a stream of visual cues that are utilized to control the visual representation of the instance of the automated assistant and / or to control the display of the client device. This stream of visual cues may be visually rendered to indicate, for example, that the instance of the automated assistant is waiting for the user to provide one or more additional utterances in furthering the dialogue session, that one or more components of the client device remain active to process one or more additional utterances, and / or other information. In particular, this stream of visual cues may be visually rendered, optionally, independent of any synthesized voice audio data and independent of any utterances to which the stream of visual cues may be considered responsive.
[0064] In particular, in an implementation of the method 200 of FIG. 2, a given assistant output may be generated and then adapted to a given persona assigned to an instance of an automated assistant using various post-processing steps. However, it should be understood that this is for illustrative purposes and is not meant to be limiting. For example, as described below with respect to FIG. 3, a given assistant output may be generated and adapted to a given persona assigned to an instance of an automated assistant without these post-processing steps. Furthermore, while the implementation of the method 200 of FIG. 2 is described with respect to a user providing speech, it should be understood that this is also for illustrative purposes and is not meant to be limiting. Rather, it should be understood that the implementation of the method 200 of FIG. 2 may also be utilized in response to a user providing typing and / or touch input.
[0065] Referring now to FIG. 3, a flow chart is shown illustrating another exemplary method 300 of dynamically adapting a given assistant output based on a given persona assigned to an automated assistant from among a plurality of disparate personas. For convenience, the operations of the method 300 are described with reference to a system that performs the operations. The system of the method 300 includes one or more processors, memories, and / or other component(s) of a computing device(s) (e.g., the client device 110 of FIG. 1, FIG. 6A, FIG. 6B, FIG. 7A, and FIG. 7B, the persona system 120, and / or the computing device 810 of FIG. 8, one or more servers, and / or other computing devices). Furthermore, although the operations of the method 300 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, and / or added.
[0066] In block 352, the system receives a stream of audio data capturing speech of a user of a client device, the speech being directed at least in part to an instance of an automated assistant running on the client device. The system may receive the stream of spoken audio data in the same or similar manner as described above with respect to the operations of block 252 of method 200 of FIG. 2. Additionally, the system may optionally process the stream of spoken audio data to identify the user who provided the speech in the same or similar manner as described above with respect to the operations of block 252 of method 200 of FIG.
[0067] In block 354, the system generates a given assistant output based on processing the stream of audio data and using the given LLM, which is responsive to the utterance and specific to a given persona assigned by the user to the automated assistant instance from among multiple heterogeneous personas. The given assistant output includes (1) a stream of text content and (2) a stream of visual cues. In generating the modified given assistant output, the system may also take into account the context (if any) of the dialogue session in which the utterance is provided by the user. As discussed above with respect to method 200 of Figure 2, a given persona may include, for example, a given vocabulary that is specific to the given persona and that is utilized to modify the stream of textual content to generate a stream of modified textual content, a given set of prosodic features that is specific to the given persona and that is utilized to subsequently synthesize the stream of modified textual content for aural presentation to a user, and / or a given set of visual cues that is specific to the given persona and that is utilized to modify the stream of visual cues to generate a stream of modified visual cues. Notably, each of the other disparate personas may also be associated with a corresponding vocabulary, a corresponding set of prosodic features, and / or a corresponding set of visual cues.
[0068] In some implementations, the LLM may be specific to a given persona assigned to an instance of the automated assistant, as shown in block 354A. In these implementations, the LLM specific to a given persona may be trained in the manner described with respect to FIG. 4. For example, the system may identify a given persona assigned to an instance of the automated assistant and retrieve the LLM specific to the given persona assigned to the instance of the automated assistant from one or more databases (e.g., on-device storage of the client device and / or storage remote from the client device but accessible to the client device via one or more networks). Additionally, the system may process the stream of audio data using the ASR model(s) to generate a stream of ASR output, such as one or more recognized terms corresponding to utterances captured in the stream of audio data. Additionally, the system may process the stream of ASR output using the NLU model(s) to generate a stream of NLU output, such as one or more intents and slot values corresponding to one or more parameters associated with one or more of the intents. In this example, the system may process the stream of ASR output, the stream of NLU output, and / or the context (if any) of the dialogue session in which the utterances were provided by the user using the LLM to generate the given assistant output. Thus, the given assistant output generated in block 354 may be specific to the given persona assigned to the instance of the automated assistant, and the LLM utilized to generate the given assistant output is specifically trained for the given persona, and therefore does not require any additional post-processing of the given assistant output to adapt the given assistant output to the given persona assigned to the instance of the automated assistant.
[0069] In some implementations, as shown in block 354B, the LLM may be generic to one or more of the multiple personas. In these implementations, the system may use the LLM to process the stream of ASR output, the stream of NLU output, and / or the context of the interaction session in which the utterances were provided by the user (if any) as described above to generate the given assistant output. However, in these implementations, the system may also use the LLM to process given persona data specific to a given persona assigned to an instance of the automated assistant along with the stream of ASR output, the stream of NLU output, and / or the context of the interaction session. The given persona data may correspond to other given persona data including, for example, a given persona token, a given persona embedding, and / or a representation of a given persona assigned to an instance of the automated assistant. As described herein, the given persona data may be curated by a developer associated with the automated assistant, generated using one or more LLMs, and / or otherwise learned using various techniques. Thus, even though the LLM is generic to one or more of the multiple personas, the given assistant output generated in block 354 may be specific to a given persona assigned to an instance of the automated assistant and does not require any additional post-processing of the given assistant output to adapt it to the given persona assigned to the instance of the automated assistant because the processing using the LLM to generate the given assistant output is adapted to the given persona through utilization of the given persona data.
[0070] In block 356, the system causes synthetic speech audio data capturing synthetic speech corresponding to the stream of textual content or the stream of modified textual content to be audibly rendered for presentation to the user. In block 358, the system causes the stream of visual cues or the stream of modified visual cues to be utilized to control the display of the client device and / or to control the visual representation of the instance of the automated assistant. The operations of blocks 356 and 358 may be performed in the same or similar manner as described with respect to the operations of blocks 258 and 260 of method 200 of FIG. 2, respectively. In various implementations, the system may synchronize the audible rendering of synthetic speech corresponding to the stream of visual cues or the stream of modified textual context for presentation to the user (e.g., as described with respect to method 200 of FIG. 2) with the utilization of the stream of visual cues or the stream of modified visual cues in controlling the display of the client device and / or to control the visual representation of the instance of the automated assistant.
[0071] In block 360, the system determines whether an additional stream of audio data has been received. If in the iteration of block 360, the system determines that an additional stream of audio data has not been received, the system continues to monitor the additional stream of audio data in block 360. If in the iteration of block 360, the system determines that an additional stream of audio data has been received, the system returns to block 354 to generate a given additional assistant output based on processing the additional stream of audio data. The system may receive the additional stream of audio data, perform subsequent processing thereon, and / or monitor the additional stream of audio data in the same or similar manner as described with respect to the operation of block 262 of method 200 of FIG. 2. Notably, in these implementations, in contrast to the implementation of method 200 of FIG. 2, a given assistant output may be generated for a given persona assigned to an instance of an automated assistant without the various post-processing steps described above with respect to method 200 of FIG. 2.
[0072] 4, a flow chart is shown illustrating an example method 400 of training a large-scale language model for use in dynamically adapting a given assistant output based on a given persona assigned to an automated assistant from among a plurality of disparate personas. For convenience, the operations of the method 400 are described with reference to a system that performs the operations. The system of the method 400 includes one or more processors, memory, and / or other component(s) of a computing device(s) (e.g., the client device 110 of FIG. 1, FIG. 6A, FIG. 6B, FIG. 7A, and FIG. 7B, the persona system 120, and / or the computing device 810 of FIG. 8, one or more servers, and / or other computing devices). Furthermore, although the operations of the method 400 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, and / or added.
[0073] At block 452, the system generates a given persona training instance that is utilized to train an instance of a given LLM specific to a given persona from among a plurality of different personas. The given LLM instance may correspond to a generic LLM that has already been trained to generate at least a stream of textual content based on, for example, processing a stream of corresponding ASR output of various utterances, a stream of corresponding NLU data of various utterances, and / or a corresponding context of a corresponding interaction session in which the various utterances are received. However, the given LLM instance may not yet be trained to generate a stream of textual content specific to the given persona and / or may not yet be trained to generate a stream of visual cues for various utterances. In other words, an instance of a given LLM may already be utilized to generate at least a stream of textual content using a generic large vocabulary, but may not have been trained to generate a stream of textual content using the given persona's large vocabulary specific to the given persona, and / or may not have been trained to generate a stream of visual cues for controlling the display of a client device on which an instance of an automated assistant to which the given persona is assigned is running, and / or for controlling a visual representation of the instance of the automated assistant. Thus, the given persona training instance described in the manner herein is one technique for enabling an instance of a given LLM to generate a stream of textual content specific to the given persona, and a stream of visual cues specific to the given persona.
[0074] In some implementations, as indicated at block 452A, the given persona training instance may be generated based on developer input (e.g., as described with respect to methods 500A and 500B of FIG. 5). In additional or alternative implementations, as indicated at block 452B, the given persona training instance may be generated based on analyzing video content from an online multimedia repository (e.g., as described with respect to method 500C of FIG. 5). In other implementations, such as implementations in which the given persona training instance is generated by one or more of the 3P systems 192 of FIG. 1, the given persona training instance may be retrieved from one or more databases.
[0075] At block 454, the system determines whether one or more conditions for training an instance of the given LLM are met. The one or more conditions for training an instance of the given LLM may include a quantity of training instances available for training the instance of the given LLM, a time of day, a day of the week, a period of time that has elapsed since the instance of the given LLM was previously trained via an instance of method 400 of FIG. 4, whether developer input has been received to train the instance of the given LLM, and / or other conditions. If, at an iteration of block 454, the system determines that the one or more conditions for training the given LLM are not met, the system may continue to monitor for the one or more conditions to be met at block 454. In particular, while the system continues to monitor for the one or more conditions to be met at block 454, the system may continue to generate additional given persona training instances that are utilized to train an instance of the given LLM specific to the given persona. If, in an iteration of block 454, the system determines that one or more conditions for training a given LLM are met, the system may proceed to block 456.
[0076] In block 456, the system trains an instance of the given LLM to be used for subsequent processing of audio data capturing speech directed to an instance of the automated assistant to which the given persona is assigned. The system may adapt the manner in which the given instance of the LLM is trained based on the manner in which the given persona training instance is generated, as described with respect to methods 500A, 500B, and 500C of FIG. 5. In block 458, the system trains an instance of the given persona to be used for subsequent processing of audio data capturing speech directed to an instance of the automated assistant to which the given persona is assigned (e.g., as described with respect to methods 300 and 400 of FIG. 3 and FIG. 4, respectively).
[0077] In particular, the system may return to the operations of block 452 to generate additional given persona training instances specific to the given persona to be utilized in further training the given instance of the given LLM specific to the given persona. The system may continue to refine the given LLM instance based on the additional given persona training instances. Additionally, the system may perform additional iterations of the method 400 of FIG. 4 in a parallel or serial manner to train additional instances of the given LLM specific to other corresponding personas that may be assigned to an instance of an automated assistant. In other words, multiple iterations of the method 400 of FIG. 4 may be utilized to generate corresponding instances of the given LLM specific to multiple different personas that may be assigned to an instance of an automated assistant.
[0078] 5, a flow chart is shown illustrating exemplary methods 500A, 500B, and 500C for generating persona training instances for use in training a large-scale language model that is used to dynamically adapt a given assistant output based on a given persona assigned to an automated assistant from among a plurality of disparate personas. For convenience, the operations of methods 500A, 500B, and 500C are described with reference to a system that performs the operations. The systems of methods 500A, 500B, and 500C include one or more processors, memory, and / or other component(s) of computing device(s) (e.g., client device 110 of FIG. 1, FIG. 6A, FIG. 6B, FIG. 7A, and FIG. 7B, persona system 120, and / or computing device 810 of FIG. 8, one or more servers, and / or other computing devices). Furthermore, the operations of methods 500A, 500B, and 500C are shown in a particular order, but this is not meant to be limiting. One or more operations may be reordered, omitted, and / or added. Further, it should be understood that each of methods 500A, 500B, and 500C provides additional or alternative techniques that may be implemented in one or more iterations of the operations of block 452 of method 400 of FIG.
[0079] In some embodiments, at block 552, the system receives developer input that annotates a stream of text content with one or more visual cue timestamps indicating when a stream of visual cues should be utilized relative to the stream of text content for a given persona associated with the auto assistant. For example, assume that the stream of text content is represented by a data structure <text_stream=“Hey there, how are you doing?”>. In this example, the developer input can associate one or more relevant visual cues to one or more portions of the stream of text content, such as by providing developer input that includes a <visual_stream= hand lift_”Hey there” / body gesture> to wave at a visual representation of an instance of the auto assistant while the “Hey there” portion of the stream of text content is being aurally presented, and a <visual_stream=smile”how are you doing?” / face gesture> to smile at the visual representation of the instance of the auto assistant while the “how are you doing” portion of the stream of text content is being aurally presented. Further, in this or other examples where one or more visual cues are already associated with the stream of text content, the developer input can include one or more visual cue timestamps indicating when the stream of visual cues should be utilized relative to the stream of text content by defining when to start, when to pause, when to end the utilization of the visual cues included in the stream of visual cues, and / or other instructions on how to utilize them.In other words, the developer input is as described above with respect to method 200 of FIG. <start=body gesture_hand lift / “Hey there” / end=body gesture_hand lift / start=face gesture_smile / “how are you doing?” / end=face gesture_smile> A synchronization data structure may be provided to synchronize the utilization of a stream of visual cues to a stream of textual content, as represented by the synchronization data structure:
[0080] In these implementations, at block 554, the system generates a given persona training instance based at least on the developer input. In some versions of these implementations, the given persona training instance may include, for example, training instance input and training instance output. Continuing with the above example, the training instance input may include, for example, a stream of ASR output for one or more corresponding utterances, which may result in a stream of text segments, a stream NLU output generated based on the stream of ASR output, and / or one or more corresponding contexts of a corresponding dialogue in which the one or more corresponding utterances were received. The training instance output may include, for example, a stream of text content and a stream of visual cues defined (e.g., via one or more visual cue timestamps) for the stream of text content. In these implementations, when training an instance of the given LLM based on the given persona training instance generated in the manner described with respect to the operation of block 554, the system may cause the instance of the given LLM to process the training instance input and generate a predicted output. The predicted output may include, for example, a predicted stream of text content and a predicted stream of visual cues defined for the predicted stream of text content. In this example, the system may compare a predicted stream of visual cues defined for a predicted stream of textual content with a stream of visual cues defined for a stream of textual content of training instance outputs determined based on developer input to generate one or more losses, and an instance of a given LLM may be updated (e.g., by backpropagation) based on one or more of the losses.In this example, the system may additionally or alternatively compare the predicted stream of textual content to the stream of textual content of the training instances output annotated by the developer to generate one or more additional or alternative losses, and the given instance of the LLM may be updated (e.g., by backpropagation) based on one or more of the additional or alternative losses. In other words, the system may cause an instance of a given LLM to be trained to generate a stream of textual content and a stream of visual cues specific to a given persona and defined for the visual cues.
[0081] In additional or alternative versions of these embodiments, the given persona training instance may include, for example, a mapping between a stream of textual content and a stream of visual cues (and / or one or more visual cue timestamps). In these embodiments, when training an instance of a given LLM based on a given persona training instance generated in the manner described with respect to the operations of block 554, the system may assign a mapping to the stream of textual content such that when the instance of the given LLM subsequently generates a stream of textual content based on subsequent utterances, the stream of visual cues and one or more streams of visual cue timestamps may be obtained and utilized to control the display of the client device and / or the visual representation assigned to the instance of the automated assistant for the given persona. In other words, the system may cause the instance of a given LLM to be trained to generate a stream of textual content and to obtain a stream of visual cues and one or more visual cue timestamps specific to the given persona and defined with respect to the visual cues.
[0082] In an additional or alternative embodiment, at block 556, the system receives developer input from a developer associated with the auto assistant to modify the screen animation as a stream of visual cues regarding a stream of text content for a given persona. For example, again, assume that the stream of text content is represented by a data structure <text_stream=“Hey there, how are you doing?”. In this example, in contrast to the operation of block 552 described above, the developer input is used to control the visualization representation of an instance of the auto assistant with respect to the text content and / or higher-level screen animation used to control the display of the client device, at a higher level (e.g., at least a higher level than the data structure described above with respect to the operation of block 552) of animated body movement gestures that can be defined. For example, the developer input may be provided via a dedicated training platform that enables the developer to modify or move the animated body parts of the visualization representation of an instance of the auto assistant with respect to a stream of text content via one or more input devices (e.g., a mouse and keyboard), and / or to modify or add screen animations used to control the display of the client device. As a non-limiting example, the developer input may include dragging to move with a waving motion of the visualization representation, providing an input of “wave” to move with a waving motion of the visualization representation, providing an input of “jump” to move the visualization representation with a jumping motion, dragging one or more parts of the face of the visualization representation to smile or frown the face of the visualization representation, providing an input of “smile” to smile the arm of the visualization representation, and / or any other type of developer input via the dedicated training platform. In particular, in these embodiments, the developer input may not need to explicitly annotate the stream of text content with one or more visual cue timestamps.
[0083] In these implementations, at block 558, the system generates the given persona training instance based at least on the developer input. In some versions of these implementations, the given persona training instance may include, for example, training instance inputs and training instance outputs as described above (e.g., with respect to the operations of block 554). However, in these implementations, a stream of visual cues and / or one or more visual cue timestamps may be derived from the developer inputs by converting the higher level animated body movement gestures and / or screen animations into the different structures described above to enable an instance of the given LLM to be trained in the same or similar manner as described above with respect to the processing of the given persona training instance in the operations of block 554. In additional or alternative versions of these implementations, the given persona training instance may include, for example, a mapping between a stream of textual content and the higher level animated body movement gestures and / or screen animations. In these implementations, when training an instance of a given LLM based on a given persona training instance generated in the manner described with respect to the operations of block 558, the system may assign mappings to the stream of textual content such that when the instance of the given LLM subsequently generates a stream of textual content based on subsequent utterances, higher level animated body movement gestures and / or screen animations can be obtained and utilized to control the display of a client device and / or a visual representation assigned to an instance of an automated assistant for the given persona.
[0084] In an additional or alternative embodiment, at block 560, the system retrieves video content from an online multimedia repository, including a stream of audio data for the auditory content of the video and a stream of visual data for the visual content of the video. The online multimedia repository may be hosted by one party described herein, such as one or more of the 1P system 191 and / or the 3P system 192 of FIG. 1. For example, the online multimedia repository may be a public online multimedia repository that allows multiple users to upload video content and share it with other users. Also, for example, the online multimedia repository may be a private online multimedia repository accessible only to a user. The video content may capture one or more people or characters emulated, for example, by a virtualized representation of an automated assistant associated with a given persona.
[0085] In these implementations, at block 562, the system processes the stream of audio data to generate a stream of textual content and processes the stream of visual data to generate a stream of visual cues. For example, the system may process the stream of audio data included in the video content using ASR model(s) to generate a stream of ASR output and may process the stream of ASR output using NLU model(s) to generate a stream of NLU output. The stream of textual content may be determined based on at least the stream of ASR output and / or the stream of NLU output as described herein (e.g., with respect to method 200 of FIG. 2, method 300 of FIG. 3, etc.). In this example, the system may also process the stream of visual data using one or more motion tracking machine learning models (e.g., machine learning model(s) trained to track gaze, mouth movement, lip movement, body movement, body posture, etc.) to generate an output indicative of how a person or character captured in the video content visually expresses themselves when speaking. In some versions of these implementations, the stream of visual cues may correspond to an output generated using one or more motion tracking machine learning models. In additional or alternative implementations, the output generated using one or more motion tracking machine learning models may be further processed (e.g., using one or more classification machine learning models) to determine higher level animated body movement gestures and / or screen animations. In particular, the auditory and visual content of the video content are already synchronized when initially processed by being both part of the video content. Thus, one or more video timestamps associated with both the auditory and visual content of the video content may be utilized to assign one or more visual cue timestamps that indicate when the stream of visual cues should be utilized relative to the stream of textual content.
[0086] In these implementations, the system generates a given persona training instance based on at least the stream of textual content and the stream of visual cues in block 564. In some versions of these implementations, the given persona training instance may include, for example, training instance inputs and training instance outputs, and may be utilized to train an instance of a given LLM for a given persona as described above (e.g., with respect to the operations of block 554).
[0087] In particular, these various methods 500A, 500B, and 500C provide different techniques for generating a persona training instance. For example, method 500A of FIG. 5 describes a process of manually defining a given persona training instance by developer input, in that the developer input defines a stream of visual cues and / or one or more visual cue timestamps relative to a stream of textual content. Also, for example, method 500B of FIG. 5 describes a process of semi-manually defining a given persona training instance by developer input, in that the developer input defines a stream of visual cues and / or one or more visual cue timestamps at a higher level relative to the stream of textual content that allows the given persona training instance to be generated. Also, for example, method 500C of FIG. 5 describes a process of automatically defining a given persona training instance by processing video content, in that no developer input is utilized in generating the given persona training instance. Each of these different methods may be used alone or in any combination to generate a given persona training instance and / or additional given persona training instances that are used to train an instance of a given LLM specific to a given persona. For example, methods 500A and 500B may be used to generate a given persona training instance to initially bootstrap an instance of a given LLM, but method 500C may be used to further train and refine the instance of the given LLM to reduce the time and cost associated with generating the given persona training instance. Alternatively, method 500C may be used to initially bootstrap an instance of a given LLM to reduce the time and cost associated with generating the given persona training instance, but methods 500A and 500B may be used to further train and refine the instance of the given LLM to ensure greater accuracy and precision in subsequent uses of the instance of the given LLM.
[0088] While methods 500A, 500B, and 500C of FIG. 5 describe specific techniques for generating a given persona training instance, it should be understood that these specific techniques are provided for illustrative purposes and are not intended to be limiting. Additionally, while methods 500A, 500B, and 500C are generally described with respect to generating a given persona training instance for use in training an instance of a given LLM, it should be understood that this is also for illustrative purposes and is not meant to be limiting. For example, methods 500A, 500B, and / or 500C may also be used to generate additional given persona training instances for use in training an instance of a given LLM specific to a given persona that can be assigned to an instance of an automated assistant. Additionally or alternatively, methods 500A, 500B, and / or 500C may also be used to generate a given additional persona training instance for use in training an additional instance of a given LLM specific to a given additional persona that can also be assigned to an instance of an automated assistant.
[0089] 6A and 6B, various non-limiting examples are shown for dynamically adapting the display of a client device on which an automated assistant is implemented based on a given persona assigned to the automated assistant from among a plurality of different personas. For illustrative purposes, assume that the client device shown in FIGS. 6A and 6B is an instance of the client device 110 of FIG. 1, and that the client device 110 includes a display 190. The display 190 may be, for example, a touch screen display including various portions. For example, the display 190 may include a first portion 190A that includes an indication of user accounts that are active on the client device 110 (e.g., as indicated by a user account symbol on the right side of the first portion 190A of the display 190) and an indication of when various components, such as one or more microphones or voice processing components of the client device 110, are active on the client device 110 (e.g., as indicated by an oval 190A1 on the first portion 190A of the display 190). Additionally, the display 190 may include a second portion 190B that includes a transcription of an interaction between a user of the client device 110 and an instance of an automated assistant implemented at least in part on the client device 110. Additionally, the display may include a third portion 190C that includes a space for visual content (e.g., a home screen) that is provided for visual presentation to the user.
[0090] It should be understood that while the display 190 shown in Figures 6A and 6B includes various disparate portions, this is for illustration purposes and is not meant to be limiting. For example, the disparate portions of the display may overlap one another (e.g., the second portion 190B of the display 190 overlaps the third portion 190C of the display 190) and / or may be omitted in some circumstances (e.g., the second portion 190B of the display 190 may be omitted when the user is not engaged in an interaction session with an instance of an automated assistant). Furthermore, while the client device 110 is shown in Figures 6A and 6B, this is for illustration purposes and is not meant to be limiting, and additional or alternative client devices having displays may implement the techniques described herein. Furthermore, while the examples described with respect to Figures 6A and 6B show specific examples, this is for illustration purposes and is not meant to be limiting, and it should be understood that the techniques described herein may be utilized to control the display of the client device 110 in various other interaction sessions related to various disparate topics.
[0091] In the example of what is described in FIGS. 6A and 6B , assume that a user makes an utterance 652 of “Hey Assistant, what time is it?”, and an instance of the automated assistant processes a stream of audio data capturing the utterance 652 (e.g., generated by one or more microphones of the client device 110) such that a synthesized speech 654 of “Good morning [User]! It's 8:30 AM. Any fun plans today?” is audibly rendered for presentation to the user via one or more speakers of the client device 110 in response to the utterance 652. Assume that the user makes an additional utterance 656 of “Yes, I'm thinking about going to the beach” in response to the synthesized speech 654, and an instance of the automated assistant processes the stream of audio data capturing the utterance 656 to generate a synthesized speech 654 of “Good morning [User]! It's 8:30 AM. Any fun plans today?” Assume that a synthesized speech 658, "Sounds fun! But if you're going to Half Moon Bay again, expect rain and chilly temperatures," is audibly rendered for presentation to the user via one or more speakers of the client device 110 in response to the utterance 656. In this example, an instance of the automated assistant may process the corresponding utterance captured in the corresponding stream of audio data to generate a corresponding stream of text segments using one or more LLMs trained in any manner described herein (e.g., with respect to method 200 of FIG. 2 and / or method 300 of FIG. 3), and optionally, in any manner described herein (e.g., with respect to method 400 of FIG. 5 and / or methods 500A, 500B, and / or 500C of FIG. 5).In some implementations, the automated assistant instance may cause a captured transcription of this interaction session to be provided for visual presentation to the user on the display 190 of the client device 110 (e.g., as shown in the second portion 190B of the display 190 in FIGS. 6A and 6B).
[0092] In this example, when processing a stream of audio data (e.g., generated by one or more microphones of the client device 110) capturing the speech 656, the instance of the automated assistant may utilize a stream of visual cues to control the display 190 of the client device 110. For example, as shown in FIG. 6A, because the speech 656 indicates that the user plans to go to the beach, the stream of visual cues generated when processing the stream of audio data capturing the speech 656 and utilized to control the display 190 of the client device 110 causes the display 190 of the client device 110 to visually present a beach scene for presentation to the user on the display 190 of the client device 110 (e.g., as shown in the third portion 190C of the display 190 of FIG. 6A). The beach scene shown in FIG. 6A may include, for example, a sandy beach with a sand castle, an ocean with waves and fish swimming in the water, a bright and sunny sky with birds flying, and / or other content that a user may typically associate with a beach. In particular, the automated assistant instance may utilize the stream of visual cues to control the display 190 of the client device 110 to visually present a beach scene in association with the stream of textual content captured in the synthetic speech 658, such as the stream of textual content "Sounds fun! But if you go to Half Moon Bay again, expect rain and chilly temperatures" synthesized to generate the synthetic speech 658. In this example, the stream of textual content utilized to generate the synthetic speech 658 and the stream of visual cues utilized to control the display 190 of the client device 110 form a given assistant output responsive to the utterance 656.
[0093] The stream of visual cues in the example of FIG. 6A may include, for example, an indication of the visual content to be presented, an indication of how the visual content should be presented and / or arranged, one or more visual cue timestamps for the visual content to be presented with respect to the stream of textual content captured in the synthetic speech 658, and / or other information to instruct the instance of the automated assistant what to display, how to display it, and when to display it. In some implementations, the visual content provided for visual presentation to the user according to the stream of visual cues may be static visual content. For example, waves, fish, and birds may not move once displayed unless otherwise indicated by the stream of visual cues. In other implementations, the visual content provided for visual presentation to the user according to the stream of visual cues may be dynamic visual content. For example, waves may appear to move and wash up on a beach, fish may appear to swim in the ocean, and birds may appear to fly in the sky, as indicated by the stream of visual cues.
[0094] 6A, the automated assistant instance may use one or more visual cue timestamps, such as an instance of time immediately prior to audibly presenting synthetic speech 658 and / or an instance of time when synthetic speech 658 is first audibly presented, to cause a beach scene to be visually provided for presentation to the user via display 190. However, when synthetic speech 658 is first audibly presented, the automated assistant instance may utilize a stream of visual cues to cause display 190 of client device 190 to adapt the beach scene to be more contextually relevant to the interaction session between the user and the automated assistant instance.
[0095] In particular, the synthetic speech 658 indicates that it will rain and chilly temperatures in Half Moon Bay, a fictional beach that the user often visits when going to the beach. This weather information may be obtained based on the automated assistant instance submitting a query for weather information based on the utterance 656 indicating that the user is going to the beach. Thus, the synthetic speech 658 may also be contextually relevant to the interaction session between the user and the automated assistant instance. However, the beach scene shown in FIG. 6A may not accurately reflect the weather information. As a result, as shown in FIG. 6B, because the weather information indicates that it will rain and chilly temperatures, the stream of visual cues generated when processing the stream of audio data capturing the utterance 656 and utilized to control the display 190 of the client device 110 may visually adapt the display 190 of the client device 110 to present a beach scene to the user on the display 190 of the client device 110 (e.g., as shown in the third portion 190C of the display 190 of FIG. 6B). The beach scene shown in Figure 6B may include, for example, a sandy beach with a sand castle transitions to remove the sand castle because people don't typically build sand castles in the rain, ocean waves may become larger than those shown in Figure 6B, fish swimming in the water may disappear to reflect the large swells and turbulence that accompany the rain, a bright and sunny sky with birds flying may transition to an overcast sky with light rain, and / or other content that a user may typically associate with a rainy beach. In particular, an instance of an automated assistant may utilize the stream of visual cues to control the display 190 of the client device 110 to visually present the beach scene in association with a stream of textual content captured in the synthetic speech 658, such as a stream of textual content "Sounds fun! But if you go to Half Moon Bay again, expect rain and chilly temperatures" synthesized to generate the synthetic speech 658.
[0096] 6A , in the example of FIG. 6B , the automated assistant instance may use one or more visual cue timestamps, such as the instance of time when the “But” portion of the synthetic speech 658 is audibly presented and / or the instance of time when the “expect rain and chilly temperatures” portion of the synthetic speech 658 is audibly presented, to cause the beach scene to be adapted for presentation to the user via the display 190. Thus, when the synthetic speech 658 is audibly presented, the automated assistant instance may utilize a stream of visual cues to control the display 190 of the client device 190 to initially present the sunny beach scene shown in FIG. 6A , but dynamically adapt the sunny beach scene to the rainy beach scene shown in FIG. 6B as it becomes more contextually relevant throughout the interaction session between the user and the automated assistant instance.
[0097] 6A and 6B are described above with respect to a particular stream of visual cues utilized to control the display 190 of the client device 110, it should be understood that this is for illustrative purposes and is not meant to be limiting. Rather, it should be understood that the examples of FIGS. 6A and 6B are provided to illustrate how an instance of an automated assistant may utilize a stream of visual cues in addition to a particular stream of textual content (e.g., captured in synthetic speech 658) in response to an utterance 656. Also, for example, the stream of visual cues may additionally or alternatively be utilized to control a visual representation of an instance of an automated assistant (e.g., as described below with respect to FIGS. 7A and 7B). Furthermore, it should be understood that while the particular stream of visual cues utilized to control the display 190 of the client device 110 is responsive only to an utterance 656, this is for simplicity and is not meant to be limiting. For example, an additional specific stream of visual cues may be generated and utilized with respect to the additional stream of textual content (e.g., captured in the synthetic speech 654) to dynamically present visual content related to a clock that reflects the current time (e.g., 8:30 a.m., as indicated by the synthetic speech 654). Additionally, while FIGS. 6A and 6B are not described with respect to an instance of an automated assistant without a specific persona, it should be understood that this is for illustrative purposes and is not meant to be limiting. Rather, it should be understood that the vocabulary utilized in generating the stream of textual content may be biased toward a given persona vocabulary of a given persona assigned to an instance of an automated assistant, and that the set of prosodic features utilized in generating the synthetic speech may be specific to a given persona assigned to an instance of an automated assistant (e.g., as described in more detail with respect to FIGS. 7A and 7B). Additionally, while FIGS. 6A and 6B are described with respect to a user providing speech as input, it should be understood that this is for illustrative purposes and is not meant to be limiting.Rather, it should be understood that the techniques of FIGS. 6A and 6B may also be employed in situations where the user provides typing and / or a mixture of typing and speech.
[0098] 7A and 7B, various non-limiting examples of dynamically adapting the visualization representation of an automated assistant based on a given persona assigned to the automated assistant from among a plurality of disparate personas are shown. For simplicity, the client devices shown in FIGS. 7A and 7B are the same client devices described above with respect to FIGS. 6A and 6B. For the example illustrating FIG. 7A, assume that the user of the client device 110 has assigned the butler persona to the automated assistant instance. For the example illustrating FIG. 7B, assume that the user of the client device 110 has assigned the pirate persona to the automated assistant instance. As described herein, these different personas assigned to the automated assistant may affect the vocabulary used by the automated assistant instance to generate a stream of text segments during an interaction session between a user and the automated assistant instance, the set of prosodic features used by the automated assistant instance to generate a corresponding synthetic speech that captures the stream of text segments during an interaction session between a user and the automated assistant instance, and / or the set of visual cues used by the automated assistant instance to generate a stream of visual cues during an interaction session between a user and the automated assistant instance.
[0099] Specifically referring to FIG. 7A (e.g., when the butler persona is assigned to the automated assistant), assume that a user of the client device 110 utters the utterance 752, "Hey Assistant, what time is it?", and an instance of the automated assistant processes a stream of audio data (e.g., generated by one or more microphones of the client device 110) capturing the utterance 752 such that a synthesized speech 754A, "Salutations [User]! It's 8:30 AM. How are you keeping on this fine morning?", is audibly rendered for presentation to the user via one or more speakers of the client device 110 in response to the utterance 752. Assume that the user responds to the synthesized speech 754A by making an additional utterance 756A, "I'm doing well and thinking about going to the beach," and that the automated assistant instance processes the stream of audio data capturing the utterance 756A such that synthesized speech 758A, "Excellent to hear sire, but I do caution you to take your umbrella and expect rain at Half Moon Bay," is audibly rendered for presentation to the user via one or more speakers of the client device 110 in response to the utterance 756A.
[0100] Specifically referring to FIG. 7B (e.g., when a pirate persona is assigned to the automated assistant), assume that a user of the client device 110 utters utterance 752, "Hey Assistant, what time is it?", and an instance of the automated assistant processes a stream of audio data (e.g., generated by one or more microphones of the client device 110) capturing the utterance 752 such that a synthesized speech 754B, "Good morning matey! It's 8:30 AM. Where are you setting sail to today?", is audibly rendered for presentation to the user via one or more speakers of the client device 110 in response to the utterance 752. Assume that the user, in response to the synthesized speech 754B, makes an additional utterance 756B, "I'm actually thinking about going to the beach today," and assume that the instance of the automated assistant processes the stream of audio data capturing the utterance 756B such that synthesized speech 758B, "It's going to be raining, but perfect weather to find some treasure at Half Moon Bay!" is audibly rendered for presentation to the user via one or more speakers of the client device 110 in response to the utterance 756B.
[0101] Thus, in these examples, the speech provided by the user of the client device 110 is generally the same between Figures 7A and 7B, but the corresponding streams of text segments utilized to generate the synthetic speech of Figures 7A and 7B are different, and the corresponding streams of visual cues utilized to control the corresponding visualized representations of the different personas are different (e.g., as shown by the first gesture(s) and / or animation(s) 190A2A performed by the butler visualized representation of the automated assistant instance in Figure 7A, and by the second gesture(s) and / or animation(s) 190A2B performed by the pirate visualized representation of the automated assistant instance in Figure 7B). In these examples, the butler persona may correspond to a more polite and conservative persona, and the pirate persona may correspond to a more adventurous and edgy persona. These different personas are exemplified by different vocabulary associated with the different personas, different sets of prosodic features associated with the different personas, and / or different sets of visual cues associated with the different personas.
[0102] For example, in FIG. 7A where a butler persona is assigned to an automated assistant instance, the automated assistant instance may generate a corresponding stream of textual content and a corresponding stream of visual cues as a corresponding given assistant output specific to the butler persona. In some implementations, the automated assistant instance may generate a corresponding given assistant output that is specific to the butler persona through the use of various post-processing operations (e.g., as described with respect to method 200 of FIG. 2) and, optionally, takes into account the use of one or more LLMs specific to the butler persona and / or utilizes butler persona data specific to the butler persona. In additional or alternative embodiments, an instance of the automated assistant may generate a corresponding given assistant output that is specific to the butler's persona through utilization of one or more LLMs specific to the butler's persona and / or utilizes butler persona data specific to the butler's persona (e.g., as described with respect to method 300 of FIG. 3).
[0103] For example, an instance of an automated assistant may process a corresponding stream of audio data capturing the utterance 752, 756A using an ASR model(s) to generate a corresponding stream of ASR output, and may process the corresponding stream of ASR output using an NLU model(s) to generate a corresponding stream of NLU output. The corresponding stream of audio data, the corresponding stream of ASR output, the corresponding stream of NLU output, and / or the context of the interaction session may be processed to generate a corresponding given assistant output (or a corresponding modified given assistant output). In determining the corresponding stream of textual content, the corresponding output generated using the LLM, or previously generated output generated using the LLM (e.g., previously generated based on the same or similar utterances given by the user), may include a corresponding probability distribution of sequences of one or more words and / or phrases across one or more vocabularies, such as a butler persona vocabulary specific to the butler persona, or a generic vocabulary that may be biased towards words and / or phrases associated with the butler persona (e.g., "Salutations" and "How are you keeping on this fine morning?" in synthetic speech 754A, and "sire" and "I do caution" in synthetic speech 758A). Additionally, in generating the corresponding synthetic speech 754A and 758A based on the corresponding stream of textual content, a set of prosodic features associated with the butler persona may be utilized to ensure that the tone, rhythm, pitch, intonation, and / or other prosodic features reflect those of the butler persona. Thus, an instance of the automated assistant may verbally reflect not only the vocabulary that an actual butler might utilize with respect to the corresponding streams of textual content, but also the manner of speech that an actual butler might utilize with respect to the composition of the corresponding streams of textual content, thereby reflecting the polite and conservative personality of the butler persona.
[0104] Additionally, in determining the stream of visual cues, the corresponding output generated using the LLM or previously generated output generated using the LLM (e.g., previously generated based on the same or similar utterances provided by the user) may include a corresponding probability distribution of sequences of tokens representing one or more animated body movement gestures that may be performed by the visualized butler with respect to the corresponding stream of textual content (e.g., as shown in third portion 190C of display 190 in FIG. 7A ) with respect to the corresponding stream of textual content, such as the visualized butler nodding or lifting his monocle, and / or waving to say "Good morning," or holding and opening an umbrella to say "Remember to bring an umbrella," and / or other animations performed with respect to the corresponding stream of textual content (e.g., instructions for controlling display 190 as described above with respect to FIGS. 6A and 6B ). These animated body gesture movements and / or other animations may be synchronized with the corresponding stream of textual content in any manner described herein and / or in other manners. Thus, the visualized butler for an instance of an automated assistant may not only visually reflect the actual butler in terms of appearance on display 190, but may also perform animated body gestural movements that match the actual body gestural movements of the actual butler in terms of the visualized butler being controlled based on a stream of visual cues, thereby further reflecting the polite and conservative personality of the butler persona.
[0105] Similarly, in FIG. 7B where the pirate persona is assigned to an instance of an automated assistant, the automated assistant instance may generate a corresponding stream of textual content and a corresponding stream of visual cues as a corresponding given assistant output specific to the pirate persona. In some implementations, the automated assistant instance may generate a corresponding given assistant output that is specific to the pirate persona through the use of various post-processing operations (e.g., as described with respect to method 200 of FIG. 2) and, optionally, through the use of one or more LLMs specific to the pirate persona, that takes into account the use of one or more LLMs specific to the pirate persona and / or utilizes pirate persona data specific to the pirate persona. In additional or alternative implementations, the automated assistant instance may generate a corresponding given assistant output that is specific to the pirate persona through the use of one or more LLMs specific to the pirate persona and / or utilizes pirate persona data specific to the pirate persona (e.g., as described with respect to method 300 of FIG. 3).
[0106] For example, an instance of an automated assistant may process a corresponding stream of audio data capturing the utterances 752, 756B using an ASR model(s) to generate a corresponding stream of ASR output, and may process the corresponding stream of ASR output using an NLU model(s) to generate a corresponding stream of NLU output. The corresponding stream of audio data, the corresponding stream of ASR output, the corresponding stream of NLU output, and / or the context of the interaction session may be processed to generate a corresponding given assistant output (or a corresponding modified given assistant output). In determining the corresponding stream of textual content, the corresponding output generated using the LLM, or previously generated output generated using the LLM (e.g., previously generated based on the same or similar utterances as given by the user), may include a corresponding probability distribution of sequences of one or more words and / or phrases across one or more vocabularies, such as a pirate persona vocabulary specific to the pirate persona, or a generic vocabulary that may be biased towards words and / or phrases associated with the pirate persona (e.g., “Ahoy matey” and “Where are you setting sail today?” in synthetic speech 754B and “treasure” in synthetic speech 758B). Additionally, in generating the corresponding synthetic speech 754 and 758B based on the corresponding stream of textual content, a set of prosodic features associated with the pirate persona may be utilized to ensure that the tone, rhythm, pitch, intonation, and / or other prosodic features reflect those of the pirate persona. Thus, an instance of the automated assistant can verbally reflect not only the vocabulary that an actual pirate might use with respect to the corresponding streams of textual content, but also the speaking style that an actual pirate might use with respect to the composition of the corresponding streams of textual content, thereby reflecting the adventurous and edgy personality of the pirate persona.
[0107] Additionally, in determining the stream of visual cues, the corresponding output generated using the LLM or previously generated output generated using the LLM (e.g., previously generated based on the same or similar utterances provided by the user) may include a corresponding probability distribution of a sequence of animated body gesture movements that may be performed by the visualized pirate with respect to the corresponding stream of textual content (e.g., as shown in third portion 190C of display 190 in FIG. 7B ), such as jumping up and down, and / or waving when saying "hey buddy," or opening a treasure chest when saying "great weather for treasure hunting," and / or other animations performed with respect to the corresponding stream of textual content (e.g., instructions for controlling display 190 as described above with respect to FIGS. 6A and 6B ). These animated body gesture movements and / or other animations may be synchronized with the corresponding stream of textual content in any manner described herein and / or other manners. Thus, the visualized pirate for an instance of an automated assistant may not only visually reflect the actual pirate in terms of appearance on display 190, but may also perform animated body gesture movements that match the actual body gesture movements of the actual pirate in terms of the visualized pirate being controlled based on a stream of visual cues, thereby further reflecting the adventurous and edgy personality of the pirate persona.
[0108] Although Figures 7A and 7B are described with respect to particular personas (e.g., a butler persona for Figure 7A and a pirate persona for Figure 7B) and with respect to particular visual representations of particular personas, it should be understood that this is for illustrative purposes and is not meant to be limiting. Rather, it should be understood that the personas that may be assigned to an instance of an automated assistant may be virtually limitless. Furthermore, although the particular personas of Figures 7A and 7 are described with respect to particular streams of text segments and particularly streams of visual cues, it should be understood that this is also for illustrative purposes and is not meant to be limiting. Rather, it should be understood that the stream of text segments can be generated based on a virtually infinite vocabulary through the use of the LLM as described herein, and that the stream of visual cues can include persona-specific animated body movement gestures and / or animations (e.g., a visualized butler lifting a monocle, a visualized pirate digging for treasure, etc.), persona-independent animated body movement gestures and / or animations (e.g., a visualized butler smiling, a visualized pirate smiling, etc.), persona-specific emotions for different scenarios (e.g., a visualized butler noting the rain and temperature, a visualized pirate excited by the rain and temperature), corresponding visual cue timestamp(s) for controlling how the stream of visual cues is utilized relative to the stream of textual content, etc. Furthermore, it should be understood that while Figures 7A and 7B are described with respect to controlling the visual representation of an instance of an automated assistant based on a stream of visual cues without controlling a display 190 (e.g., as described with respect to Figures 6A and 6B), this is for brevity and is not meant to be limiting. Rather, it should be understood that these techniques may be combined, for example, the visualized battler of Figure 7A and the visualized pirate of Figure 7B may appear to be located in the sunny beach scene shown in Figure 6A adapted to the rainy beach scene shown in Figure 6B.Additionally, while Figures 7A and 7B are described with respect to a user providing speech as input, it should be understood that this is for purposes of illustration and is not meant to be limiting. Rather, it should be understood that the techniques of Figures 7A and 7B may also be employed where the user provides typing and / or a mixture of typing and speech.
[0109] 8, there is shown a block diagram of an example computing device 810 that may optionally be utilized to perform one or more aspects of the techniques described herein. In some implementations, one or more of the client device, the cloud-based automated assistant component(s), and / or other component(s) may include one or more components of the example computing device 810.
[0110] Computing device 810 typically includes at least one processor 814 that communicates with a number of peripheral devices via a bus subsystem 812. These peripheral devices may include a storage subsystem 824, including, for example, a memory subsystem 825 and a file storage subsystem 826, user interface output devices 820, user interface input devices 822, and a network interface subsystem 816. The input and output devices enable user interaction with computing device 810. The network interface subsystem 816 provides an interface to external networks and is coupled to corresponding interface devices in other computing devices.
[0111] The user interface input devices 822 may include a keyboard, a pointing device such as a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touch screen integrated into a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 810 or a communications network.
[0112] The user interface output devices 820 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 810 to a user or to another machine or computing device.
[0113] Storage subsystem 824 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 824 may include logic to perform selected aspects of the methods disclosed herein, as well as logic to implement the various components illustrated in FIG.
[0114] These software modules are typically executed by the processor 814 alone or in combination with other processors. The memory 825 used by the storage subsystem 824 may include several memories, including a main random access memory (RAM) 830 for storing instructions and data during program execution, and a read-only memory (ROM) 832 in which fixed instructions are stored. The file storage subsystem 826 may provide persistent storage for program files and data files, and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of an embodiment may be stored by the file storage subsystem 826 in the storage subsystem 824 or on other machines accessible by the processor(s) 814.
[0115] Bus subsystem 812 provides a mechanism that allows the various components and subsystems of computing device 810 to communicate with each other as intended. Although bus subsystem 812 is illustrated generally as a single bus, alternative implementations of bus subsystem 812 may use multiple buses.
[0116] Computing device 810 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computing device 810 shown in Figure 8 is intended only as a specific example to illustrate some implementations. Many other configurations of computing device 810 can have more or fewer components than the computing device shown in Figure 8.
[0117] In situations where the systems described herein may collect or otherwise monitor personal information about a user or utilize personal information and / or monitoring information, the user may be provided with an opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social behavior or activity, occupation, user preferences, or the user's current geographic location) or whether and / or how it receives content from a content server that may be more relevant to the user. Additionally, certain data may be handled in one or more ways such that personally identifiable information is removed before it is stored or used. For example, the user's identity may be handled such that information that personally identifies the user cannot be determined, or if geographic location information (such as at the city, zip code, or state level) is obtained, the user's geographic location may be generalized such that the user's specific geographic location cannot be determined. Thus, the user may control how information about the user is collected and / or how the information is used.
[0118] In some embodiments, a method implemented by one or more processors is provided, comprising: receiving a stream of audio data capturing an utterance of a user of a client device, the stream of audio data being generated by one or more microphones of the client device, the utterance being directed at least in part to an instance of an automated assistant executing on the client device; and generating a given assistant output responsive to the utterance based on processing the stream of audio data. The given assistant output includes (i) a stream of textual content and (ii) a stream of visual cues for controlling a display of the client device in response to the utterance and / or for controlling a visual representation of the instance of the automated assistant that is visually rendered for presentation to the user via the display of the client device. The method further includes modifying the given assistant output responsive to the utterance based on a given assistant persona assigned by the user to the instance of the automated assistant from among a plurality of heterogeneous assistant personas to generate a modified given assistant output. The modified given assistant output includes (i) a modified stream of textual content that is different from the stream of textual content, and (ii) a modified stream of visual cues for controlling a display of a client device in response to utterances and / or for controlling a visual representation of an instance of the automated assistant that is different from the stream of visual cues.The method further includes, in response to receiving a stream of audio data capturing speech of a user of the client device, causing synthetic speech audio data capturing synthetic speech corresponding to the stream of modified text context to be audibly rendered for presentation to the user via one or more speakers of the client device, and causing the stream of modified visual cues to be utilized to control a display of the client device and / or to control a visual representation of an instance of the automated assistant.
[0119] These and other implementations of the technology disclosed herein can optionally include one or more of the following features.
[0120] In some implementations, the method may further include synchronizing the auditory rendering of synthetic speech corresponding to the stream of modified textual context for presentation to a user with the use of the stream of modified visual cues in controlling a display of a client device and / or in controlling a visual representation of an instance of an automated assistant.
[0121] In some versions of those implementations, the method may further include annotating the modified text context stream with one or more visual cue timestamps indicating when the stream of visual cues should be utilized to control the display of the client device and / or to control the visual representation of the instance of the automated assistant before the synthetic speech audio data capturing the synthetic speech corresponding to the stream of modified text context is audibly rendered for presentation to the user. Synchronizing the aural rendering of the synthetic speech corresponding to the stream of modified text context and the utilization of the stream of visual cues in controlling the display of the client device and / or to control the visual representation of the instance of the automated assistant may be based on the one or more visual cue timestamps.
[0122] In some further versions of those embodiments, the one or more visual cue timestamps may include at least a start visual cue timestamp indicating when to start utilizing a given visual cue included in the stream of visual cues in controlling the display of the client device and / or to control the visual representation of an instance of the automated assistant, and a stop visual cue timestamp indicating when to stop utilizing a given visual cue included in the stream of visual cues in controlling the display of the client device and / or to control the visual representation of an instance of the automated assistant.
[0123] In some implementations, generating a given Assistant output responsive to the utterance based on processing the stream of audio data may include processing the stream of audio data capturing the utterance using an automatic speech recognition (ASR) model to generate a stream of ASR output, processing the stream of ASR output using a natural language understanding (NLU) model to generate a stream of NLU output, and determining a given Assistant output responsive to the utterance based on at least the stream of NLU output.
[0124] In some versions of those embodiments, determining a given assistant output responsive to the utterance based on the stream of NLU output may include processing the stream of ASR output and / or the stream of NLU output using a large language model (LLM) to determine a stream of text content and a stream of visual cues to be included in the given assistant output.
[0125] In additional or alternative versions of these embodiments, determining a given assistant output responsive to the utterance based on the stream of NLU output may include processing the stream of ASR output and / or the stream of NLU output using a large language model (LLM) output previously generated based on previous instances of the utterance to determine a stream of textual content and a stream of visual cues to be included in the given assistant output.
[0126] In additional or alternative versions of these embodiments, determining a given assistant output responsive to the speech based on the stream of NLU output may include generating one or more structuring requests based on the stream of ASR output and / or the stream of NLU output, sending one or more structuring requests to one or more first party agents and / or one or more third party agents, and determining a stream of textual content and a stream of visual cues to be included in the given assistant output based on content received in response to the one or more structuring requests.
[0127] In some implementations, modifying the given assistant output responsive to the spoken output based on a given assistant persona assigned to the instance of the automated assistant to generate a modified given assistant output may include retrieving, from one or more databases, given persona data specific to the given persona assigned to the instance of the automated assistant, and processing the stream of textual content and the stream of visual cues together with the persona data specific to the given persona assigned to the instance of the automated assistant to generate a modified stream of textual content that differs from the stream of textual content and a modified stream of visual cues that differs from the stream of visual cues.
[0128] In some versions of those implementations, the persona data specific to a given persona assigned to an instance of the automated assistant may include a given persona token specific to a given persona assigned to an instance of the automated assistant and / or a given embedding specific to a given persona assigned to an instance of the automated assistant.
[0129] In additional or alternative versions of these embodiments, processing the stream of textual content and the stream of visual cues together with persona data specific to a given persona assigned to the instance of the automated assistant to generate a modified stream of textual content and a modified stream of visual cues that differ from the stream of textual content and the stream of visual cues that differ from the stream of visual cues may include using a large-scale language model (LLM) to process the stream of textual content and the stream of visual cues together with persona data specific to a given persona assigned to the instance of the automated assistant to generate a modified stream of textual content and a modified stream of visual cues that differ from the stream of textual content and the stream of visual cues.
[0130] In additional or alternative versions of these embodiments, processing the stream of textual content and the stream of visual cues together with persona data specific to a given persona assigned to the instance of the automated assistant to generate a modified stream of textual content and a modified stream of visual cues that differ from the stream of textual content and the stream of visual cues that differ from the stream of visual cues may include processing the stream of textual content and the stream of visual cues together with persona data specific to a given persona assigned to the instance of the automated assistant to generate a modified stream of textual content and a modified stream of visual cues that differ from the stream of textual content and the stream of visual cues that differ from the stream of visual cues using a large-scale language model (LLM) output previously generated using an LLM based on previous instances of the utterance.
[0131] In some implementations, a given persona assigned to an instance of an automated assistant may be associated with a first vocabulary of a plurality of disparate vocabularies used to modify the stream of textual content to generate a stream of modified textual content, a first set of prosodic features of a plurality of disparate sets of prosodic features used to generate synthetic speech audio data that captures synthetic speech corresponding to the stream of modified textual context that is audibly rendered for presentation to the user, and / or a first set of visual cues of a plurality of disparate sets of visual cues used to modify the stream of visual cues to generate a stream of modified visual cues.
[0132] In some versions of those embodiments, a modified stream of textual content different from the stream of textual content may be modified using a first vocabulary, and a modified stream of visual cues different from the stream of visual cues for controlling a display of a client device in response to speech and / or for controlling a visual representation of an instance of an automated assistant may be modified using a first set of visual cues.
[0133] In some further versions of those embodiments, the method may further include processing the stream of modified textual content using a text-to-speech (TTS) model based on the first set of prosodic features to generate synthetic speech audio data.
[0134] In additional or alternative versions of those embodiments, the method may further include receiving an additional stream of audio data capturing an additional utterance of an additional user of the additional client device, the additional stream of audio data being generated by one or more additional microphones of the additional client device, the additional utterance being directed at least in part to an additional instance of an automated assistant executing on the additional client device, the additional utterance being identical to the utterance, and generating a given assistant output responsive to the additional utterance based on processing the additional audio data stream. The given assistant output may include (i) a stream of textual content and (ii) a stream of visual cues for controlling an additional display of the additional client device in response to the additional utterance and / or for controlling an additional visual representation of the additional instance of the automated assistant that is visually rendered for presentation to the additional user via the additional display of the additional client device. The method may further include modifying the given assistant output responsive to the additional utterance based on a given additional assistant persona assigned by the additional user to the additional instance of the automated assistant in addition to the given assistant persona from among a plurality of disparate assistant personas to generate a modified given additional assistant output. The modified given additional assistant output may include (i) a modified additional stream of textual content different from the stream of textual content and different from the modified stream of textual content, and (ii) a modified additional stream of visual cues for controlling an additional display of the additional client device in response to the additional utterance and / or for controlling an additional visual representation of the additional instance of the automated assistant different from the stream of visual cues and different from the modified stream of visual cues.The method may further include, in response to receiving an additional stream of audio data capturing an additional speech of an additional user of the additional client device, causing additional synthetic speech audio data capturing additional synthetic speech corresponding to the modified additional stream of text context to be audibly rendered for presentation to the additional user via one or more additional speakers of the additional client device, and causing the modified additional stream of visual cues to be utilized to control an additional display of the additional client device and / or to control an additional visual representation of the additional instance of the automated assistant.
[0135] In some versions of those embodiments, a given additional persona assigned to an additional instance of the automated assistant may be associated with a second vocabulary of the plurality of heterogeneous vocabularies that is added to the first vocabulary and used to modify the additional stream of textual content to generate a modified additional stream of textual content; a second set of prosodic features of the plurality of heterogeneous sets of prosodic features that is added to the first set of prosodic features and used to generate additional synthetic speech audio data that captures additional synthetic speech corresponding to the modified additional stream of textual context that is audibly rendered for presentation to the additional user; and / or a second set of visual cues of the plurality of heterogeneous sets of visual cues that is added to the first set of visual cues and used to modify the additional stream of visual cues to generate a modified additional stream of visual cues.
[0136] In some implementations, the method may further include receiving an additional stream of audio data capturing an additional utterance of an additional user of the client device, the additional stream of audio data being generated by one or more microphones of the client device, the additional utterance being directed at least in part to an instance of an automated assistant executing on the client device, the additional utterance being identical to the utterance, and generating a given assistant output responsive to the additional utterance based on processing the stream of audio data. The given assistant output may include (i) a stream of text content, and (ii) a stream of visual cues for controlling a display of the client device in response to the additional utterance and / or for controlling an additional visual representation of the instance of the automated assistant that is visually rendered for presentation to the additional user via the display of the client device. The method may further include modifying the given assistant output responsive to the additional utterance based on a given additional assistant persona assigned by the additional user to the given assistant persona plus the additional instance of the automated assistant from among a plurality of heterogeneous assistant personas to generate a modified given additional assistant output. The modified given additional assistant output may include (i) a modified additional stream of textual content that is different from the stream of textual content and different from the modified stream of textual content, and (ii) a modified additional stream of visual cues for controlling a display of the client device in response to the additional utterance and / or for controlling an additional visual representation of an instance of the automated assistant that is different from the stream of visual cues and different from the modified stream of visual cues.The method may further include, in response to receiving an additional stream of audio data capturing an additional speech of an additional user of the client device, causing additional synthetic speech audio data capturing additional synthetic speech corresponding to the modified additional stream of text context to be audibly rendered for presentation to the additional user via one or more speakers of the client device, and causing the modified additional stream of visual cues to be utilized to control a display of the client device and / or to control an additional visual representation of an instance of the automated assistant.
[0137] In some implementations, a user of a client device may assign a given persona to an instance of an automated assistant while initially configuring an automated assistant account for the instance of the automated assistant or while interacting with the assistant settings of an automated assistant application for the instance of the automated assistant.
[0138] In some implementations, the stream of modified visual cues may be used to control the display of a client device in response to speech.
[0139] In some versions of these implementations, the stream of modified visual cues may be further utilized to control the visual representation of an instance of the automated assistant.
[0140] In additional or alternative versions of these implementations, the stream of modified visual cues utilized to control the client device's display in response to speech may include one or more display animations that cause the client device's display to dynamically adapt while the synthesized speech audio data is being audibly rendered for presentation to the user.
[0141] In some implementations, the stream of modified visual cues may be used to control the visual representation of an instance of an automated assistant.
[0142] In some versions of these implementations, the stream of modified visual cues may be further utilized to control the display of the client device in response to speech.
[0143] In additional or alternative versions of these embodiments, the stream of modified visual cues utilized to control the visual representation of the instance of the automated assistant may include one or more animated body gestural movements performed by the visual representation of the instance of the automated assistant while the synthetic speech audio data is audibly rendered for presentation to the user.
[0144] In some embodiments, a method implemented by one or more processors is provided, comprising: receiving a stream of audio data capturing an utterance of a user of a client device, the stream of audio data being generated by one or more microphones of the client device, the utterance being directed at least in part to an instance of an automated assistant executing on the client device; and generating a given assistant output responsive to the utterance and specific to a given persona assigned to the instance of the automated assistant from among a plurality of disparate personas based on processing the stream of audio data and using a given large-scale language model (LLM). The given assistant output includes (i) a stream of textual content specific to the given persona of the instance of the automated assistant, and (ii) a stream of visual cues for controlling a display of the client device in response to the utterance and / or for visually rendering for presentation to a user via a display of the client device and for controlling a visual representation of the instance of the automated assistant that is specific to the given persona assigned to the instance of the automated assistant. The method further includes, in response to receiving a stream of audio data capturing speech of a user of the client device, causing synthetic speech audio data capturing synthetic speech corresponding to the stream of text context to be audibly rendered for presentation to the user via one or more speakers of the client device, and causing the stream of visual cues to be utilized to control a display of the client device and / or to control a visual representation of an instance of the automated assistant.
[0145] These and other implementations of the technology disclosed herein can optionally include one or more of the following features.
[0146] In some implementations, the method may further include identifying a given persona assigned by a user to an instance of the automated assistant from among a plurality of disparate personas, and selecting a given LLM associated with the given persona assigned by the user to the automated assistant from among a plurality of disparate LLMs.
[0147] In some implementations, the method may further include identifying a given persona assigned by a user to an instance of the automated assistant from among a plurality of disparate personas, and selecting given persona data specific to the given persona assigned to the instance of the automated assistant. The given persona data may be processed using the given LLM to generate a given assistant output.
[0148] In some versions of those implementations, the persona data specific to a given persona assigned to an instance of the automated assistant may include a given persona token specific to a given persona assigned to an instance of the automated assistant and / or a given embedding specific to a given persona assigned to an instance of the automated assistant.
[0149] In some embodiments, a method implemented by one or more processors is provided, comprising: receiving a stream of audio data capturing an utterance of a user of a client device, the stream of audio data being generated by one or more microphones of the client device, the utterance being directed at least in part to an instance of an automated assistant executing on the client device; and generating a given assistant output responsive to the utterance based on processing the stream of audio data. The given assistant output includes (i) a stream of textual content and (ii) a stream of visual cues for controlling a display of the client device in response to the utterance and / or for controlling a visual representation of the instance of the automated assistant that is visually rendered for presentation to the user via the display of the client device. The method further includes modifying the given assistant output responsive to the utterance based on a given assistant persona assigned by the user to the instance of the automated assistant from among a plurality of heterogeneous assistant personas to generate a modified given assistant output. The modified given assistant output includes (i) a stream of textual content and (ii) a stream of modified visual cues for controlling a display of a client device in response to speech and / or for controlling a visual representation of an instance of an automated assistant that is different from the stream of visual cues. In response to receiving a stream of audio data capturing speech of a user of the client device, the method further includes causing synthetic speech audio data capturing synthetic speech corresponding to the stream of text context to be audibly rendered for presentation to the user via one or more speakers of the client device, and causing the stream of modified visual cues to be utilized to control a display of the client device and / or control a visual representation of an instance of an automated assistant.
[0150] In some embodiments, a method implemented by one or more processors is provided, comprising: receiving, from a developer associated with an automated assistant, developer input related to one or more visual cues utilized to control a display of a client device with respect to a stream of textual content and / or to control a visual representation of an instance of the automated assistant with respect to the stream of textual content for a given persona assignable to the automated assistant; generating, based at least on the developer input, a given persona training instance utilized to further train an instance of a given large-scale language model (LLM) specific to the given persona from among a plurality of heterogeneous personas, the given LLM having been pre-trained to generate a stream of textual content; training the instance of the given LLM based at least on the given persona training instance; and causing the instance of the given LLM to be utilized for subsequent processing of audio data capturing speech directed to the instance of the automated assistant to which the given persona is assigned.
[0151] These and other implementations of the technology disclosed herein can optionally include one or more of the following features.
[0152] In some implementations, the method may further include receiving, from a developer associated with the automated assistant, additional developer input related to one or more additional visual cues utilized to control the display of the client device with respect to the additional stream of textual content and / or to control an additional visual representation of the additional instance of the automated assistant with respect to the additional stream of textual content for a given additional persona assignable to the automated assistant; generating a given additional persona training instance utilized to further train an additional instance of the given LLM specific to the given additional persona from among the plurality of disparate personas based at least on the additional developer input; training the additional instance of the given additional LLM based at least on the given additional persona training instance; and causing the additional instance of the given additional LLM to be utilized for subsequent processing of additional audio data capturing additional utterances directed to the additional instance of the automated assistant to which the given additional persona is assigned.
[0153] In some implementations, developer input may annotate a stream of textual content with one or more visual cue timestamps that indicate when the stream of visual cues should be used to control the display of a client device with respect to the stream of textual content and / or to control the visual representation of an instance of an automated assistant with respect to the stream of textual content.
[0154] In some versions of these embodiments, the one or more visual cue timestamps may include at least a start visual cue timestamp indicating when to start using a given visual cue included in the stream of visual cues in controlling the display of the client device with respect to the stream of text content and / or controlling the visual representation of an instance of an automated assistant with respect to the stream of text context, and a stop visual cue timestamp indicating when to stop using a given visual cue included in the stream of visual cues in controlling the display of the client device with respect to the stream of text context and / or controlling the visual representation of an instance of an automated assistant with respect to the stream of text context.
[0155] In some implementations, developer input may modify the screen animation on the client device's display with respect to the stream of textual content and / or cause the visual representation of the instance of the automated assistant to perform one or more animated body gesture movements with respect to the stream of textual content.
[0156] In some embodiments, a method implemented by one or more processors is provided, comprising: obtaining video content from an online multimedia repository, the video content including a stream of audio data for auditory content of the video and a stream of visual data for visual content of the video; processing the stream of audio data for the auditory content of the video using an automatic speech recognition model to generate a stream of textual content corresponding to one or more utterances captured in the stream of audio data for the auditory content of the video; processing the stream of visual data for the visual content of the video using one or more motion tracking machine learning models to generate a stream of visual cues; generating a given persona training data instance based on processing the stream of audio data and based on processing the stream of video data, the given persona training data instance being utilized to further train an instance of a given large-scale language model (LLM) specific to a given persona embodied in the video content from among a plurality of heterogeneous personas; training the given LLM instance based at least on the given persona training instance; and causing the given LLM instance to be utilized for later processing additional audio data capturing additional utterances directed to an instance of an automated assistant to which the given persona is assigned.
[0157] Further, some embodiments include one or more processors (e.g., central processing unit(s) (CPU(s)), graphics processing unit(s) (GPU(s)), and / or tensor processing unit(s) (TPU(s)) of one or more computing devices. The one or more processors are operable to execute instructions stored in associated memory, the instructions configured to cause performance of any of the methods described above. Some embodiments also include one or more non-transitory computer-readable storage media storing computer instructions executable by the one or more processors to perform any of the methods described above. Some embodiments also include a computer program product including instructions executable by the one or more processors to perform any of the methods described above.
Claims
1. 1. A method implemented by one or more processors, comprising: Receiving a stream of audio data capturing speech of a user of a client device, the stream of audio data being generated by one or more microphones of the client device, the speech being directed at least in part to an instance of an automated assistant executing on the client device; generating a given assistant output responsive to the utterance based on processing the stream of audio data, the given assistant output including (i) a stream of textual content and (ii) a stream of visual cues for controlling a display of the client device in response to the utterance and / or for controlling a visual representation of the instance of the automated assistant that is visually rendered for presentation to the user via the display of the client device; Modifying the given assistant output responsive to the utterance based on a given assistant persona assigned by the user to the instance of the automated assistant from among a plurality of disparate assistant personas to generate a modified given assistant output, the modified given assistant output including (i) a modified stream of textual content different from the stream of textual content, and (ii) a modified stream of visual cues different from the stream of visual cues for controlling the display of the client device in response to the utterance and / or for controlling the visual representation of the instance of the automated assistant; in response to receiving the stream of audio data capturing the speech of the user of the client device, causing synthesized speech audio data capturing synthesized speech corresponding to the stream of modified textual context to be audibly rendered for presentation to the user via one or more speakers of the client device; causing the modified stream of visual cues to be utilized to control the display of the client device and / or to control the visual representation of the instance of the automated assistant; A method comprising:
2. 2. The method of claim 1, further comprising synchronizing an auditory rendering of the synthetic speech corresponding to the stream of modified textual context for presentation to the user with the use of the stream of modified visual cues in controlling the display of the client device and / or in controlling the visual representation of the instance of the automated assistant.
3. before the synthetic speech audio data capturing the synthetic speech corresponding to the stream of modified text context is aurally rendered for presentation to the user; annotating the stream of modified textual context with one or more visual cue timestamps indicating when the stream of visual cues should be utilized to control the display of the client device and / or to control the visual representation of the instance of the automated assistant; 3. The method of claim 2, wherein synchronizing the auditory rendering of the synthetic speech corresponding to the stream of modified text context and the use of the stream of visual cues in controlling the display of the client device and / or for controlling the visual representation of the instance of the automated assistant is based on the one or more visual cue timestamps.
4. 4. The method of claim 3, wherein the one or more visual cue timestamps include at least a start visual cue timestamp indicating when to start utilizing a given visual cue included in the stream of visual cues in controlling the display of the client device and / or to control the visual representation of the instance of the automated assistant, and a stop visual cue timestamp indicating when to stop utilizing the given visual cue included in the stream of visual cues in controlling the display of the client device and / or to control the visual representation of the instance of the automated assistant.
5. generating the given assistant output responsive to the utterance based on processing the stream of audio data; processing the stream of audio data capturing the speech using an automatic speech recognition (ASR) model to generate a stream of ASR output; processing the stream of ASR outputs using a natural language understanding (NLU) model to generate a stream of NLU outputs; Determining the given assistant output responsive to the utterance based on at least the stream of NLU outputs; The method according to any one of claims 1 to 4, comprising:
6. Determining the given assistant output responsive to the utterance based on the stream of NLU outputs; 6. The method of claim 5, comprising processing the stream of ASR outputs and / or the stream of NLU outputs using a large language model (LLM) to determine the stream of text content and the stream of visual cues included in the given assistant output.
7. Determining the given assistant output responsive to the utterance based on the stream of NLU outputs; 6. The method of claim 5, comprising processing the stream of ASR outputs and / or the stream of NLU outputs using a large language model (LLM) output previously generated based on previous instances of the utterance to determine the stream of text content and the stream of visual cues to be included in the given assistant output.
8. Determining the given assistant output responsive to the utterance based on the stream of NLU outputs; generating one or more structured requests based on the stream of ASR outputs and / or the stream of NLU outputs; sending the one or more structuring requests to one or more first party agents and / or one or more third party agents; Determining the stream of text content and the stream of visual cues to be included in the given assistant output based on content received in response to the one or more structuring requests; The method of claim 5 , comprising:
9. Modifying the given assistant output responsive to the utterance based on the given assistant persona assigned to the instance of the automated assistant to generate a modified given assistant output; Obtaining, from one or more databases, given persona data specific to the given persona assigned to the instance of the automated assistant; Processing the stream of textual content and the stream of visual cues together with persona data specific to the given persona assigned to the instance of the automated assistant to generate the modified stream of textual content that is different from the stream of textual content and the modified stream of visual cues that is different from the stream of visual cues; The method according to any one of claims 1 to 8, comprising:
10. 10. The method of claim 9, wherein the persona data specific to the given persona assigned to the instance of the automated assistant includes a given persona token specific to the given persona assigned to the instance of the automated assistant and / or a given embedding specific to the given persona assigned to the instance of the automated assistant.
11. processing the stream of textual content and the stream of visual cues together with persona data specific to the given persona assigned to the instance of the automated assistant to generate the modified stream of textual content that is different from the stream of textual content and the modified stream of visual cues that is different from the stream of visual cues; 11. The method of claim 9 or 10, comprising using a large language model (LLM) to process the stream of textual content and the stream of visual cues together with the persona data specific to the given persona assigned to the instance of the automated assistant to generate the modified stream of textual content that is different from the stream of textual content and the modified stream of visual cues that is different from the stream of visual cues.
12. processing the stream of textual content and the stream of visual cues together with persona data specific to the given persona assigned to the instance of the automated assistant to generate the modified stream of textual content that is different from the stream of textual content and the modified stream of visual cues that is different from the stream of visual cues; The method of claim 9 or 10, comprising processing the stream of textual content and the stream of visual cues together with persona data specific to the given persona assigned to the instance of the automated assistant using a large language model (LLM) output previously generated using an LLM based on previous instances of the utterance to generate the modified stream of textual content different from the stream of textual content and the modified stream of visual cues different from the stream of visual cues.
13. The given persona assigned to the instance of the automated assistant, a first vocabulary of a plurality of disparate vocabularies utilized in modifying the stream of textual content to generate the modified stream of textual content; a first set of prosodic features of a plurality of disparate sets of prosodic features utilized in generating the synthetic speech audio data capturing the synthetic speech corresponding to the stream of modified text context that is audibly rendered for presentation to the user; and / or a first set of visual cues of a plurality of disparate sets of visual cues that are utilized to modify the stream of visual cues to generate the modified stream of visual cues; The method according to any one of claims 1 to 12,
14. The method of claim 13, wherein the modified stream of textual content, which differs from the stream of textual content, is modified using the first vocabulary, and the modified stream of visual cues, which differs from the stream of visual cues for controlling the display of the client device in response to the utterance and / or for controlling the visual representation of the instance of the automated assistant, is modified using the first set of visual cues.
15. 15. The method of claim 14, further comprising processing the stream of modified textual content based on the first set of prosodic features using a text-to-speech (TTS) model to generate the synthetic speech audio data.
16. Receiving an additional stream of audio data capturing an additional utterance of an additional user of an additional client device, the additional stream of audio data being generated by one or more additional microphones of the additional client device, the additional utterance being directed at least in part to an additional instance of the automated assistant executing on the additional client device, the additional utterance being identical to the utterance; generating the given assistant output responsive to the additional utterance based on processing the additional stream of audio data, the given assistant output including (i) the stream of textual content and (ii) the stream of visual cues responsive to the additional utterance for controlling an additional display of the additional client device and / or for controlling an additional visualized representation of the additional instance of the automated assistant visually rendered for presentation to the additional user via the additional display of the additional client device; Modifying the given assistant output responsive to the additional utterance based on a given additional assistant persona assigned by the additional user to the additional instance of the automated assistant in addition to the given assistant persona from among the plurality of disparate assistant personas to generate a modified given additional assistant output, the modified given additional assistant output including: (i) a modified additional stream of textual content that is different from the stream of textual content and different from the modified stream of textual content; and (ii) a modified additional stream of visual cues for controlling the additional display of the additional client device in response to the additional utterance and / or for controlling the additional visual representation of the additional instance of the automated assistant that is different from the stream of visual cues and different from the modified stream of visual cues; in response to receiving the additional stream of audio data capturing the additional speech of the additional user of the additional client device; causing additional synthetic speech audio data capturing additional synthetic speech corresponding to the modified additional stream of textual context to be audibly rendered for presentation to the additional user via one or more additional speakers of the additional client device; causing the modified additional stream of visual cues to be utilized to control the additional display of the additional client device and / or to control the additional visual representation of the additional instance of the automated assistant; 15. The method of claim 13 or claim 14, further comprising:
17. The given additional persona assigned to the additional instance of the automated assistant, a second vocabulary of the plurality of disparate vocabularies added to the first vocabulary for use in modifying the stream of additional textual content to generate the modified stream of additional textual content; a second set of prosodic features of the plurality of disparate sets of prosodic features that are added to the first set of prosodic features and that are utilized in generating the additional synthetic speech audio data capturing the additional synthetic speech corresponding to the stream of the modified additional text context that is audibly rendered for presentation to the additional user; and / or a second set of visual cues among the plurality of disparate sets of visual cues that are added to the first set of visual cues and that are utilized to modify the stream of additional visual cues to generate the modified stream of additional visual cues; The method of claim 16, wherein the
18. Receiving an additional stream of audio data capturing an additional utterance of an additional user of the client device, the additional stream of audio data being generated by the one or more microphones of the client device, the additional utterance being directed at least in part to the instance of the automated assistant executing on the client device, and the additional utterance being identical to the utterance; generating the given assistant output responsive to the additional utterance based on processing the additional stream of audio data, the given assistant output including (i) the stream of textual content and (ii) the stream of visual cues for controlling the display of the client device in response to the additional utterance and / or for controlling an additional visualized representation of the instance of the automated assistant visually rendered for presentation to the additional user via the display of the client device; Modifying the given assistant output responsive to the additional utterance based on a given additional assistant persona assigned by the additional user to the additional instance of the automated assistant in addition to the given assistant persona from among the plurality of disparate assistant personas to generate a modified given additional assistant output, the modified given additional assistant output including: (i) a modified additional stream of textual content that is different from the stream of textual content and different from the modified stream of textual content; and (ii) a modified additional stream of visual cues for controlling the display of the client device in response to the additional utterance and / or for controlling the additional visual representation of the instance of the automated assistant that is different from the stream of visual cues and different from the modified stream of visual cues; in response to receiving the additional stream of audio data capturing the additional speech of the additional user of the client device; causing additional synthetic speech audio data capturing additional synthetic speech corresponding to the modified additional stream of textual context to be audibly rendered for presentation to the additional user via the one or more speakers of the client device; causing the modified additional stream of visual cues to be utilized to control the display of the client device and / or to control the additional visual representation of the instance of the automated assistant; The method of any one of claims 1 to 17, further comprising:
19. The method of any one of claims 1 to 18, wherein the user of the client device assigns the given persona to the instance of the automated assistant while initially configuring an automated-assistant account for the instance of the automated assistant or while interacting with assistant settings of an automated-assistant application for the instance of the automated assistant.
20. The method of any preceding claim, wherein the stream of modified visual cues is utilized to control the display of the client device in response to the speech.
21. 21. The method of claim 20, wherein the modified stream of visual cues is further utilized to control the visual representation of the instance of the automated assistant.
22. 21. The method of claim 20, wherein the stream of modified visual cues utilized to control the display of the client device in response to the speech includes one or more display animations that cause the display of the client device to dynamically adapt while the synthesized speech audio data is being audibly rendered for presentation to the user.
23. The method of any one of claims 1 to 22, wherein the stream of modified visual cues is utilized to control the visual representation of the instance of the automated assistant.
24. 24. The method of claim 23, wherein the stream of modified visual cues is further utilized to control the display of the client device in response to the speech.
25. 24. The method of claim 23, wherein the stream of modified visual cues utilized to control the visual representation of the instance of the automated assistant includes one or more animated body gestural movements performed by the visual representation of the instance of the automated assistant while the synthetic speech audio data is audibly rendered for presentation to the user.
26. 1. A method implemented by one or more processors, comprising: Receiving a stream of audio data capturing speech of a user of a client device, the stream of audio data being generated by one or more microphones of the client device, the speech being directed at least in part to an instance of an automated assistant executing on the client device; generating a given assistant output responsive to the utterance and specific to a given persona assigned to the instance of the automated assistant from among a plurality of heterogeneous personas based on processing the stream of audio data and using a given large-scale language model (LLM), the given assistant output including: (i) a stream of textual content specific to the given persona of the instance of the automated assistant; and (ii) a stream of visual cues for controlling a display of the client device in response to the utterance and / or for controlling a visual representation of the instance of the automated assistant that is visually rendered for presentation to the user via the display of the client device and that is specific to the given persona assigned to the instance of the automated assistant; in response to receiving the stream of audio data capturing the speech of the user of the client device, causing synthesized speech audio data capturing synthesized speech corresponding to the stream of textual context to be audibly rendered for presentation to the user via one or more speakers of the client device; causing the stream of visual cues to be utilized to control the display of the client device and / or to control the visual representation of the instance of the automated assistant; A method comprising:
27. Identifying the given persona assigned by the user to the instance of the automated assistant from among a plurality of disparate personas; Selecting a given LLM associated with the given persona assigned by the user to the automated assistant from among a plurality of disparate LLMs; 27. The method of claim 26, further comprising:
28. Identifying the given persona assigned by the user to the instance of the automated assistant from among a plurality of disparate personas; Selecting given persona data specific to the given persona assigned to the instance of the automated assistant; Further comprising:
27. The method of claim 26, wherein the given persona data is processed using the given LLM in generating the given assistant output.
29. 29. The method of claim 28, wherein the persona data specific to the given persona assigned to the instance of the automated assistant includes a given persona token specific to the given persona assigned to the instance of the automated assistant and / or a given embedding specific to the given persona assigned to the instance of the automated assistant.
30. 1. A method implemented by one or more processors, comprising: Receiving a stream of audio data capturing speech of a user of a client device, the stream of audio data being generated by one or more microphones of the client device, the speech being directed at least in part to an instance of an automated assistant executing on the client device; generating a given assistant output responsive to the utterance based on processing the stream of audio data, the given assistant output including (i) a stream of textual content and (ii) a stream of visual cues for controlling a display of the client device in response to the utterance and / or for controlling a visual representation of the instance of the automated assistant that is visually rendered for presentation to the user via the display of the client device; Modifying the given assistant output responsive to the utterance based on a given assistant persona assigned by the user to the instance of the automated assistant from among a plurality of disparate assistant personas to generate a modified given assistant output, the modified given assistant output including (i) the stream of textual content and (ii) a modified stream of visual cues for controlling the display of the client device in response to the utterance and / or for controlling the visual representation of the instance of the automated assistant that is different from the stream of visual cues; in response to receiving the stream of audio data capturing the speech of the user of the client device, causing synthesized speech audio data capturing synthesized speech corresponding to the stream of textual context to be audibly rendered for presentation to the user via one or more speakers of the client device; causing the modified stream of visual cues to be utilized to control the display of the client device and / or to control the visual representation of the instance of the automated assistant; A method comprising:
31. 1. A method implemented by one or more processors, comprising: receiving developer input from a developer associated with an automated assistant, for a given persona assignable to the automated assistant, related to one or more visual cues utilized to control a display of a client device with respect to a stream of textual content and / or to control a visual representation of an instance of the automated assistant with respect to the stream of textual content; generating a given persona training instance that is utilized to further train an instance of a given large-scale language model (LLM) specific to the given persona from among a plurality of heterogeneous personas based at least on the developer input, the given LLM having been pre-trained to generate the stream of text content; training the instance of the given LLM based at least on the given persona training instance; The instance of the given LLM is utilized to subsequently process audio data capturing speech directed to the instance of the automated assistant to which the given persona is assigned; A method comprising:
32. receiving additional developer input from the developer associated with the automated assistant, for a given additional persona assignable to the automated assistant, related to one or more additional visual cues utilized to control the display of the client device with respect to an additional stream of textual content and / or to control an additional visual representation of an additional instance of the automated assistant with respect to the additional stream of textual content; generating, based at least on the additional developer input, a given additional persona training instance that is utilized for further training to be utilized for training an additional instance of the given LLM specific to the given additional persona from among the plurality of disparate personas; training the additional instance of the given additional LLM based at least on the given additional persona training instance; The additional instance of the given additional LLM is utilized to subsequently process additional audio data capturing additional utterances directed to the additional instance of the automated assistant to which the given additional persona is assigned; 32. The method of claim 31 , further comprising:
33. The method of claim 31 or claim 32, wherein the developer input annotates the stream of textual content with one or more visual cue timestamps that indicate when the stream of visual cues should be used to control the display of the client device with respect to the stream of textual content and / or to control the visual representation of the instance of the automated assistant with respect to the stream of textual content.
34. 34. The method of claim 33, wherein the one or more visual cue timestamps include at least a start visual cue timestamp indicating when to start utilizing a given visual cue included in the stream of visual cues in controlling the display of the client device with respect to the stream of text content and / or controlling the visual representation of the instance of the automated assistant with respect to the stream of text context, and a stop visual cue timestamp indicating when to stop utilizing the given visual cue included in the stream of visual cues in controlling the display of the client device with respect to the stream of text context and / or controlling the visual representation of the instance of the automated assistant with respect to the stream of text context.
35. 32. The method of claim 31 , wherein the developer input modifies a screen animation on the display of the client device with respect to the stream of textual content and / or causes the visual representation of the instance of the automated assistant to perform one or more animated body gesture movements with respect to the stream of textual content.
36. 1. A method implemented by one or more processors, comprising: obtaining video content from an online multimedia repository, the video content including a stream of audio data for an auditory content of the video and a stream of visual data for a visual content of the video; processing the stream of audio data for the auditory content of the video using an automatic speech recognition model to generate a stream of textual content corresponding to one or more utterances captured in the stream of audio data for the auditory content of the video; processing the stream of visual data about the visual content of the video using one or more motion tracking machine learning models to generate a stream of visual cues; generating, based on processing the stream of audio data and based on processing the stream of video data, a given persona training data instance that is utilized to further train an instance of a given large-scale language model (LLM) specific to a given persona embodied in the video content from among a plurality of heterogeneous personas; training the instance of the given LLM based at least on the given persona training instance; The instance of the given LLM is utilized to subsequently process additional audio data capturing additional utterances directed to the instance of the automated assistant to which the given persona is assigned; and A method comprising:
37. 1. A system comprising: At least one processor; a memory storing instructions that, when executed, cause said at least one processor to perform operations corresponding to any one of claims 1 to 36; A system comprising:
38. A non-transitory computer readable storage medium storing instructions that, when executed, cause at least one processor to perform operations corresponding to any one of claims 1-36.
Citation Information
Patent Citations
Interactive operation-supporting system, interactive operation-supporting method and recording medium
JP2002041276A
User interface / entertainment devices that simulate personal interactions and respond to the user's emotional state and / or personality
JP2004513445A
Information presentation system, information presentation device and information presentation program
JP2005196645A
Voice output controller and voice output control program
JP2017219746A
Voice interface system
JP2020166074A