Chat robot development technology based on unstructured description
By receiving unstructured free-form natural language input and using large language models to generate chat robots, the generalization ability and resource occupation problems of chat robot systems in the prior art when dealing with human languages is solved, and efficient task execution under limited resource conditions is achieved.
Patent Information
- Application Number
- CN202280102323.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-05
- Filing Date
- 2022-12-06
- Publication Date
- 2025-07-04
AI Technical Summary
When existing chatbot systems handle unstructured freeform natural language input, they have functional limitations, making it difficult to generalize the nuances of human voice, and require a large amount of computing resources and storage space, resulting in poor scalability.
By receiving unstructured freeform natural language input, using large language models (LLMs) to generate chatbots without relying on predefined intent patterns or sample libraries, combined with fine-tuning techniques to generate chatbots suitable for specific tasks, which can be quickly deployed and performed on clients or remote systems.
It realizes the rapid generation and deployment of chatbots under limited computing resources, can efficiently perform complex tasks, reduce computing resource usage, and improve the ability to generalize human languages.
Smart Images

Figure CN120266122A_ABST
Abstract
Description
Background Art
[0001] Humans can engage in human-machine conversations via various computing devices using interactive software applications known as "chatbots", "voicebots", "automated assistants", "interactive personal assistants", "intelligent personal assistants", "conversational agents", etc. As an example, these chatbots can correspond to machine learning models or combinations of different machine learning models and can be used to perform various tasks on behalf of a user. For example, some of these chatbots can talk to various people to perform actions on behalf of another person or entity. In some of these cases, the conversation can include voice-based conversations, such as conversations conducted locally at a computing device, conversations conducted remotely on multiple computing devices via a telephone network, or conversations in other voice-based scenarios. In other cases, the conversation can include text-based conversations, such as conversations conducted via text or SMS messaging, email, and / or other text-based scenarios.
[0002] However, the functionality of some of these chatbots may be limited in various ways. For example, the functionality of some of these chatbots may be restricted by predefined intent patterns used by the chatbot to perform actions. In other words, if a person participating in a given conversation with a given chatbot provides an utterance that is determined to include an intent not defined by the predefined intent pattern, the given chatbot may fail. Further, to update these chatbots, existing intent patterns can be modified or new intent patterns can be added. As another example, the functionality of some of these chatbots may be restricted by the corpus of examples used to train the chatbot. In other words, if a person participating in a given conversation with a given chatbot provides an utterance that is not included in the given corpus of examples, the given chatbot may fail. Further, to update these chatbots, existing examples in the corpus can be modified or new examples can be added. However, in both of these examples, there are almost infinite intent patterns and / or examples that may need to be previously defined to make the robot robust to various nuances of human speech and reduce the cases of failure.
[0003] It should be noted that a large amount of computing resources are required to manually define and / or manually refine such intent patterns and / or examples. Further, even if a large number of intent patterns and / or examples are defined, a large amount of memory is required to store and / or utilize a large number of intent patterns for these chatbots, and / or to train these chatbots based on a large number of examples in the corpus. Therefore, the intent patterns for rule-based chatbots and the examples for example-based chatbots are not actually scalable to the extent of learning the nuances of human speech. Summary of the Invention
[0004] The implementation involves: receiving an unstructured free-form natural language input; generating a chatbot based on the unstructured free-form natural language input and in response to receiving the unstructured free-form natural language input; and causing the chatbot to perform a task associated with an entity on behalf of the user. In some versions of those implementations, the unstructured free-form natural language input conveys details of the task to be performed, but does not define any corresponding dialogue state diagram (e.g., does not define any dialogue states or any dialogue state transitions). Further, the unstructured free-form natural language input is not provided according to any pattern. Nevertheless, the unstructured free-form natural language input can be used to fine-tune a machine learning model (ML) that is already capable of being used for more generalized conversations, and / or can be used as an input across ML models without fine-tuning the ML model. As a result, the chatbot can be generated and deployed in a quick and efficient manner to perform tasks on behalf of the user.
[0005] For example, assume that an unstructured free-form natural language input corresponding to the spoken utterance "ask the restaurant I ate dinner at yesterday if they found my red leather jacket" is received at a user's client device. The processor may use various automatic speech recognition (ASR), natural language understanding (NLU), and / or fulfillment techniques to determine that the unstructured free-form natural language input includes a task to be performed by the chatbot. In this example, the task may include, for instance, submitting a query of "did you find [the user's] red leather jacket" to "the restaurant [the user] ate dinner at yesterday". Further, the processor may generate the chatbot to participate in a corresponding conversation with a representative of the "restaurant" to submit the query, and determine responsive content to be provided to the user for presentation to the client device based on the corresponding conversation with the representative of the "restaurant". Although the above example is described with respect to a relatively simple query (e.g., submitting a query), it should be understood that this is for illustrative purposes, and the same or similar techniques may be used to perform relatively complex tasks, which may include multiple other tasks or subtasks that may be performed during the corresponding conversation.
[0006] In various implementations, the corresponding conversation can be a voice-based conversation in which the chatbot participates in the corresponding conversation by making a phone call or locally at the client device. In these implementations, the chatbot can additionally or alternatively be referred to as a voicebot. In other implementations, the corresponding conversation can be a text-based conversation in which the chatbot participates in the corresponding conversation using text messaging or SMS services, email services, or other text-based services. Thus, the chatbot can be deployed in various environments to participate in corresponding conversations with various entities to perform various tasks. In some implementations, the chatbot generated to perform various tasks can be fine-tuned in an "instant" manner such that the chatbot is generated in response to receiving an unstructured free-form natural language input. Although some chatbots can be used (e.g., in parallel or serially) to participate in corresponding conversations with multiple different entities, the corresponding chatbot can be generated based on each instance of the unstructured free-form natural language input provided by the user that includes the task to be performed on behalf of the user. In other implementations, the generated chatbots may not be fine-tuned in an "instant" manner, but they can be identified and utilized in response to receiving an unstructured free-form natural language input.
[0007] In various implementations, the processor can be implemented locally at the user's client device, where an unstructured free-form natural language input is received that conveys details of the task to be performed. In these implementations, the processor can obtain a previously trained large language model (LLM) from the on-device storage of the client device as an ML model that is already capable of being used for more generalized conversations. Further, the processor can fine-tune the previously trained LLM based on the unstructured free-form natural language input to generate a fine-tuned LLM. Additionally, the processor uses the fine-tuned LLM as a chatbot to participate in the corresponding conversation on behalf of the user. In other versions of these implementations, the processor can obtain a previously trained LLM from the on-device storage of the client device as an ML model that is already capable of being used for more generalized conversations, but avoid fine-tuning the previously trained LLM based on the unstructured free-form natural language input.
[0008] In other implementations, the processor can be implemented remote from the user's client device (e.g., at a remote system such as a high-performance server or a high-performance server cluster). In these implementations, the processor can obtain a previously trained large language model (LLM) from the remote storage of the remote system as an ML model that is already capable of being used to conduct more generalized conversations. Further, the processor can generate a fine-tuned LLM and use the fine-tuned LLM as a chatbot to represent the user in corresponding conversations. Continuing with the above example, the processor can cause the chatbot to process corresponding data such that a query is submitted during the corresponding conversation to determine whether the user left his or her jacket at the restaurant where he or she had dinner the previous night. In some versions of these implementations, the processor can obtain a previously trained LLM from the remote storage of the remote system as an ML model that is already capable of being used to conduct more generalized conversations, but avoid fine-tuning the previously trained LLM based on unstructured free-form natural language input.
[0009] Notably, the previously trained LLM can correspond to existing LLMs such as LaMDA, BERT, Meena, GPT-3, and / or any other previously trained LLM. These previously trained LLMs have been previously trained on a large amount of diverse data and are capable of participating in corresponding conversations with users in a natural and intuitive manner. However, these LLMs have multiple ML layers and hundreds of millions to hundreds of billions of ML parameters. Thus, in implementations where the fine-tuned chatbot is locally generated at the client device, the previously trained LLM that is obtained and fine-tuned can be a sparsified version of the previously trained LLM. In contrast, in implementations where the fine-tuned chatbot is generated remote from the client device, the previously trained LLM that is obtained and fine-tuned can be an unsparsified version of the previously trained LLM. Compared to the almost infinite resources of the remote system, due to various hardware constraints and / or software constraints at the client device, the sparsified version of the previously trained LLM can have fewer ML layers, fewer ML parameters, masked weights, and / or other sparsification aspects to reduce the size of the previously trained LLM.
[0010] In some implementations, and in enabling a chatbot to participate in a corresponding conversation, the processor may use a fine-tuned LLM corresponding to the chatbot to process task data (e.g., using an ASR model, NLU model, and / or fulfillment model or rules and output generated based on processing unstructured free-form natural language input), audio data of any verbal input representing an entity, text data predicted to correspond to the audio data of any verbal input representing an entity, any conversation context data, and / or any other data described herein, to generate an output. In other implementations, and in enabling a chatbot to participate in a corresponding conversation, the processor may use a previously trained (e.g., untuned) LLM to process natural language descriptions, task data, audio data of any verbal input representing an entity, text data predicted to correspond to the audio data of any verbal input representing an entity, any conversation context data, and / or any other data described herein included in the unstructured free-form natural language input, to generate an output. For example, the output may be a probability distribution over a vocabulary or sequence of terms and / or phrases. Based on the probability distribution over the vocabulary or sequence of terms and / or phrases, the processor may select an instance of the text data corresponding to the text and / or speech to be provided by the chatbot.
[0011] In implementations where the corresponding conversation is a text-based conversation, the processor may cause an instance of the text data to be visually rendered for presentation to the representative of the entity at the client device and / or an additional client device of the entity. However, in implementations where the corresponding conversation is a voice-based conversation, the processor may cause the chatbot to use a text-to-speech (TTS) model to process an instance of the text data corresponding to generating an instance of the synthesized speech audio data, the instance of the synthesized speech audio data capturing the synthesized speech corresponding to the text data. Further, the processor may cause the instance of the synthesized speech audio data to be visually rendered for presentation to the representative of the entity at the client device and / or an additional client device of the entity. Notably, in implementations where the chatbot corresponds to a previously trained LLM fine-tuned based on unstructured free-form natural language input, the chatbot is capable of generating a conversation output focusing on the task to be performed based on the unstructured free-form natural language input. Further, in implementations where the chatbot corresponds to a previously trained LLM that is not fine-tuned, the chatbot is still capable of generating a conversation output focusing on the task specified by the unstructured free-form natural language input, since the unstructured free-form natural language input is still applied as input across the previously trained LLM that is not fine-tuned.
[0012] In various implementations, the processor may cause responsive content to be provided for presentation to a user of a client device that provided unstructured free-form natural language input. The responsive content may be determined based on one or more responses provided by a representative of an entity during a corresponding conversation. Further, the responsive content may include, for example, corresponding results of one or more tasks determined during the corresponding conversation, a corresponding summary of the corresponding conversation, and / or other content. Continuing with the above example, the chatbot may determine, in response to a response provided by a representative of an entity and during a corresponding conversation, whether the user left his or her red leather jacket at the restaurant.
[0013] In various implementations, and during a corresponding conversation, the chatbot may engage in a corresponding conversation with a representative of an entity using one or more peripheral behaviors. These peripheral behaviors may include, for example, a greeting behavior that enables the chatbot to identify the user and / or identify itself as a chatbot, a hold behavior that enables the chatbot to pause and resume the corresponding conversation, a bailout behavior that enables the chatbot to terminate the corresponding conversation with the representative of the entity, and / or other peripheral behaviors. These peripheral behaviors are some non-limiting examples of why a previously trained LLM enables the chatbot to perform general aspects of a conversation and unstructured free-form natural language input need not specify these general aspects of the conversation that the chatbot is capable of performing. However, a fine-tuned chatbot fine-tuned based on unstructured free-form natural language input enables the chatbot to perform tasks on behalf of the user while still being able to perform these general aspects of the conversation.
[0014] In various implementations, in response to determining that one or more conditions are met, the processor may discard a chatbot generated based on unstructured free-form natural language input. The processor may discard the chatbot based on, for example, whether the chatbot successfully performed a task associated with an entity, whether the chatbot will be used to perform a task associated with an additional entity, whether a threshold duration has elapsed since the chatbot was generated, whether, in an implementation where the chatbot is locally generated at the client device, the chatbot consumed a threshold amount of the device storage of the client device, whether, in an implementation where the chatbot is locally generated at the client device, the threshold amount of the device storage of the client device is available while the chatbot is stored in the device storage, and / or based on other conditions. In other words, in an implementation where the chatbot is locally generated at the client device, in determining whether to discard the chatbot, the system may balance the performance of the chatbot and how the chatbot affects the client device.
[0015] By using the techniques described herein, various technical advantages can be achieved. As a non-limiting example, the techniques described herein enable a processor of a client device and / or a remote system to generate a chatbot based on unstructured free-form natural language input to perform tasks specified by a user, and / or to utilize an existing chatbot based on unstructured free-form natural language input to perform tasks specified by a user. These tasks can be specified by a natural language description provided by the user. This enables the process to generate and deploy chatbots in a fast and efficient manner to perform tasks on behalf of the user. Further, in some examples, the chatbot can be discarded after performing the task on behalf of the user. Thus, in implementations where the chatbot is implemented locally at the client device, the techniques described herein balance the current and / or future performance of the chatbot with respect to how the chatbot can affect the performance of the client device. Further, in implementations where the chatbot is implemented away from the client device (e.g., by a remote system communicatively coupled to the client device), tasks can still be performed on behalf of the user while computational resources are saved at the client device.
[0016] The above description is provided as an overview of only some of the implementations disclosed herein. Those implementations, as well as other implementations, are described in more detail herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A block diagram depicting an example environment that illustrates various aspects of the present disclosure and in which the implementations disclosed herein can be implemented.
[0018] Figure 2 A flowchart depicting an example process flow for generating a chatbot and enabling the chatbot to participate in a conversation with an entity according to various implementations.
[0019] Figure 3 A flowchart depicting an example method for locally generating a chatbot at a client device and enabling the chatbot to participate in a conversation with an entity according to various implementations.
[0020] Figure 4 A flowchart depicting an example method for remotely generating a chatbot at a remote system and enabling the chatbot to participate in a conversation with an entity according to various implementations.
[0021] Figure 5A and Figure 5B Depicts various non-limiting example interactions in which corresponding unstructured free-form natural language input is used to generate a corresponding chatbot, and the corresponding chatbot performs a corresponding task based on the corresponding unstructured free-form natural language input according to various implementations.
[0022] Figure 6A and Figure 6B depicts additional non - limiting example interactions in which corresponding unstructured free - form natural language inputs are used to generate corresponding chatbots, and in which the corresponding chatbots perform corresponding tasks based on the corresponding unstructured free - form natural language inputs, according to various implementations.
[0023] Figure 7 depicts an example architecture of a computing device according to various implementations. Detailed Description
[0024] Turning now to Figure 1 , a block diagram depicts an example environment that illustrates various aspects of the present disclosure and in which the implementations disclosed herein can be implemented. Client device 110 is shown in Figure 1 and, in various implementations, includes a user input engine 120, a rendering engine 130, an on - device machine learning (ML) model engine 140, and a chatbot development engine client 150. The client device 110 can be, for example, a stand - alone device (e.g., having a microphone, visual components, speakers, a display, and / or other user interface components), a laptop computer, a desktop computer, a tablet, a wearable computing device, an in - vehicle computing device, and / or any other client device capable of implementing the chatbot development engine client 150.
[0025] The user input engine 120 can detect various types of user input at the client device 110. In some examples, the user input detected at the client device 110 can include verbal input detected via a microphone of the client device 110. In these examples, the microphone of the client device 110 can generate audio data that captures the verbal utterances included in the verbal input. In other examples, the user input detected at the client device 110 can include touch input detected via a user interface input device (e.g., a touch - sensitive display) of the client device 110, and / or typed input detected via a user interface input device (e.g., a touch - sensitive display and / or a keyboard) of the client device 110. In these examples, the user interface input device of the client device 110 can generate text data that captures the touch input and / or the typed input. It is noted that the unstructured free - form natural language input described herein can be provided by a user of the client device 110 as any combination of verbal input, touch input, and / or typed input.
[0026] The rendering engine 130 may cause responsive content and / or other output to be visually rendered for presentation to a user at the client device 110 (e.g., via a touch-sensitive display or other user interface output device), and / or to be aurally rendered for presentation to a user at the client device 110 (e.g., via a speaker or other user interface output device). The responsive content and / or other output may include, for example, various types of user interfaces associated with the chatbot development engine client 150 that can be visually rendered via the user interface of the client device 110, such as unstructured free-form natural language input provided by a user of the client device 110 that conveys details of a task to be performed by the chatbot on behalf of the user of the client device, a transcription of a corresponding conversation performed by the chatbot on behalf of the user of the client device 110, various prompts associated with the corresponding conversation performed by the chatbot on behalf of the client device 110, the results and / or summary of the corresponding conversation performed by the chatbot on behalf of the client device 110, and / or any other responsive content or output that can be visually and / or aurally rendered for presentation to a user at the client device 110.
[0027] In various implementations, the on-device ML model engine 140 may include an automatic speech recognition (ASR) engine 141, a natural language understanding (NLU) engine 142, an execution engine 143, and a text-to-speech (TTS) engine 144. As described in more detail below, these on-device ML model engines of the on-device ML model engine 140 may utilize various on-device ML models (e.g., stored in the on-device ML model database 140A) to process various user inputs (e.g., received via the user input engine 120) and generate various outputs (e.g., to be visually and / or aurally rendered via the rendering engine 130 for presentation to a user). This, in turn, enables the chatbot development engine client 150 to use the on-device ML model engine 140 to process various user inputs received at the client device 110 and generate various outputs to be provided for presentation to a user at the client device 110.
[0028] Further, the client device 110 is shown in Figure 1 to be communicatively coupled to a remote system 160 via one or more networks 199 (e.g., any combination of Wi-Fi, Bluetooth, or other local area networks (LANs); Ethernet, the Internet, or other wide area networks (WANs); and / or other networks). In various implementations, the remote system 160 includes a remote system ML model engine 170 and a chatbot development engine 180. The remote system 160 may be, for example, a high-performance server, a cluster of high-performance servers, and / or any other computing device remote from the client device 110.
[0029] In various implementations, the remote ML model engine 170 may include an ASR engine 171, an NLU engine 172, a fulfillment engine 173, and a TTS engine 174. As described in more detail below, these remote ML model engines of the remote engine 170 may utilize various remote ML models (e.g., stored in the remote ML model database 170A) to process various user inputs (e.g., received via the user input engine 120) in the same or similar manner as the on-device ML model engine 140 and based on data provided to the remote system 160 by the client device 110 via one or more networks 199 and generate various outputs (e.g., to be visually and / or aurally rendered via the rendering engine 130 for presentation to the user). This, in turn, enables the chatbot development engine 180 to use the remote ML model engine 170 to process various user inputs received at the client device 110 and generate various outputs to be provided for presentation to the user at the client device 110. In implementations where the remote ML model engine 170 is used to process various user inputs received at the client device 110 and generate various outputs to be provided for presentation to the user at the client device 110, the various user inputs received at the client device 110 may be transmitted (e.g., via one or more networks 199) to the remote system 160, and the various user outputs may be transmitted (e.g., via one or more networks 199) back to the client device 110.
[0030] Notably, the chatbot development engine client 150 of the client device may communicate with the chatbot development engine 180 via one or more networks 199. From the perspective of a user interacting with the client device 110, the chatbot development engine client 150 and the chatbot development engine 180 form a logical instance of the chatbot development platform. Although the chatbot development platform is depicted in Figure 1 as being implemented in a distributed manner via one or more networks 199 (e.g., via the use of the chatbot development engine client 150 and the chatbot development engine 180), it should be understood that this is for illustrative purposes only and is not meant to be limiting. For example, the chatbot development platform may alternatively be implemented specifically at the client device 110. As another example, the chatbot development platform may alternatively be implemented specifically at the remote system 160, but the client device 110 is still used to enable the user to interact with the chatbot development platform.
[0031] A user (e.g., a user of the client device 110) can utilize a chatbot development platform to train a chatbot as described herein to be deployed to conduct corresponding conversations on behalf of the user and / or on behalf of a third party associated with the user (e.g., via the third-party system 192). It is noted that the chatbot development platform can be provided by a first party, and the user can utilize the chatbot development platform to generate a chatbot for himself or herself or for a third party associated with the user. As used herein, the term first party refers to the entity that publishes the chatbot development platform, and the term third party refers to an entity that is different from the entity associated with the first party and does not publish the chatbot development system. Thus, the user of the client device 110 who interacts with the chatbot development platform can also be referred to as a third-party developer.
[0032] The corresponding conversations described herein as being conducted by the chatbot and on behalf of the user of the client device 110 can include various types of conversations, such as voice-based conversations and text-based conversations. Voice-based conversations can include, for example, corresponding conversations during an automated telephone call (e.g., Voice over Internet Protocol (VoIP), Public Switched Telephone Network (PSTN), and / or other telephone communication protocols) and corresponding conversations between the client device and an additional client device 191, corresponding conversations that the chatbot and other entities and / or users participate in locally at a given client device (e.g., in a scenario where the client device 110 is a shared client accessible by multiple users), and / or corresponding conversations in any other voice-based scenario where the chatbot is deployed to conduct a corresponding conversation with the user. Text-based conversations can include, for example, corresponding conversations during text or SMS messaging, email, and / or in any other text-based scenario where the chatbot is deployed to conduct a corresponding conversation with the user.
[0033] As described above, the chatbot development platform can use the on-device ML model engine 140 and / or the remote system ML model engine 170 to process various user inputs received at the client device 110 and generate various outputs to be provided for presentation to the user at the client device 110. Each sub-engine of the on-device ML model engine 140 and / or the remote system ML model engine 170 can be configured to perform one or more functions. Notably, the remote system ML model engine 170 includes a remote-based counterpart of the sub-engines of the on-device ML model engine 140. In various implementations, the utilization of the on-device ML model engine 140 can be prioritized, at least in part due to latency considerations, network bandwidth privacy considerations, and / or other considerations. In these implementations, when one or more sub-engines of the on-device ML model engine 140 fail, the remote system ML model engine 170 can be utilized. In other implementations, the utilization of the remote ML model engine 170 can be prioritized, at least in part due to computational considerations at the client device 110, hardware considerations at the client device 110, software considerations at the client device 110, and / or other considerations. In yet other implementations, the on-device ML model engine 140 and the remote system ML model engine 170 can be utilized in combination with each other.
[0034] For example, the ASR engine 141 and / or 171 can use an ASR model (e.g., a recurrent neural network (RNN) model, a transformer model, and / or any other type of ML model capable of performing ASR) stored in the corresponding ML model database to process audio data captured from spoken utterances and generated by the microphone of the client device 110 to generate an ASR output. Further, the NLU engine 142 and / or 172 can use an NLU model (e.g., long short-term memory (LSTM), gated recurrent unit (GRU), and / or any other type of RNN or other ML model capable of performing NLU) and / or NLU rules stored in the corresponding ML model database to process the ASR output (or other typed input or touch input received via the user input engine 120 of the client device 110) to generate an NLU output. Additionally, the fulfillment engine 143 and / or 173 can use a fulfillment model and / or fulfillment rules stored in the corresponding ML model database to process the NLU data to generate a fulfillment output. Additionally, the TTS engine 144 and / or 174 can use a TTS model stored in the corresponding ML model database to process text data (e.g., text formulated by the chatbot) to generate synthesized voice audio data including computer-generated synthesized speech.
[0035] In various implementations, the ASR output may include, for example, multiple speech hypotheses for the verbal input (e.g., word hypotheses and / or transcription hypotheses) based on the processing of the audio data, and a particular speech hypothesis may be optionally selected as the recognized text for the verbal input based on a corresponding value (e.g., a probability value, a log-likelihood value, and / or other value) associated with each of the multiple speech hypotheses. In various implementations, the ASR model stored in the respective ML model database is an end-to-end speech recognition model such that the ASR engine 141 and / or 171 can directly use the model to generate the multiple speech hypotheses. For example, the ASR model can be an end-to-end model for generating each of the multiple speech hypotheses on a character-by-character basis (or other token-by-token basis). A non-limiting example of such an end-to-end model for generating the recognized text on a character-by-character basis is the Recurrent Neural Network Transducer (RNN-T) model. The RNN-T model is in the form of a sequence-to-sequence model that does not employ an attention mechanism. In other implementations, the ASR model is not an end-to-end speech recognition model such that the ASR engine 141 and / or 171 can alternatively generate the predicted phonemes (and / or other representations). For example, the predicted phonemes (and / or other representations) can then be used by the ASR engine 141 and / or 171 to determine the multiple speech hypotheses that match the predicted phonemes. In doing so, the ASR engine 141 and / or 171 can optionally employ a decoding graph, a dictionary, and / or other resources. In various implementations, the corresponding transcription can be rendered at the client device 110 (e.g., rendered in association with the training instance input, the training instance output, the corresponding feature emphasis input, the demonstrative talk, and / or other aspects of the chatbot development platform).
[0036] In various implementations, the NLU output can include, for example, the annotated recognized text, which includes one or more (e.g., all) annotations for one or more of the terms of the recognized text. For example, the NLU engines 142 and / or 172 can include a part-of-speech tagger (not depicted) configured to annotate terms using the syntactic roles of the terms. Additionally or alternatively, the NLU engines 142 and / or 172 can include an entity tagger (not depicted) configured to annotate entity references in one or more segments of the recognized text, such entity references being references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, locations (real and imaginary), etc. In some implementations, data about entities can be stored in one or more databases, such as stored in a knowledge graph (not depicted). In some implementations, the knowledge graph can include nodes representing known entities (and in some cases, entity attributes), and edges connecting the nodes and representing relationships between the entities. The entity tagger can annotate references to entities at a high granularity level (e.g., such that all references to an entity category, such as people, can be identified) and / or at a lower granularity level (e.g., such that all references to a specific entity, such as a specific person, can be identified). The entity tagger can rely on the content of the natural language input to resolve specific entities and / or can optionally communicate with a knowledge graph or other entity database to resolve specific entities. Additionally or alternatively, the NLU engines 142 and / or 172 can include a coreference resolver (not depicted) configured to group or "cluster" references to the same entity based on one or more contextual cues. For example, the coreference resolver can be used to resolve the term "them" in the input "buy them" to "buy theater tickets" based on the "theater tickets" mentioned in a client device notification rendered immediately before receiving the natural language input "buy them". In some implementations, one or more components of the NLU engines 142 and / or 172 can rely on annotations from one or more other components of the NLU engines 142 and / or 172. For example, in some implementations, the entity tagger can rely on annotations from the coreference resolver in annotating all mentions of a specific entity. Also, for example, in some implementations, the coreference resolver can rely on annotations from the entity tagger in clustering references to the same entity. Also, for example, in some implementations, the coreference resolver can rely on the user data of the user of the client device 110 (e.g., stored in the user data database 110A) in coreference resolution and / or entity resolution.User data can include, for example, historical location data, historical time data, user preference data, user account data, calendar information, email data, and / or any other user data accessible at the client device 110.
[0037] In various implementations, the fulfillment output can include, for example, one or more tasks to be performed by the chatbot and on behalf of the user of the client device 110. As described herein (e.g., with respect to Figure 2 , Figure 3 , Figure 4 , Figure 5A , Figure 5B , Figure 6A and Figure 6B ), the user can provide unstructured free-form natural language input that includes one or more tasks to be performed by the chatbot and on behalf of the user of the client device 110. The one or more tasks may require the chatbot to participate in a corresponding conversation with an entity (or a representative of the entity). It is noted that the unstructured free-form natural language input can convey details of the one or more tasks to be performed by the chatbot without any corresponding dialogue state diagram to be utilized by the chatbot in conducting the corresponding conversation. Nevertheless, and by leveraging the chatbot development engine client 150 and / or the chatbot development engine 180, a chatbot can be generated and deployed in response to receiving the unstructured free-form natural language input to perform the one or more tasks. Thus, it should be understood that the fulfillment output can be based on the one or more tasks to be performed by the chatbot and can be based on responsive content determined during a corresponding conversation with an entity (or a representative of the entity).
[0038] In various implementations, the TTS engine 144 and / or 174 can generate synthesized speech audio data that captures computer-generated synthesized speech. The synthesized speech audio data can be rendered at the client device 110 via the speakers of the client device 110, and / or at an additional client device via the corresponding speakers of the additional client device 191 (e.g., a client device associated with an entity and / or a representative of the entity). The synthesized speech can include any output generated by the chatbots described herein and can include, for example, synthesized speech generated as part of a conversation between the user of the client device 110 and the chatbot, synthesized speech generated as part of a conversation between an entity (or a representative of the entity) and the chatbot, and / or other synthesized speech.
[0039] Although Figure 1Descriptions are made with a single user for a single client device, but it should be understood that this is for illustrative purposes and is not meant to be restrictive. For example, one or more additional client devices of the user may also implement the techniques described herein. For example, client device 110, one or more additional client devices, and / or any other computing device of the user may form a device ecosystem that can employ the techniques described herein. These additional client devices and / or computing devices may communicate (e.g., via one or more networks 199) with client device 110 and / or remote system 160. As another example, a given client device may be utilized by multiple users in a shared setting (e.g., a group of users, a household, etc.).
[0040] In various implementations, the chatbot development engine 180 may include a chatbot identification engine 181, a chatbot fine-tuning engine 182, a task / entity identification engine 183, a conversation engine 184, a conversation context engine 185, a responsive content engine 186, and a peripheral behavior engine 187, as Figure 1 depicted. Although the chatbot development engine 180 is depicted as having specific sub-engines, it should be understood that this is for illustrative purposes and is not meant to be restrictive. For example, one or more of the sub-engines depicted Figure 1 may be combined, while one or more of the other sub-engines depicted Figure 1 may be omitted. Further, although the chatbot development engine client 150 is not depicted as including any sub-engines, it should be understood that this is for the sake of brevity and is not meant to be restrictive. For example, the chatbot development engine client 150 may include the same sub-engines or a subset thereof as described for the chatbot development engine 180. Additional descriptions of the chatbot development engine 180 and its various sub-engines are provided Figure 2 herein.
[0041] Now referring Figure 2 , an example process flow 200 for generating a chatbot and enabling the chatbot to participate in a conversation with an entity is depicted. For illustrative purposes, it is assumed that a user of client device 110 from Figure 1 provides unstructured free-form natural language input 201 as input at client device 110. Client device 110 may receive the unstructured free-form natural language input 201 via the user input engine 120 of client device 110. Client device 110 may cause the unstructured free-form natural language input 201 to be processed using various sub-engines of the on-device ML model engine 140 and / or using various sub-engines of the remote ML model engine 170.
[0042] Note that in processing the unstructured free-form natural language input 201, the client device 110 may identify one or more features 202 based on the output generated by one or more of the various sub-engines of the on-device ML model engine 140 and / or one or more of the various sub-engines of the remote ML model engine 170. The one or more features 202 may include, for example, an ASR output in the case where the unstructured free-form natural language input 201 is an oral input, an NLU output in the case where the unstructured free-form natural language input 201 is an oral input or a typed input, and / or a fulfillment output in the case where the unstructured free-form natural language input 201 is an oral input or a typed input, such as entities, intents, slot values, tasks associated with the entities to be performed by the chatbot, and / or other features.
[0043] For example purposes, further assume that the unstructured free-form natural language input 201 is an oral utterance of "Call Example Hotel and ask if there is a pet fee for small dogs". In this example, the ASR output may be the recognized text corresponding to the oral utterance (e.g., the recognized text of "call Example Hotel and ask if there is a pet fee for small dogs"), and the NLU output may be a "call" intent associated with a task of submitting a query of "is there is a pet fee for small dogs?" to a representative of "Example Hotel" with a slot value of "Example Hotel". Thus, in some of these examples, and in response to receiving the unstructured free-form natural language input 201, the client device 110 may generate a chatbot that is configured to call the phone number associated with "Example Hotel" during a training phase (e.g., enclosed by box 280A in Figure 2 to submit a query of "is there is a pet fee for small dogs" to a representative of "Example Hotel".
[0044] During a training phase, the chatbot identification engine 181 may identify a chatbot 203 (e.g., stored in the chatbot database 180A). The chatbot 203 may be a previously trained ML model or a combination of various previously trained ML models that can be fine-tuned based on an unstructured free-form natural language input 201 and / or one or more features 202 extracted from the unstructured free-form natural language input 201. For example, the chatbot 203 may correspond to a previously trained large language model (LLM) such as LaMDA, BERT, Meena, GPT-3, and / or another previously trained LLM. It is noted that these previously trained LLMs have been previously trained on a large amount of different data and are generally generative ML models capable of participating in a corresponding conversation with a user in a more natural and intuitive manner. These LLMs have multiple ML layers and hundreds of millions to hundreds of billions of ML parameters and are capable of generalizing corresponding conversations with users. For example, and as described in more detail herein (e.g., with respect to Figure 5A , Figure 5B and Figure 6B ), text data may be provided as an input across these previously trained LLMs to generate an LLM output, such as a probability distribution over a vocabulary, and a response to the text data may be generated based on the probability distribution over the vocabulary. Due to the multiple ML layers and hundreds of millions to hundreds of billions of ML parameters, it should be noted that LLMs are generally not conducive to local implementation at the client device 110, such as when the chatbot database 180A is local to the client device 110 (e.g., stored in the device storage of the client device 110). Nevertheless, various sparsification techniques can be utilized to reduce the amount of ML layers and / or the amount of ML parameters utilized by these LLMs such that a sparsified version of the previously trained LLM can be locally implemented at the client device 110 while mitigating a reduction in the accuracy and / or recall rate of the previously trained LLM due to sparsification. These sparsification techniques may include, but are not limited to, folding and / or combining multiple layers among the multiple ML layers of the previously trained LLM, pruning multiple layers among the multiple ML layers of the previously trained LLM, masking the weights of the previously trained LLM, pruning the weights of the previously trained LLM, and / or other sparsification techniques. However, when the chatbot database 180A is remote from the client device 110 (e.g., stored in a remote storage of the remote system 160), an unsparsified version of the previously trained LLM can be remotely implemented at the remote system 160. Thus, the chatbot 203 may be identified locally at the client device 110 and / or remotely at the remote system 160 (e.g., remote from the client device 110 that received the unstructured free-form natural language input 201).
[0045] Further, during the training phase, the chatbot fine-tuning engine 182 can utilize various fine-tuning techniques to generate a fine-tuned chatbot 204 by fine-tuning the chatbot 203 and based on the unstructured free-form natural language input 201 and / or one or more features 202 extracted from the unstructured free-form natural language input 201 (and the fine-tuned chatbot 204 can optionally be stored in the chatbot database 180A). These fine-tuning techniques can include, but are not limited to, instruction tuning, few-shot learning, and / or other fine-tuning techniques, and the fine-tuning performed can vary based on the unstructured free-form natural language input 201 provided by the user. In other words, the previously trained LLM corresponding to the chatbot 203 can be further trained based on the unstructured free-form natural language input 201 and / or one or more features 202 extracted from the unstructured free-form natural language input 201 such that the previously trained LLM that is fine-tuned and corresponds to the fine-tuned chatbot 204 is suitable for performing tasks on behalf of the user. By fine-tuning the chatbot 203, the resulting fine-tuned chatbot 204 leverages the generalization ability of the previously trained LLM while also being suitable for performing tasks associated with the entity on behalf of the user. Thus, the chatbot 203 can be fine-tuned to locally generate the fine-tuned chatbot 204 at the client device 110 and / or remotely generate the fine-tuned chatbot 204 at a remote system 160 (e.g., away from the client device 110 that received the unstructured free-form natural language input 201). The fine-tuned chatbot 204 can then be utilized during the inference phase (e.g., enclosed by the box 280B in Figure 2 ).
[0046] Although Figure 2 is described with respect to fine-tuning the chatbot 203 based on the unstructured free-form natural language input 201 and / or one or more features 202 of the unstructured free-form natural language input 201 to generate the fine-tuned chatbot 204, it should be understood that this is just one implementation considered herein. For example, in other implementations, the chatbot 203 may not be fine-tuned such that the chatbot 203 can then be utilized by the client device 110 and / or the remote system 160 during the inference phase (e.g., enclosed by the box 280B in Figure 2 ).
[0047] During the inference phase, the task / entity identification engine 183 can determine task data 205 and entity data 206 to be used by the chatbot to perform tasks during a corresponding conversation. Continuing the above example of an oral utterance where the unstructured free-form natural language input 201 is "Call Example Hotel and ask if there is a pet fee for small dogs", the task data 205 can include, for example, the task of submitting a query of "is there a pet fee for small dogs" to a representative of "Example Hotel". Further, the entity data 206 can include a corresponding identifier for "Example Hotel", such as a phone number for initiating a call to a representative of "Example Hotel". In various implementations, information about tasks and / or entities can be stored in the task / entity database 180B and / or other data sources accessible by the client device 110. Although the task data 205 and the entity data 206 are described as including specific data, it should be understood that this is for illustrative purposes and is not meant to be restrictive. For example, the task data 205 can include any data related to any task that can be specified by a user of the client device 110 in the unstructured free-form natural language input 201. Further, the entity data 106 can include an indication of any entity and / or any corresponding identifier for the entity. Additionally, although the task data 205 in the above example includes only a single task and the entity data 206 identifies only a single entity, it should be understood that this is for illustrative purposes and is not meant to be restrictive. For example, the task data 205 can include any data for multiple tasks that can be specified by a user of the client device 110 in the unstructured free-form natural language input 201. Further, the entity data 206 can identify multiple entities belonging to a particular type of entity (e.g., if the unstructured free-form natural language input 201 instead corresponds to the oral utterance "Call hotels around me and ask if there is a pet fee for small dogs", where "hotels" is a particular type of entity and can identify specific hotels located near the user).
[0048] Note that the task / entity identification engine 183 can utilize the entity data 206 to initiate a corresponding conversation with a representative of the entity (e.g., by calling the phone number associated with "Example Hotel" to make an outbound call to a given additional client device 191A via the chatbot 203 or the fine-tuned chatbot 204), and can provide the task data 205 to the conversation engine 184 to enable the chatbot 203 or the fine-tuned chatbot 204 to participate in the corresponding conversation with the representative of the entity. In various implementations, the conversation context engine 185 can provide the conversation context data 207 to the conversation engine 184 and, in addition, enable the chatbot 203 or the fine-tuned chatbot 204 to participate in a more contextually relevant conversation with the representative of the entity task data 205. In these implementations, the conversation context data 207 can represent (e.g., as a vector or other data structure) the initial context information for the corresponding conversation or the subsequent context information determined during the corresponding conversation (e.g., determined based on the data stored in the chatbot activity database 180C).
[0049] Continuing the above example where the unstructured free-form natural language input 201 is the spoken utterance "Call Example Hotel and ask if there is a pet fee for small dogs", the conversation context engine 185 can generate conversation context data 207 indicating that "[user] intends on staying at Example Hotel", information associated with the user's intention to stay at the client device 110 (e.g., check-in or check-out date / time from the user's email account at the client device 110, "Example Hotel" loyalty reward number from the user's email account at the client device 110, or "Example Hotel" software application accessible at the client device 110), "[user] would like to bring his or her small dog", and / or other context information that can be inferred based on the unstructured free-form natural language input 201 and / or other data accessible at the client device 110 (e.g., via the user data database 110A).
[0050] Further, during the inference phase and after initiating a corresponding conversation (e.g., using entity data 206), and in an implementation where the chatbot corresponds to the fine-tuned chatbot 204, the conversation engine 184 may initially use the fine-tuned chatbot 204 to process task data 205 (and optionally conversation context data 207) to generate an output, such as a probability distribution over a sequence of words or phrases. The conversation engine 184 may generate conversation data 208 based on the output generated using the fine-tuned chatbot 204. The conversation data 208 may include, for example, one or more instances of synthesized speech audio data in an implementation where the corresponding conversation is a voice-based conversation, and one or more instances of text data in an implementation where the corresponding conversation is a text-based conversation. In various implementations, and as Figure 2 depicted, the conversation data 208 may be transmitted to a given additional client device 191A such that the conversation data 208 is rendered audibly and / or visually at the given additional client device 191A. However, in other implementations, such as when the fine-tuned chatbot 204 participates in the corresponding conversation locally at the client device 110 (e.g., when the client device 110 is deployed in a shared setting), the conversation data 208 may be rendered audibly and / or visually at the client device 110.
[0051] Alternatively, during the inference phase and in an implementation where the chatbot corresponds to chatbot 203 (e.g., rather than the fine-tuned chatbot 204), the conversation engine 184 may initially use chatbot 203 to process the unstructured free-form natural language input 201, one or more features 202 determined based on processing the unstructured free-form natural language input 201, task data 205 (and optionally any other data described herein) to generate an output, such as a probability distribution over a sequence of words or phrases. The conversation engine 184 may generate conversation data 208 based on the output generated using chatbot 203. The conversation data 208 may include, for example, an instance of synthesized speech audio data in an implementation where the corresponding conversation is a voice-based conversation, and an instance of text data in an implementation where the corresponding conversation is a text-based conversation. In various implementations, and as Figure 2As depicted, the conversation data 208 can be transmitted to a given additional client device 191A such that the conversation data 208 is rendered audibly and / or visually at the given additional client device 191A. However, in other implementations, such as when the chatbot 203 participates in the corresponding conversation locally at the client device 110 (e.g., when the client device 110 is deployed in a shared setting), the conversation data 208 can be rendered audibly and / or visually at the client device 110. In other words, rather than fine-tuning the chatbot 203 during the training phase, the chatbot 203 can be activated during the inference phase based on the unstructured free-form natural language input 201 and / or one or more features 202 determined based on processing the unstructured free-form natural language input 201. This enables the client device 110 and / or the remote system 160 to save computing resources while still effectively deploying the chatbot to participate in the corresponding conversation.
[0052] Continuing the above example where the unstructured free-form natural language input 201 is the spoken utterance "Call Example Hotel and ask if there is a pet fee for small dogs", the conversation data 208 can include the synthesized speech audio data that is to be rendered audibly at a given additional client device 191A and includes the synthesized speech "Hi, this is a chatbot calling on behalf of [user], is there is a pet fee for small dogs?". Notably, the synthesized speech includes the query "is there a pet fee for small dogs?" and solicits a response from a representative of the entity answering the phone call initiated by the chatbot 203 or the fine-tuned chatbot 204. Thus, the response data 209 can include a response to the query included in the synthesized speech. The response data 209 can include audio data provided by the representative of the entity. In these implementations, the responsive content engine 186 can utilize the ML model engines 140 and / or 170 to process using various ML models to determine whether the response data 209 includes a response indicating that the task has been successfully executed.
[0053] For example, further assume that the response data 209 includes audio data capturing an oral statement from a representative of the entity that says "yes, our pet fee for small dogs is $25 per night". In this case, the responsive content engine 186 can cause the audio data provided by the representative of the entity to be processed (e.g., using input parsing of an ASR model, an NLU model, and / or fulfillment rules) to determine that the representative of the entity provided a response in response to the query. In other words, the responsive content engine 186 can determine the responsive content 210 provided by the representative of the entity and in response to the query corresponding to the task in this example (e.g., the pet fee for small dogs is $25 per night). Further, the responsive content engine 186 can provide the responsive content 210 to the rendering engine 130 so that the client device can audibly and / or visually provide the rendered responsive content 211 for presentation to the user. Thus, the implementations described herein enable a user to provide unstructured free-form natural language input 201 so that a fine-tuned bot 204 is generated and used to perform the tasks included in the unstructured free-form natural language input 201.
[0054] As described herein (e.g., with respect to Figure 5A , Figure 5B and Figure 6B ), the chatbot 203 and the fine-tuned chatbot 204 can have various peripheral behaviors that can be implemented by the chatbot 203 and the fine-tuned chatbot 204 by leveraging the peripheral behavior engine 187. These peripheral behaviors can include, but are not limited to: a greeting behavior that enables the chatbot 203 and the fine-tuned chatbot 204 to identify the user of the client device 110 and / or identify itself as a chatbot, a remote procedure call (RPC) behavior that enables the chatbot 203 and the fine-tuned chatbot 204 to search one or more databases during a corresponding conversation, a hold behavior that enables the chatbot 203 and the fine-tuned chatbot 204 to pause and resume a corresponding conversation, an exit behavior that enables the chatbot 203 and the fine-tuned chatbot 204 to prompt the user of the client device 110 to join a corresponding conversation and / or otherwise terminate a corresponding conversation when requested by a representative of the entity, a clarification behavior that enables the chatbot 203 and the fine-tuned chatbot 204 to clarify and / or repeat information previously provided during a corresponding conversation, and / or those other peripheral behaviors that can be invoked by the chatbot 203 and the fine-tuned chatbot 204 under corresponding conditions for invoking other peripheral behaviors.
[0055] Although Figure 2Described is a telephone call between a chatbot 203 or a fine-tuned chatbot 204 (e.g., implemented locally at the client device 110 and / or remotely at the remote system 160) and a representative of an entity (e.g., accessible at a given additional client device 191A), but it should be understood that this is not meant to be limiting. Instead, it should be understood that the techniques described herein can be used to fine-tune a chatbot that can be deployed to participate in voice-based conversations and text-based conversations across multiple computing devices and / or at a single computing device.
[0056] Now turning to Figure 3 , a flowchart of an example method 300 is depicted that shows generating a chatbot locally at a client device and having the chatbot participate in a conversation with an entity. For convenience, the operations of method 300 are described with reference to the system performing the operations. This system of method 300 includes at least one processor, memory, and / or other components of a client device (e.g., Figure 1 the client device 110, Figure 7 the computing device 710, and / or other client devices). Further, although the operations of method 300 are shown in a specific order, this is not intended to be limiting. One or more operations can be reordered, omitted, and / or added.
[0057] At block 352, the system receives unstructured free-form natural language input from a user of the client device, the unstructured free-form natural language input including one or more tasks associated with an entity. The unstructured free-form natural language input received from the user of the client device can include, for example, oral input received via a microphone of the client device, typed input received via a touch-sensitive display of the client device, and / or touch input received via a touch-sensitive display of the client device. Further, the unstructured free-form natural language input can convey details of one or more tasks to be performed by the chatbot and on behalf of the user of the client device. It is noted that the unstructured free-form natural language input is unstructured, meaning that the user does not need to provide the free-form natural language input according to any pattern or specific manner.
[0058] At block 354, the system locally generates, at the client device, a chatbot for performing one or more tasks associated with an entity on behalf of a user, at least based on an unstructured free-form natural language input. In some implementations, and as indicated at block 354A, the system obtains a previously trained large language model (LLM) locally stored at the client device. Further, in these implementations, and as indicated at block 354B, the system causes the previously trained LLM locally stored at the client device to be fine-tuned based on the unstructured free-form natural language input to generate a fine-tuned LLM. Additionally, in these implementations, and as indicated at block 354C, the system utilizes the fine-tuned LLM as the chatbot. The system can generate a chatbot for performing one or more tasks associated with an entity on behalf of a user in the same or similar manner as described above with respect to Figure 2 (e.g., in an implementation where the training phase is locally implemented at the client device and described with respect to block 280A). Notably, in these implementations, the system is being locally implemented at the client device, and as a result, due to various hardware and / or software constraints of the client device, the previously trained LLM may be a sparser version of the previously trained LLM that may otherwise be available (e.g., for a remote system).
[0059] At block 356, the system causes the chatbot to perform one or more tasks associated with an entity on behalf of a user. In some implementations, and as indicated at block 356A, the system causes the chatbot to participate in a corresponding conversation with the entity by: rendering multiple instances of the synthesized speech audio data for presentation to a representative associated with the entity, and / or rendering multiple instances of text data for presentation to the representative of the entity. Further, in these implementations, and as indicated at block 356B, the system determines responsive content in response to one or more instances of the synthesized speech audio data and / or one or more instances of the text data. The system can cause the chatbot to perform one or more tasks associated with an entity on behalf of a user in the same or similar manner as described above with respect to Figure 2 (e.g., in an implementation where the inference phase is locally implemented at the client device and described with respect to block 280B) and with respect to Figure 5A 、Figure 5C and Figure 6B described.
[0060] At block 358, the system determines whether the chatbot has successfully performed one or more tasks associated with the entity. The system can determine whether the chatbot has successfully performed one or more tasks associated with the entity based on, for example, the responsive content in response to one or more instances of the synthesized speech audio data and / or one or more instances of the text data. Herein (e.g., with respect toFigure 5A , Figure 5B and FIG. 5C) describe various non - limiting examples of determining whether a chatbot has successfully performed one or more tasks associated with an entity.
[0061] If, at an iteration of block 358, the system determines that the chatbot has successfully performed one or more tasks associated with an entity, the system can proceed to block 360. At block 360, the system causes responsive content to be provided for presentation to a user of the client device. The responsive content can be determined based on, for example, one or more responses provided by a representative of the entity during the corresponding conversation. The system proceeds to block 364. Block 364 is described in more detail below.
[0062] If, at an iteration of block 358, the system determines that the chatbot has not successfully performed one or more tasks associated with an entity, the system can proceed to block 362. At block 362, the system prompts the user to join the corresponding conversation with a representative of the entity. In other words, if the chatbot has not successfully performed one or more tasks associated with an entity, the chatbot can prompt the user to join the corresponding conversation to ensure that one or more tasks are still performed. The system proceeds to block 364.
[0063] At block 364, the system determines whether to discard the chatbot. The system can determine whether to discard the chatbot based on, for example, whether the chatbot has successfully performed one or more tasks associated with an entity, whether the chatbot is to be used to perform one or more tasks associated with an additional entity, whether a threshold duration has elapsed since the chatbot was generated, whether a threshold amount stored on the client device has been consumed by the chatbot, whether the threshold amount of the client device's storage is available while the chatbot is stored in the device's storage, and / or based on other conditions. In other words, in determining whether to discard the chatbot, the system can balance the performance of the chatbot and how the chatbot affects the client device.
[0064] If, at an iteration of block 364, the system determines not to discard the chatbot, the system can proceed to block 366. At block 366, the system continues to utilize the chatbot. For example, if one or more tasks associated with an entity are not successfully executed, but the one or more tasks can be successfully executed with respect to an additional entity, the system can continue to utilize the chatbot. For example, if the unstructured free-form natural language input corresponds to the spoken utterance "find me a local plumber that is available as soon as possible", but the responsiveness content determined based on calling the first local plumber indicates that the first local plumber is not available, the system can continue to utilize the chatbot to call a second local plumber. In such a case, the system can call the first local plumber and the second local plumber in a serial or parallel manner.
[0065] If, at an iteration of block 364, the system determines to discard the chatbot, the system can proceed to block 368. At block 368, the system discards the chatbot. For example, the system can discard the chatbot in response to determining that one or more tasks associated with an entity are successfully executed and, optionally, in response to determining that there are no additional entities with which to engage in a corresponding conversation. As another example, if the chatbot consumes more than a threshold amount of memory resources at the client device, or if the chatbot makes less than a threshold amount of memory resources available at the client device, the system can discard the chatbot.
[0066] The system can perform another iteration of method 300 based on additional unstructured free-form natural language input received by the system. Although Figure 3 is described with respect to the system being locally implemented at the user's client device, it should be understood that this is for purposes of example and is not meant to be limiting. For example, and as described below with respect to Figure 4 the system can be implemented by a remote system that is remote from the user's client device that provided the unstructured free-form natural language input.
[0067] Further, although Figure 3Method 300 is described with respect to a previously trained LLM (e.g., at block 354) that is fine-tuned as a chatbot based on unstructured free-form natural language input, but it should be understood that this is for illustrative purposes and is not meant to be limiting. In additional or alternative implementations, the previously trained LLM can be used as a chatbot without any fine-tuning. In these implementations, and in causing the chatbot to participate in a given corresponding conversation with a given additional user (e.g., at block 356), the system can initiate the previously trained LLM based on unstructured free-form natural language input. This enables the chatbot to participate in the corresponding conversation without any explicit fine-tuning during the training phase.
[0068] Now turning to Figure 4 , a flowchart of an example method 400 is depicted that shows remotely generating a chatbot at a remote system and causing the chatbot to participate in a conversation with an entity. For convenience, the operations of method 400 are described with reference to the system performing the operations. This system of method 400 includes at least one processor, memory, and / or other components of a remote system (e.g., Figure 1 remote system 160 of Figure 7 , computing device 710 of
[0069] and / or other remote systems). Further, although the operations of method 400 are shown in a particular order, this is not meant to be limiting. One or more operations can be reordered, omitted, and / or added.
[0070] At block 454, the system remotely generates, at a remote system, a chatbot for performing one or more tasks associated with an entity on behalf of a user, at least based on indications of unstructured free-form natural language input. In some implementations, and as indicated at block 454A, the system obtains a previously trained large language model (LLM) remotely stored at the remote system. Further, in these implementations, and as indicated at block 454B, the system causes the previously trained LLM remotely stored at the remote system to be fine-tuned based on the unstructured free-form natural language input to generate a fine-tuned LLM. Additionally, in these implementations, and as indicated at block 454C, the system utilizes the fine-tuned LLM as the chatbot. The system can generate a chatbot for performing one or more tasks associated with the entity on behalf of the user in the same or a similar manner as described above with respect to Figure 2 (e.g., in an implementation where the training phase is remotely implemented at the remote system and described with respect to block 280A). Notably, in these implementations, the system is implemented away from the client device and at the remote system, and as a result, due to few hardware and / or software constraints at the remote system, the previously trained LLM can be an unsparsified version of the previously trained LLM that is more robust than a sparsified version of the previously trained LLM.
[0071] At block 456, the system causes the chatbot to perform one or more tasks associated with the entity on behalf of the user. In some implementations, and as indicated at block 456A, the system causes the chatbot to participate in a corresponding conversation with the entity by: rendering multiple instances of the synthesized speech audio data for presentation to a representative associated with the entity, and / or rendering multiple instances of the text data for presentation to the representative of the entity. Further, in these implementations, and as indicated at block 456B, the system determines responsive content in response to one or more of the instances of the synthesized speech audio data and / or one or more of the instances of the text data. The system can cause the chatbot to perform one or more tasks associated with the entity on behalf of the user in the same or a similar manner as described above with respect to Figure 2 (e.g., in an implementation where the inference phase is remotely implemented at the remote system and described with respect to block 280B) and with respect to Figure 5A Figure 5C and Figure 6B described. Notably, in these implementations, and in participating in the corresponding conversation, in some cases, the system can communicate directly with an additional client device of the representative of the entity without interacting with the client device during the corresponding conversation.
[0072] At block 458, the system determines whether the chatbot has successfully executed one or more tasks associated with an entity. The system can determine whether the chatbot has successfully executed one or more tasks associated with an entity based on, for example, responsive content in one or more instances of the synthesized speech audio data and / or one or more instances of the text data. Various non-limiting examples of determining whether the chatbot has successfully executed one or more tasks associated with an entity are described herein (e.g., with respect to Figure 5A , Figure 5B and FIG. 5C).
[0073] If, at an iteration of block 458, the system determines that the chatbot has successfully executed one or more tasks associated with an entity, the system can proceed to block 460. The responsive content can be determined based on, for example, one or more responses provided by a representative of the entity during the corresponding conversation. At block 460, the system transmits the responsive content to the user's client device. Transmitting the responsive content to the user's client device causes the responsive content to be provided for presentation to the user of the client device. The system proceeds to block 464. Block 464 is described in more detail below.
[0074] If, at an iteration of block 458, the system determines that the chatbot has not successfully executed one or more tasks associated with an entity, the system can proceed to block 462. In other words, if the chatbot has not successfully executed one or more tasks associated with an entity, the chatbot can prompt the user to join the corresponding conversation to ensure that one or more tasks are still executed. At block 462, the system transmits a prompt for the user to join the corresponding conversation to the user's client device. Transmitting the prompt to the user's client device causes the prompt to be provided for presentation to the user of the client device. The system proceeds to block 464.
[0075] At block 464, the system determines whether to discard the chatbot. The system can determine whether to discard the chatbot based on, for example, whether the chatbot has successfully executed one or more tasks associated with an entity, whether the chatbot is to be used to execute one or more tasks associated with an additional entity, whether a threshold duration has elapsed since the chatbot was generated, and / or based on other conditions. It is noted that in these implementations, the system may not have to balance the performance of the chatbot and how the chatbot affects the client device when determining whether to discard the chatbot because the chatbot may not be locally generated and / or implemented at the user's client device.
[0076] If, at an iteration of block 464, the system determines not to discard the chatbot, the system can proceed to block 466. At block 466, the system continues to utilize the chatbot. For example, if one or more tasks associated with an entity were not successfully executed, but the one or more tasks can be successfully executed with respect to an additional entity, the system can continue to utilize the chatbot (e.g., as described in block 364 with respect to Figure 3 ). If, at an iteration of block 464, the system determines to discard the chatbot, the system can proceed to block 468. At block 468, the system discards the chatbot. For example, the system can discard the chatbot in response to determining that one or more tasks associated with the entity were successfully executed and optionally in response to determining that there are no additional entities with which to engage in a corresponding conversation.
[0077] The system can perform another iteration of method 400 based on an additional indication of additional unstructured free-form natural language input received by the system. Although Figure 4 is described with respect to the system being remote from the user's client device and being implemented at a remote system, it should be understood that this is for illustrative purposes and is not meant to be limiting. For example, and as described above with respect to Figure 3 , the system can be implemented at the client device of the user providing the unstructured free-form natural language input. As another example, the system can be implemented in a distributed manner at both the client device and the remote system. For example, the chatbot can be locally generated at the client device but implemented at the remote system such that the computational resources of the client device are not used to have the chatbot engage in a corresponding conversation. Also, for example, the chatbot can be generated at the remote system but implemented at the client device such that the computational resources of the client device are only used to have the chatbot engage in a corresponding conversation.
[0078] Further, although Figure 4 method 400 of
[0079] is also described with respect to the chatbot being a previously trained LLM that is fine-tuned based on unstructured free-form natural language input (e.g., at block 454), it should be understood that this is for illustrative purposes and is not meant to be limiting. In additional or alternative implementations, the previously trained LLM can be used as the chatbot without any fine-tuning. In these implementations, and in having the chatbot engage in a given corresponding conversation with a given additional user (e.g., at block 456), the system can initiate the previously trained LLM based on the unstructured free-form natural language input. This enables the chatbot to engage in a corresponding conversation without any explicit fine-tuning during the training phase. Figure 5A and Figure 5B, depicts various non - limiting example interactions in which corresponding unstructured free - form natural language inputs are used to generate corresponding chatbots, and the corresponding chatbots perform corresponding tasks based on the corresponding unstructured free - form natural language inputs. It should be noted that the interactions 500A and 500B described with respect to Figure 5A and Figure 5B can be implemented across multiple computing devices so that the chatbot performs the corresponding task to be executed. For example, the corresponding unstructured free - form natural language inputs described in the examples with respect to Figure 5A and Figure 5B can be received at the user's client device (e.g., the client device 110 of Figure 1 ), the chatbots described in the examples with respect to Figure 5A and Figure 5B can be generated at the user's client device (e.g., the client device 110 of Figure 1 ) and / or at a remote system (e.g., the remote system 160 from Figure 1 ), the chatbots described in the examples with respect to Figure 5A and Figure 5B can be implemented at the user's client device (e.g., the client device 110 of Figure 1 ) and / or at a remote system (e.g., the remote system 160 from Figure 1 ) and communicate with a representative of an entity via additional computing devices acting as representatives. Each of these computing devices can include corresponding components such as user - interface input components (e.g., microphones, vision components, presence sensors, touch - sensitive displays, keyboards, hardware buttons, software buttons, etc.), user - interface output components (e.g., touch - sensitive displays, speakers, monitors, projectors, etc.), network interfaces, and / or other components. Thus, although the interactions 500A and 500B of Figure 5A and Figure 5B are respectively depicted within a single interface, it should be understood that this is for the purpose of illustrating the various techniques described herein and is not meant to be limiting.
[0080] Specifically referring to Figure 5A, assuming that an on-device conversation with a user of a client device (e.g., Jane Doe) is initiated as indicated by interaction 500A1. In various implementations, the on-device conversation with Jane Doe can be initiated as part of a conversation between Jane Doe and an automated assistant that is executed at least in part at the client device. In these implementations, interaction 500A1 can be a voice-based interaction or a text-based interaction. For example, Jane Doe can invoke an automated assistant that is executed at least in part at the client device (e.g., by actuation of a software or hardware button, by saying a particular term or phrase such as "Assistant," "Hey Assistant," and / or by other means), and provide unstructured free-form natural language input 552A1 as spoken input. As another example, Jane Doe can access an automated assistant application that is accessible at the client device and associated with the automated assistant, and provide unstructured free-form natural language input 552A1 as typed input.
[0081] for Figure 5A For purposes of the example in , assume that Jane Doe provides an unstructured free-form natural language input 552A1 of "Ask the restaurant I ate dinner at yesterday if they found my red leather jacket" as verbal input. In this example, the automated assistant can use an ASR model to process the audio data that captures the verbal input to generate an ASR output, such as recognized text corresponding to the verbal input (e.g., the recognized text of "Ask the restaurant I ate dinner at yesterday if they found my red leather jacket"). Further, the automated assistant can use an NLU model to process the ASR output to generate NLU output, such as an intent, slot values of parameters associated with the intent, and / or other NLU outputs. It is noteworthy that the verbal input explicitly includes an intent of [submitting a query] (etc.) with a slot value of "did you find [Jane Doe's] redleather jacket" (etc.) for a [query content] parameter (e.g., task data). However, spoken input only implicitly identifies entities (e.g., “the restaurant I ate dinner at yesterday”).
[0082] Thus, in this example, the automated assistant can utilize data from various sources to parse entities implicitly included in the verbal input. (e.g., user data and / or other information of Jane Doe) For example, the automated assistant can utilize Jane Doe's historical location information to identify the restaurant where she ate dinner yesterday, utilize calendar information including a reservation to identify the restaurant where she ate dinner yesterday, utilize software application information including a reservation to identify the restaurant where she ate dinner yesterday, utilize email information including a confirmation email to identify the restaurant where she ate dinner yesterday, and / or utilize other information. As a result, further assume that "the restaurant I ate dinner at yesterday" corresponds to the restaurant entity of "Hypothetical Café" (e.g., entity data). Thus, the automated assistant can provide a response 554A1 of "Okay, calling Hypothetical Café, I’ll let you know if I need anything or if they found it" to be presented audibly and / or visually to Jane Doe to indicate that the automated assistant will perform the task of submitting a query to the entity of "Hypothetical Café".
[0083] In this example, and based on the unstructured free-form natural language input 552A1, the automated assistant can cause the client device and / or the remote system to generate a chatbot to perform the task of submitting a query to a representative of "Hypothetical Café". In some implementations, the chatbot can correspond to a previously trained LLM that is fine-tuned based on the unstructured free-form natural language input 552A1 using various fine-tuning techniques (e.g., as described with respect to Figure 2 ). In other implementations, the chatbot can correspond to a previously trained LLM that is not fine-tuned based on the unstructured free-form natural language input 552A1 but is initiated based on one or more of the multiple unstructured free-form natural language inputs. Further assume that the chatbot will participate in a corresponding conversation with a representative of "Hypothetical Café" as part of a phone call to perform the task of submitting a query to a representative of "Hypothetical Café". In this example, the automated assistant can determine the corresponding identifier associated with "Hypothetical Café", such as a phone number (e.g., entity data) that can be used to initiate a corresponding conversation with a representative of "Hypothetical Café".
[0084] Thus, and as shown in interaction 500A2, the automated assistant can enable a chatbot to be implemented at a computing device, such as a client device in an implementation where the chatbot is locally generated at the client device, or a remote system in an implementation where the chatbot is generated away from the client device. This enables the chatbot to make a phone call to a phone number associated with "Hypothetical Café" to facilitate the task of submitting a query to a representative of "Hypothetical Café". Further assume that after making the phone call, the chatbot and the representative of "Hypothetical Café" engage in a corresponding conversation. For example, further assume that the representative of "Hypothetical Café" answers the incoming phone call and provides the verbal input 552A2 of "Hello, this is John Smith at Hypothetical Café, how may I help you?". In this example, the automated assistant can enable the chatbot to process the intent of [submitting a query] (etc.) with a slot value of "did you find [Jane Doe’s] red leather jacket" for the [query content] parameter (e.g., task data), capture the audio data of the verbal input 552A2, the text data corresponding to the captured audio data of the verbal input 552A2 (e.g., determined using an ASR model), and / or any contextual conversation data to generate an instance of the synthesized voice audio data. The instance of the synthesized voice audio data can be audibly rendered at the client device of the representative of "Hypothetical Cafe" and can capture the synthesized voice 554A2 of "Hi, I’m a virtual assistant calling on behalf of Jane Doe to see if anyone found a red leather jacket yesterday".
[0085] In this example, and in instances of generating the synthesized speech audio data, the automated assistant can cause this data to be applied as input across a previously trained LLM that is fine-tuned and / or initiated based on unstructured free-form natural language input 552A1 to generate an output such as a probability distribution over a vocabulary of terms and / or phrases. Based on the probability distribution over the vocabulary of terms and / or phrases, the automated assistant can cause the chatbot to select text data corresponding to the synthesized speech 554A2. Further, the automated assistant can cause the chatbot to process the text data corresponding to the synthesized speech 554A2 using a TTS model to generate an instance of the synthesized speech audio data that is audibly rendered at a client device of a representative of "Hypothetical Cafe". Additionally, at least in part because the previously trained LLM is fine-tuned and / or initiated based on the unstructured free-form natural language input 552A1 provided by Jane Doe during interaction 500A1, the automated assistant is able to cause the chatbot to generate an output and / or select text data corresponding to the synthesized speech 554A2. Accordingly, the automated assistant is able to cause the chatbot to perform the task of submitting a query to a representative of an entity at "Hypothetical Café" on behalf of Jane Doe.
[0086] Further assume that the representative of "Hypothetical Cafe" responds to the synthesized speech 554A2 with an oral input 556A2 of "Please hold while I go check lost and found". In this example, the automated assistant can cause the chatbot to process the audio data capturing the oral input 556A2, the text data corresponding to the audio data capturing the oral input 556A2 (e.g., determined using an ASR model), and / or any context conversation data to generate an additional instance of the synthesized speech audio data. The additional instance of the synthesized speech audio data can be audibly rendered at a client device of the representative of "Hypothetical Cafe" and can capture the synthesized speech 558A2 of "Okay". Further, the automated assistant can cause the chatbot to monitor for additional oral inputs from the representative of "Hypothetical Cafe" to indicate that the representative has rejoined the phone call after placing the chatbot on hold.
[0087] Further assume that the representative of "Hypothetical Cafe" returns from holding by providing the verbal input 560A2 of "I found her red leather jacket, I’ll keep it at the hostess stand for her". In this example, the automated assistant can cause the chatbot to process the audio data capturing the verbal input 560A2, the text data corresponding to the audio data capturing the verbal input 560A2 (e.g., determined using an ASR model), and / or any contextual conversation data to generate a further additional instance of the synthesized speech audio data. The further additional instance of the synthesized speech audio data can be audibly rendered at the client device of the representative of "Hypothetical Cafe" and can capture the synthesized speech 562A2 of "Thank you, I will let her know". Further, the automated assistant can cause the chatbot to terminate the phone call with the representative of "Hypothetical Cafe".
[0088] In this example, in response to determining that the task has been successfully completed, the automated assistant can cause the chatbot to terminate the corresponding conversation with the representative of "Hypothetical Café". The automated assistant or the chatbot can determine that the task has been successfully completed based on, for example, the representative of "Hypothetical Café" responding to the query with the confirmation that Jane Doe did in fact leave her red leather jacket at Hypothetical Café the previous night and that she can pick it up at the hostess stand. Thus, based on the response to the query provided by the representative of "Hypothetical Café", the automated assistant can determine the responsive content 552A3 that can be provided for presentation to Jane Doe at interaction 500A3. The responsive content 552A3 can include the result of the task performed by the chatbot, such as "Your red leather jacket is at Hypothetical Café, they’ll leave it at the hostess stand for you". Thus, interaction 500A3 can be a notification generated for presentation to the user, or can be provided for presentation to the user during a subsequent conversation session between Jane Doe and the automated assistant that is at least partially executed at her client device.
[0089] It should be noted that in interaction 500A2, the chatbot implements various peripheral behaviors. For example, the chatbot introduces itself as "virtual assistant calling on behalf of Jane Doe" in the synthesized voice 554A2 by using a greeting behavior that enables the chatbot to identify Jane Doe and identify itself as a chatbot; the chatbot places itself on hold after the synthesized voice 558A2 by using a hold behavior that enables the chatbot to pause and resume the corresponding conversation; and the voice terminates the phone call after providing the synthesized voice 562A2 by using an exit behavior to terminate the corresponding conversation with the representative of "Hypothetical Café". These peripheral behaviors are a non-limiting example of structuring free-form natural language input 552A1 using a previously trained LLM to enable the chatbot to perform generalization of conversations before fine-tuning and / or starting the chatbot to enable the chatbot to perform generalization of conversations. Further, it should be noted that other peripheral behaviors can be implemented by the chatbot, and those Figure 5A described are for illustrative purposes and are not meant to be restrictive.
[0090] In addition, in various implementations, the chatbot generated to perform the task of submitting a query to the representative of "Hypothetical Café" can be discarded. For example, the chatbot can be discarded in response to determining that the task has been successfully completed. As another example, the chatbot can be discarded in response to determining that the chatbot may not be used to participate in any additional corresponding conversations due to the task being successfully completed. However, in various implementations, the chatbot may not always successfully perform the task during the corresponding conversation.
[0091] Specifically referring to Figure 5B and assuming that a conversation on the user's (e.g., Jane Doe) device with the client device is initiated in the same or similar manner as described with respect to Figure 5A as indicated by interaction 500B1. For Figure 5BFor the purposes of the example in, assume that Jane Doe provides the unstructured free-form natural language input 552B1 of "Tell Bobby Jones the locksmith that I transferred the money to his quick cash account" as an oral input. In this example, the automated assistant can use an ASR model to process the audio data capturing the oral input to generate an ASR output. Further, the automated assistant can use an NLU model to process the ASR output to generate an NLU output. Notably, the oral input explicitly includes an intent of a [notification entity] (etc.) having a slot value for a [notification content] parameter (e.g., task data) such as "I transferred the money to [your] quick cash account" (etc.), and the oral input explicitly includes an [entity] (e.g., entity data) of [Bobby Jones the locksmith]. Thus, the automated assistant can provide a response 554A2 of "Okay, I’ll let Bobby Jones know" to be presented audibly and / or visually to Jane Doe to indicate that the automated assistant will perform the task of notifying the entity of "Bobby Jones".
[0092] In this example, and based on the unstructured free-form natural language input 552B1, the automated assistant can cause the client device and / or the remote system to generate a chatbot to perform the task of notifying the representative of "Bobby Jones" (e.g., where "Bobby Jones" is an entity and where "Bobby Jones" is his own representative). In this example, the automated assistant can determine a corresponding identifier associated with "Bobby Jones", such as a phone number or a contact entry that can be used to initiate a corresponding conversation with the representative of "Bobby Jones". Thus, and as shown in interaction 500B2, the automated assistant can cause the chatbot to be implemented at the computing device in the same or a similar manner as described with respect to Figure 5A This enables the chatbot to make a phone call to the phone number associated with "Bobby Jones" to facilitate the performance of the task of notifying "Bobby Jones".
[0093] Further assume that after a phone call is made, the chatbot and the representative of "Bobby Jones" participate in the corresponding conversation. For example, further assume that the representative of "Bobby Jones" answers an incoming phone call, and the chatbot generates an instance of the synthesized voice audio data in the same or similar manner as described regarding Figure 5A and the synthesized voice audio data is audibly rendered at the client device of the representative of "Bobby Jones". In this example, further assume that the instance of the synthesized voice audio data captures the synthesized voice 552B2 of "Hi, I’m a virtual assistant calling on behalf of Jane Doe, I wanted to let you know that she transferred the money she owed you to your quick cash account".
[0094] Further assume that the representative of "Bobby Jones" responds to the synthesized voice 552B2 with the verbal input 554B2 of "I don’t know what you’re talking about, I don’t know any Jane Doe?". In this example, the automated assistant can cause the chatbot to at least process the audio data capturing the verbal input 554B2 to generate an additional instance of the synthesized voice audio data. The additional instance of the synthesized voice audio data can be audibly rendered at the client device of the representative of "Bobby Jones" and can capture the synthesized voice 556B2 of "In that case, I’ll reach out to Jane". In other words, based on the unsuccessful execution of a task (e.g., based on determining that "Bobby Jones" does not expect money to be transferred to his quick cash account and / or based on determining that "Bobby Jones" does not know who Jane Doe is), the automated assistant can cause the chatbot to utilize an exit behavior that enables the chatbot to prompt Jane Doe to join the corresponding conversation.
[0095] Accordingly, and as shown at interaction 500B3, the automated assistant can cause the chatbot to generate and provide a prompt 552B3 of "Hi Jane, I need you to talk to Bobby Jones the locksmith about the payment" for presentation to Jane Doe. The prompt can include a specific reason why the chatbot was unable to successfully perform the task (e.g., "Bobby Jones does not know who you are"). As a result, and as shown at interaction 500B4, Jane Doe can respond to prompt 500B3 by joining the corresponding conversation and providing an oral input 552B1 of "Hi Bobby, sorry for the confusion, this is Jane Doe, the payment I transferred to you was on behalf of my father, John Doe, and the work you did for him last week". Accordingly, the automated assistant can cause the chatbot to participate in the corresponding conversation with the representative of "Bobby Jones", but in cases where the chatbot does not successfully complete the task, can cause the chatbot to prompt Jane Doe to join the corresponding conversation.
[0096] Notably, in Figure 5A and Figure 5B the examples, for the initial interaction with the corresponding representative (e.g., Figure 5A interaction 500A2 of Figure 5B and interaction 500B2 of Figure 5B ), Jane Doe is not an active participant in the corresponding conversation between the chatbot and the representative. However, the automated assistant can cause the chatbot to prompt Jane Doe to join the corresponding conversation when needed (e.g., as shown in the example of
[0097] AlthoughFigure 5A and Figure 5B The corresponding conversation is described as a telephone call, but it should be understood that this is for illustrative purposes and is not meant to be restrictive. For example, the corresponding conversation can be a text-based conversation via any text-based platform or service through which a chatbot can participate in the corresponding conversation with an entity or its representative (e.g., text or SMS messaging, email, and / or other text-based platforms). Further, although Figure 5A and Figure 5B the corresponding representative in the examples of Figure 5A and Figure 5B is a human, it should be understood that this is for illustrative purposes and is not meant to be restrictive. For example, the corresponding representative can be a corresponding additional chatbot deployed on behalf of the corresponding entity. In these cases, the chatbot can participate in the corresponding conversation with the corresponding additional chatbot. Additionally, although Figure 5A and Figure 5B the unstructured free-form natural language input of
[0098] is described as a single sentence, it should be understood that this is for illustrative purposes and is not meant to be restrictive. It is noted that since users provide different unstructured free-form natural language inputs for different tasks, Figure 6A and Figure 6B the chatbots generated and implemented in Figure 6A and Figure 6B can be different chatbots. Figure 6A and Figure 6B the corresponding unstructured free-form natural language input described in the examples of Figure 1 can be received at the user's client device (e.g., Figure 6A and Figure 6B client device 110 of Figure 1 ), the chatbots described in the examples of Figure 1 can be generated at the user's client device (e.g., Figure 6A and Figure 6B client device 110 of Figure 1 and / or at a remote system (e.g., remote system 160 fromFigure 1 implemented at the remote system 160) and communicate with the representative of the entity via additional computing devices of the representative. Each of these computing devices may include corresponding components such as user interface input components (e.g., microphones, visual components, presence sensors, touch-sensitive displays, keyboards, hardware buttons, software buttons, etc.), user interface output components (e.g., touch-sensitive displays, speakers, monitors, projectors, etc.), network interfaces, and / or other components. Thus, although Figure 6A and Figure 6B the interactions 600A and 600B are depicted within a single interface respectively, it should be understood that this is for the purpose of illustrating the various techniques described herein and is not meant to be restrictive.
[0099] Specifically referring to Figure 6A , assume that a conversation with a user (e.g., Jane Doe) of a client device is initiated in the same or similar manner as described with respect to Figure 5A and Figure 5B as indicated by interaction 600A1. However, compared to the examples in Figure 5A and Figure 5B , assume that Jane Doe provides multiple verbal inputs 652A1, 654A1, 656A1, 658A1, and 660A1 as shown in interaction 600A. Although the multiple verbal inputs 652A1, 654A1, 656A1, 658A1, and 660A1 convey details of a more complex task (e.g., booking catering services for a luncheon) than those described with respect to Figure 5A and Figure 5B (e.g., submitting a query in Figure 5A and providing a notification in Figure 5B ), the automated assistant can still generate a chatbot to book the catering services for the luncheon and cause the chatbot to participate in the corresponding conversation with a representative of "Hypothetical Café".
[0100] In this example, based on a task of Figure 6A including multiple subtasks, the task can be considered more complex than those from Figure 5A and Figure 5BThe tasks are more complex. These subtasks can include, for example, determining whether "Hypothetical Café" is available to provide luncheon catering on a specified date / time and for a specified number of people (based on verbal input 652A1 and verbal input 654A1); determining whether "Hypothetical Café" is available to provide luncheon catering with specified dietary restrictions or menu requests (based on verbal input 654A1); determining whether "Hypothetical Café" is available to provide luncheon catering at a specified price (based on verbal input 658A1); and if "Hypothetical Café" is available to provide luncheon catering at a specified price, proactively paying for the luncheon (based on verbal input 658A1). It is worth noting that each of these subtasks includes different intents and different slot values of parameters associated with different intents. As a result, in Figure 6A 's example, since Figure 6A 's tasks are more complex, the chatbot can be fine-tuned and / or initiated with more task data than the chatbots in Figure 5A and Figure 5B . Therefore, the automated assistant can provide a response 662A1 of "Okay, I’ll call Hypothetical Café and let you know how it goes" to be presented audibly and / or visually to Jane Doe to indicate that the automated assistant will perform the task of booking a luncheon with the entity "Hypothetical Café".
[0101] Specifically referring to Figure 6B , and as shown in interaction 600B1, the automated assistant can cause the chatbot to behave in a manner similar to that regarding Figure 5A and Figure 5BImplemented in the same or similar manner at a computing device. Further assume that after a phone call is made, the chatbot and a representative of "Hypothetical Café" participate in a corresponding conversation. For example, further assume that a representative of "Hypothetical Café" answers an incoming phone call and provides the verbal input 652B1 of "Hello, this is John Smith at Hypothetical Café, how may I help you?". In this example, the automated assistant can process the task data determined based on the interaction 600A1, capture the audio data of the verbal input 652B1, the text data corresponding to the captured audio data of the verbal input 652B1 (e.g., determined using an ASR model), and / or any contextual conversation data, so as to generate an instance of the synthesized speech audio data in the same or similar manner as described with respect to Figure 5A and Figure 5B An instance of the synthesized speech audio data can be audibly rendered at the client device of the representative of "Hypothetical Cafe", and the synthesized speech 654B1 of "Hi, I’m a virtual assistant calling on behalf of Jane Doe to see if you can cater her luncheon at noon on 12 / 12 / 2022" can be captured.
[0102] Further assume that a representative of "Hypothetical Cafe" uses the verbal input 656B1 of "Okay, tell me a bit more about the luncheon" to respond to the synthesized speech 654B1. In this example, the automated assistant can cause the chatbot to process task data, capture the audio data of the verbal input 656B1, the text data corresponding to the captured audio data of the verbal input 656B1 (e.g., determined using an ASR model), and / or any contextual conversation data to generate additional instances of the synthesized speech audio data. The additional instances of the synthesized speech audio data can be audibly rendered at the client device of the representative of "Hypothetical Cafe" and can capture the synthesized speech 658B1 of "There will be 50 people, some with gluten intolerance and other…", but the chatbot is interrupted by the verbal input 660B1 of "I’m sorry did you say 50 people or 15 people?" provided by the representative of "Hypothetical Café". In this example, the automated assistant can cause the chatbot to repeat the statement "50 people" in the synthesized speech 662B1 by using a clarification behavior that enables the chatbot to clarify and / or repeat information previously provided during the corresponding conversation.
[0103] Further assume that a representative of "Hypothetical Cafe" responds to the synthesized speech 662B1 with an oral input 664B1 of "50 people got it, please go on". In this example, the automated assistant can enable the chatbot to resume the corresponding conversation by generating a further additional instance of the synthesized speech that includes information not provided to the representative of "Hypothetical Café". For example, a further additional instance of the synthesized speech audio data can be audibly rendered at the client device of the representative of "Hypothetical Cafe" and can capture the synthesized speech 666B1 of "Some of the people have gluten intolerance and others are vegan, Jane would like a sandwich platter with several different options, a salad bar, and other side options". Further assume that a representative of "Hypothetical Cafe" responds to the synthesized speech 666B1 with an oral input 668B1 of "Perfect, we can cater the luncheon for $225". As a result, the automated assistant can enable the chatbot to complete the task of booking the luncheon based on the successful execution of all other subtasks of the task by generating yet another additional instance of the synthesized speech audio data. The yet another additional instance of the synthesized speech audio data can be audibly rendered at the client device of the representative of "Hypothetical Cafe" and can capture the synthesized speech 670B1 of "Excellent, here is Jane Doe’s credit card information [*provide credit card information*], please forward the details to janedoe@exampleurl.com".
[0104] In this example, in response to determining that a task (or all subtasks) has been successfully completed, the automated assistant can cause the chatbot to terminate the corresponding conversation with a representative of "Hypothetical Café". The automated assistant or chatbot can determine that the task has been successfully completed based on, for example, confirmation from a representative of "Hypothetical Café" to cater the luncheon. Thus, based on the successful reservation of the luncheon with the representative of "Hypothetical Café", the automated assistant can determine responsive content 652B2 that can be provided to be presented to Jane Doe at interaction 600B2. Responsive content 652B2 can include the result of the task performed by the chatbot, such as "Hypothetical Café is scheduled to cater your luncheon, John Smith will forward the details to your email". Thus, interaction 600B2 can be a summary of interaction 600B1 that is generated to be presented to the user, or can be provided to be presented to the user during a subsequent conversation session between Jane Doe and the automated assistant that is at least partially executed at her client device.
[0105] Turning now to Figure 7 , Figure 7 is a block diagram of an example computing device 710 that can optionally be used to perform one or more aspects of the techniques described herein. In some implementations, one or more of the client device, remote system components, and / or other components can include one or more components of the example computing device 710.
[0106] Computing device 710 typically includes at least one processor 714 that communicates with a plurality of peripheral devices via a bus subsystem 712. These peripheral devices can include a storage subsystem 724 (including, for example, a memory subsystem 725 and a file storage subsystem 726), a user interface output device 720, a user interface input device 722, and a network interface subsystem 716. The input and output devices allow a user to interact with computing device 710. The network interface subsystem 716 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.
[0107] The user interface input device 722 can include a keyboard, a pointing device (such as a mouse, trackball, touchpad, or graphics tablet), a scanner, a touch screen incorporated into a display (e.g., a touch-sensitive display), an audio input device (such as a voice recognition system, a microphone), and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into the computing device 710 or onto a communication network.
[0108] The user interface output device 720 can include a display subsystem, a printer, a fax machine, or a non-visual display (such as an audio output device). The display subsystem can include a cathode ray tube (CRT), a flat panel device (such as a liquid crystal display (LCD)), a projection device, or some other mechanism for creating a visible image. The display subsystem can also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways of outputting information from the computing device 710 to a user or to another machine or computing device.
[0109] The storage subsystem 724 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 724 can include selected aspects for performing the methods disclosed herein and the logic for implementing Figure 1 and Figure 2 the various components depicted.
[0110] These software modules are typically executed by the processor 714, either alone or in conjunction with other processors. The memory 725 used in the storage subsystem 724 can include multiple memories, including a main random access memory (RAM) 730 for storing instructions and data during program execution and a read-only memory (ROM) 732 in which fixed instructions are stored. The file storage subsystem 726 can provide persistent storage for program and data files and can include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical disk drive, or a removable media cartridge. The modules implementing the functionality of certain implementations can be stored in the storage subsystem 724 by the file storage subsystem 726 or in other machines accessible by the processor 714.
[0111] The bus subsystem 712 provides the mechanism for enabling the various components and subsystems of the computing device 710 to communicate with each other as expected. Although the bus subsystem 712 is schematically shown as a single bus, alternative implementations of the bus subsystem 712 can use multiple buses.
[0112] The computing device 710 can have different types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computing device 710 depicted Figure 7 is only intended as a specific example for the purpose of illustrating some implementations. Many other configurations of the computing device 710 are possible, which have more or fewer components compared to the computing device depicted Figure 7 herein.
[0113] In cases where the systems described herein collect or otherwise monitor personal information about a user or can utilize personal information and / or the information monitored, the user may be provided with the following opportunities: to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographical location), or to control whether and / or how content that may be more relevant to the user is received from a content server. Additionally, certain data can be processed in one or more ways before it is stored or used so that personally identifiable information is removed. For example, a user's identity can be processed so that the user's personally identifiable information cannot be determined, or in the case of obtaining geographical location information, the user's geographical location can be generalized (such as to the city, zip code, or state level) so that the user's specific geographical location cannot be determined. Thus, the user can control how information about the user is collected and / or used.
[0114] In some implementations, a method implemented by one or more processors of a client device is provided, and the method includes: receiving, at the client device, unstructured free-form natural language input from a user of the client device; and in response to receiving the unstructured free-form natural language input including one or more tasks associated with an entity, locally generating, at the client device and at least based on the unstructured free-form natural language input, a chatbot for performing, on behalf of the user, one or more tasks associated with the entity. The unstructured free-form natural language input includes one or more tasks associated with the entity. The method further includes causing the chatbot to perform, on behalf of the user, one or more tasks associated with the entity. Causing the chatbot to perform, on behalf of the user, one or more tasks associated with the entity includes: causing the chatbot to participate in a corresponding conversation with the entity; during the corresponding conversation with the entity: causing the chatbot to render multiple instances of the synthesized speech audio data for presentation to a representative of the entity; and receiving responsive content in response to at least a given instance of the synthesized speech audio data. At least the given instance of the synthesized speech audio data among the multiple instances of the synthesized speech audio data conveys details of one or more tasks associated with the entity. The method further includes causing the responsive content to be provided for presentation to the user of the client device.
[0115] These and other implementations of the techniques disclosed herein may optionally include one or more of the following features.
[0116] In some implementations, generating a chatbot for performing, on behalf of the user, one or more tasks associated with an entity may include: obtaining a previously trained large language model (LLM); causing the previously trained LLM to be fine-tuned based on the unstructured free-form natural language input to generate a fine-tuned LLM; and using the fine-tuned LLM as the chatbot.
[0117] In some versions of those implementations, causing the chatbot to render a given instance of the synthesized speech audio data for presentation to a representative may include: using a fine-tuned LLM to process one or more features of an unstructured free-form natural language input to generate a given instance of text data that conveys details of one or more tasks associated with an entity; using a text-to-speech (TTS) model to process the given instance of text data that conveys details of one or more tasks associated with the entity to generate the given instance of the synthesized speech audio data; and transmitting the given instance of the synthesized speech audio data from the client device and to an additional client device of the representative. Transmitting the given instance of the synthesized speech audio data to the additional client device may cause the additional client device to audibly render the given instance of the synthesized speech audio data via one or more speakers of the additional client device for presentation to the representative.
[0118] In some further versions of those implementations, the method may further include using a fine-tuned LLM and also one or more of the features of the unstructured free-form natural language input to process the corresponding context of the corresponding conversation to generate a given instance of text data that conveys details of one or more tasks associated with the entity.
[0119] In additional or alternative further versions of those implementations, the method may further include, in response to the given instance of the synthesized speech audio data being audibly rendered via one or more speakers of the additional client device for presentation to the representative: receiving, at the client device and from the additional client device, a given instance of response audio data that includes responsive content in response to at least the given instance of the synthesized speech audio data; using an automatic speech recognition (ASR) model to process the given instance of the response audio data to generate a given instance of response text data; and determining, based on the given instance of the response text data, whether execution of one or more of the tasks was successfully completed during the corresponding conversation.
[0120] In some yet further versions of those implementations, causing the responsive content to be provided for presentation to the user of the client device may be in response to determining that one or more of the tasks were successfully completed during the corresponding conversation.
[0121] In additional or alternative yet further versions of those implementations, the method may further include, in response to determining that one or more of the tasks were not successfully completed during the corresponding conversation: causing the chatbot to render an indication that is prompting the user to join the corresponding conversation for presentation to the representative of the entity; generating a prompt to request the user to join the corresponding conversation; and causing the prompt to request the user to join the corresponding conversation to be provided for presentation to the user at the client device.
[0122] In even further versions of those implementations, the prompt can further include specific reasons for why one or more of the tasks in the task were not successfully completed during the corresponding conversation.
[0123] In additional or alternative further versions of those implementations, the method can further include extracting one or more of the features from the unstructured free-form natural language input before processing one or more of the features of the unstructured free-form natural language input using the fine-tuned LLM.
[0124] In some further versions of those implementations, one or more of the features can be explicitly included in the unstructured free-form natural language input, and extracting one or more of the features explicitly included in the unstructured free-form natural language input from the unstructured free-form natural language input can include using an input parser to extract one or more of the features explicitly included in the unstructured free-form natural language input.
[0125] In additional or alternative further versions of those implementations, one or more of the features can be implicitly included in the unstructured free-form natural language input, and extracting one or more of the features implicitly included in the unstructured free-form natural language input from the unstructured free-form natural language input can include using an input parser to identify one or more of the features implicitly included in the unstructured free-form natural language input; and using a coreference resolver to extract one or more of the features implicitly included in the unstructured free-form natural language input.
[0126] In even further versions of those implementations, the coreference resolver can access user data locally generated at the client device to extract one or more of the features implicitly included in the unstructured free-form natural language input, and the user data can include one or more of the following: historical location data, historical time data, user preference data, user account data, calendar information, or email data.
[0127] In additional or alternative versions of those implementations, a previously trained LLM can be stored in the device storage of the client device, and the previously trained LLM that can be stored in the device storage of the client device can be a sparsified version of a global previously trained LLM available at a remote system communicatively coupled to the client device.
[0128] In some further versions of those implementations, the fine-tuned LLM can be stored in the device storage of the client device.
[0129] In yet further versions of those implementations, the method can further include, after causing the chatbot to perform one or more tasks associated with an entity on behalf of the user: discarding the fine-tuned LLM from the device storage of the client device; and refraining from discarding the previously trained LLM from the device storage of the client device.
[0130] In some implementations, the method can further include: receiving at the client device additional unstructured free-form natural language input from a user of the client device; and in response to receiving the additional unstructured free-form natural language input including one or more additional tasks associated with the entity or an additional entity: locally generating at the client device, at least based on the additional natural language input, an additional chatbot for performing one or more additional tasks associated with the entity or the additional entity on behalf of the user. The additional unstructured free-form natural language input can include one or more additional tasks associated with the entity or the additional entity. The method can further include causing the additional chatbot to perform one or more additional tasks associated with the entity or the additional entity on behalf of the user. Causing the additional chatbot to perform one or more additional tasks associated with the entity or the additional entity on behalf of the user can include: causing the additional chatbot to participate in an additional corresponding conversation with the entity or the additional entity; during the additional corresponding conversation with the entity or the additional entity: causing the chatbot to render multiple additional instances of the synthesized speech audio data for presentation to a representative of the entity or an additional representative of the additional entity; and receiving additional responsive content in response to at least a given additional instance of the synthesized speech audio data; and causing the additional responsive content to be provided for presentation to the user of the client device. At least a given additional instance of the synthesized speech audio data among the multiple additional instances of the synthesized speech audio data conveys additional details of one or more additional tasks associated with the entity or the additional entity.
[0131] In some implementations, the method can further include identifying an entity associated with one or more tasks based on the unstructured free-form natural language input; determining a corresponding identifier for the entity associated with the one or more tasks; and using the corresponding identifier for the entity associated with the one or more tasks to cause the chatbot to participate in a corresponding conversation with the entity.
[0132] In some versions of those implementations, the corresponding identifier for an entity associated with one or more tasks can be the corresponding phone number for the entity, and causing the chatbot to participate in a corresponding conversation with the entity using the corresponding identifier for the entity associated with one or more tasks can include causing the chatbot to initiate an automated phone call on behalf of the user using the corresponding phone number for the entity to perform one or more tasks associated with the entity.
[0133] In some implementations, causing the chatbot to participate in a corresponding conversation with the entity can include causing the chatbot to answer a phone call received at the client device and from a representative of the entity; and causing the chatbot to participate in the corresponding conversation with the entity as part of the phone call.
[0134] In some implementations, the user may not be an active participant in the corresponding conversation between the chatbot and the representative.
[0135] In some implementations, during the corresponding conversation with the entity, the method can further include receiving a request from the representative for the user to join the corresponding conversation; and in response to receiving the request for the user to join the corresponding conversation: generating a prompt requesting the user to join the corresponding conversation; and causing the prompt requesting the user to join the corresponding conversation to be provided for presentation to the user at the client device.
[0136] In some implementations, the entity can be explicitly identified in an unstructured free-form natural language input.
[0137] In some implementations, the entity may not be explicitly identified in the unstructured free-form natural language input, the entity can be a specific type of entity, the specific type of the entity can be explicitly identified in the unstructured free-form natural language input, and one or more tasks can be associated with the specific type of the entity.
[0138] In some versions of those implementations, one or more tasks may also be associated with additional entities that are other than and of a particular type belonging to the entity. Causing the chatbot to perform one or more tasks associated with the entity on behalf of the user may further include causing the chatbot to engage in an additional corresponding conversation with the additional entity; during the additional corresponding conversation with the additional entity: causing the chatbot to render multiple additional instances of the synthesized speech audio data for presentation to an additional representative of the additional entity; and receiving additional responsive content in response to at least a given additional instance of the synthesized speech audio data. At least a given additional instance of the synthesized speech audio data among the multiple additional instances of the synthesized speech audio data may convey details of one or more tasks that are also associated with the additional entity. The method may further include causing the additional responsive content to be provided for presentation to the user of the client device.
[0139] In some further versions of those implementations, the chatbot may conduct the corresponding conversation with the entity and the additional corresponding conversation with the additional entity in a parallel manner.
[0140] In additional or alternative further versions of those implementations, the chatbot may conduct the additional corresponding conversation with the additional entity after the corresponding conversation with the entity and in response to determining that one or more of the tasks were not successfully completed during the corresponding conversation.
[0141] In some implementations, the representative of the entity may be a human representative. In other implementations, the representative of the entity may be an additional chatbot trained to conduct the corresponding conversation on behalf of the entity.
[0142] In some implementations, the unstructured free-form natural language input may be typed or oral input that conveys details of one or more tasks associated with the entity without defining a corresponding dialogue state diagram, a dialogue state of the corresponding dialogue state diagram, or a dialogue state transition of the corresponding dialogue state diagram to be utilized in the execution of the one or more tasks associated with the entity.
[0143] In some versions of those implementations, the unstructured free-form natural language input may be typed input that conveys details of one or more tasks associated with the entity in one or more sentences, and the responsive content to be provided for presentation to the user of the client device may be the result of the execution of one or more tasks associated with the entity.
[0144] In additional or alternative versions of those implementations, the unstructured free-form natural language input can be typed input that conveys details of one or more tasks associated with an entity in one or more paragraphs, and the responsive content to be provided for presentation to a user of the client device can be a summary of a corresponding conversation with the entity.
[0145] In some implementations, a method implemented by one or more processors of a remote system is provided, and the method includes: receiving, at the remote system and from a client device, an indication of unstructured free-form natural language input from a user of the client device; and in response to receiving the indication of unstructured free-form natural language input that includes one or more tasks associated with an entity: remotely generating, at the remote system and at least based on the indication of the natural language input, a chatbot for performing one or more tasks associated with the entity on behalf of the user. The unstructured free-form natural language input includes one or more tasks associated with the entity. The method further includes causing the chatbot to perform one or more tasks associated with the entity on behalf of the user. Causing the chatbot to perform one or more tasks associated with the entity on behalf of the user includes: causing the chatbot to participate in a corresponding conversation with the entity; during the corresponding conversation with the entity: causing the chatbot to render multiple instances of the synthesized speech audio data for presentation to a representative of the entity; and receiving responsive content in response to at least a given instance of the synthesized speech audio data. At least the given instance of the synthesized speech audio data among the multiple instances of the synthesized speech audio data conveys details of one or more tasks associated with the entity. The method further includes transmitting, from the remote system and to the client device, an indication of the responsive content. Transmitting the indication of the responsive content to the client device causes the client device to provide the responsive content for presentation to a user of the client device.
[0146] These and other implementations of the techniques disclosed herein optionally can include one or more of the following features.
[0147] In some implementations, generating a chatbot for performing one or more tasks associated with an entity on behalf of a user can include: obtaining a previously trained large language model (LLM); causing the previously trained LLM to be fine-tuned based on the unstructured free-form natural language input to generate a fine-tuned LLM; and using the fine-tuned LLM as the chatbot.
[0148] In some versions of those implementations, the previously trained LLM can be stored in a remote storage of the remote system, and the previously trained LLM that can be stored in the remote storage of the remote system can be an unsparsified version of a global previously trained LLM that is available at the remote system communicatively coupled to the client device.
[0149] In some further versions of those implementations, the fine-tuned LLM can be stored in a remote storage of a remote system.
[0150] In yet further versions of those implementations, the method can further include, after causing the chatbot to perform one or more tasks associated with an entity on behalf of a user: discarding the fine-tuned LLM from the remote storage of the remote system; and refraining from discarding the previously trained LLM from the remote storage of the remote system.
[0151] In some implementations, a method implemented by one or more processors of a client device is provided, and the method includes: receiving, at the client device, unstructured free-form natural language input from a user of the client device; and in response to receiving an indication of unstructured free-form natural language input that includes one or more tasks associated with an entity: remotely generating, at a remote system and at least based on the indication of the natural language input, a chatbot for performing one or more tasks associated with the entity on behalf of the user. The unstructured free-form natural language input includes one or more tasks associated with the entity. The method further includes causing the chatbot to perform one or more tasks associated with the entity on behalf of the user. Causing the chatbot to perform one or more tasks associated with the entity on behalf of the user includes causing the chatbot to participate in a corresponding conversation with the entity; during the corresponding conversation with the entity: causing the chatbot to render multiple instances of text data for presentation to a representative of the entity; and receiving responsive content in response to at least a given instance of the text data. At least the given instance of the text data among the multiple instances of the text data conveys details of one or more tasks associated with the entity. The method further includes causing the responsive content to be provided for presentation to the user of the client device.
[0152] These and other implementations of the techniques disclosed herein optionally may include one or more of the following features.
[0153] In some implementations, the method can further include identifying an entity associated with one or more tasks based on the unstructured free-form natural language input; determining a corresponding identifier for the entity associated with the one or more tasks; and using the corresponding identifier for the entity associated with the one or more tasks to cause the chatbot to participate in the corresponding conversation with the entity.
[0154] In versions of those implementations, the corresponding identifier for an entity associated with one or more tasks can be one or more of the following: the corresponding telephone number for the entity, the corresponding email address for the entity, or the corresponding username for the entity. Using the corresponding identifier for an entity associated with one or more tasks to enable a chatbot to participate in a corresponding conversation with the entity can include enabling the chatbot to use the corresponding identifier for the entity to initiate an automated text-based messaging session on behalf of the user to perform one or more tasks associated with the entity.
[0155] In some implementations, enabling a chatbot to participate in a corresponding conversation with an entity can include: enabling the chatbot to respond to a text message, email, or other text-based message received at the client device and from a representative of the entity; and enabling the chatbot to participate in the corresponding conversation with the entity to facilitate the text message, email, or other text-based message.
[0156] In some implementations, a method implemented by one or more processors of a remote system is provided, and the method includes: receiving, at the remote system and from a client device, an indication of an unstructured free-form natural language input from a user of the client device; and in response to receiving the indication of the unstructured free-form natural language input including one or more tasks associated with an entity: remotely generating, at the remote system and at least based on the indication of the natural language input, a chatbot for performing one or more tasks associated with the entity on behalf of the user. The unstructured free-form natural language input includes one or more tasks associated with the entity. The method further includes causing the chatbot to perform one or more tasks associated with the entity on behalf of the user. Causing the chatbot to perform one or more tasks associated with the entity on behalf of the user includes: enabling the chatbot to participate in a corresponding conversation with the entity; during the corresponding conversation with the entity: enabling the chatbot to render multiple instances of the synthesized speech audio data for presentation to the representative of the entity; and receiving responsive content in response to at least a given instance of the synthesized speech audio data. At least a given instance of the synthesized speech audio data among the multiple instances of the synthesized speech audio data conveys details of one or more tasks associated with the entity. The method further includes transmitting, from the remote system and to the client device, an indication of the responsive content. Transmitting the indication of the responsive content to the client device causes the client device to provide the responsive content for presentation to the user of the client device.
[0157] In some implementations, a method implemented by one or more processors of a client device is provided, and the method includes: receiving, at the client device, unstructured free-form natural language input from a user of the client device, the unstructured free-form natural language input including one or more tasks associated with an entity; in response to receiving the unstructured free-form natural language input including a natural language description of a corresponding dialogue state diagram, identifying a chatbot for performing, on behalf of the user, one or more tasks associated with the entity; and causing the chatbot to perform, on behalf of the user, one or more tasks associated with the entity. Causing the chatbot to perform, on behalf of the user, one or more tasks associated with the entity includes causing the chatbot to participate in a corresponding conversation with the entity; during the corresponding conversation with the entity: causing the chatbot to render multiple instances of the synthesized speech audio data for presentation to a representative of the entity, wherein at least a given instance of the synthesized speech audio data among the multiple instances of the synthesized speech audio data conveys details of one or more tasks associated with the entity; and receiving responsive content in response to at least the given instance of the synthesized speech audio data. The method further includes causing the responsive content to be provided for presentation to the user of the client device.
[0158] These and other implementations of the techniques disclosed herein optionally may include one or more of the following features.
[0159] In some implementations, identifying the chatbot for performing, on behalf of the user, one or more tasks associated with the entity may include obtaining a previously trained large language model (LLM); and causing the previously trained LLM to be used as the chatbot.
[0160] In some versions of those implementations, the previously trained LLM may be stored in the device storage of the client device, and the previously trained LLM that may be stored in the device storage of the client device may be a sparsified version of a global previously trained LLM available at a remote system communicatively coupled to the client device.
[0161] In some versions of those implementations, the method may further include avoiding causing the previously trained LLM to be fine-tuned based on the unstructured free-form natural language input.
[0162] In some versions of those implementations, causing a chatbot to render a given instance of synthesized speech out of multiple instances of synthesized speech for presentation to a representative of an entity can include: using a previously trained LLM to process unstructured free-form natural language input and task data reflecting details of one or more tasks associated with the entity to generate an instance of text data reflecting a given behavior for a given implicit dialogue state; using a text-to-speech (TTS) model to process the given instance of text data reflecting the given behavior for the given implicit dialogue state to generate the given instance of synthesized speech; and transmitting the given instance of synthesized speech from a client device and to an additional client device of the representative of the entity. Transmitting the given instance of synthesized speech to the additional client device can cause the additional client device to audibly render the given instance of synthesized speech via one or more speakers of the additional client device for presentation to the representative of the entity.
[0163] In addition, some implementations include one or more processors of one or more computing devices (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)), where the one or more processors are operable to execute instructions stored in an associated memory, and where the instructions are configured to cause the execution of any of the methods described above. Some implementations also include one or more non-transitory computer-readable storage media that store computer instructions that can be executed by the one or more processors to execute any of the methods described above. Some implementations also include a computer program product that includes instructions that can be executed by the one or more processors to execute any of the methods described above.
[0164] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein.
Claims
1. A method implemented by one or more processors of a client device, the method comprising: Receiving, at the client device, unstructured free-form natural language input from a user of the client device, the unstructured free-form natural language input including one or more tasks associated with an entity; In response to receiving the unstructured free-form natural language input including the one or more tasks associated with the entity: Locally generating, at the client device and at least based on the unstructured free-form natural language input, a chatbot for performing, on behalf of the user, the one or more tasks associated with the entity; and Causing the chatbot to perform, on behalf of the user, the one or more tasks associated with the entity, wherein causing the chatbot to perform, on behalf of the user, the one or more tasks associated with the entity includes: Causing the chatbot to participate in a corresponding conversation with the entity; During the corresponding conversation with the entity: Causing the chatbot to render multiple instances of the synthesized speech audio data for presentation to a representative of the entity, wherein at least a given instance of the synthesized speech audio data among the multiple instances of the synthesized speech audio data conveys details of the one or more tasks associated with the entity; and Receiving responsive content in response to at least the given instance of the synthesized speech audio data; and Causing the responsive content to be provided for presentation to the user of the client device.
2. The method of claim 1, wherein generating the chatbot for performing, on behalf of the user, the one or more tasks associated with the entity includes: Obtaining a previously trained large language model (LLM); Causing the previously trained LLM to be fine-tuned based on the unstructured free-form natural language input to generate a fine-tuned LLM; And Utilizing the fine-tuned LLM as the chatbot.
3. The method of claim 2, wherein causing the chatbot to render the given instance of the synthesized speech audio data for presentation to the representative includes: Using the fine-tuned LLM to process one or more features of the unstructured free-form natural language input to generate a given instance of text data that conveys the details of the one or more tasks associated with the entity; Using a text-to-speech (TTS) model to process the given instance of the text data that conveys the details of the one or more tasks associated with the entity to generate the given instance of the synthesized speech audio data; And Transmitting, from the client device and to an additional client device of the representative, the given instance of the synthesized speech audio data, wherein transmitting the given instance of the synthesized speech audio data to the additional client device causes the additional client device to audibly render the given instance of the synthesized speech audio data via one or more speakers of the additional client device for presentation to the representative.
4. The method according to claim 3, further comprising: processing the corresponding context of the corresponding conversation using the fine-tuned LLM and one or more of the features of the unstructured free-form natural language input to generate the given instance of the text data conveying the details of the one or more tasks associated with the entity.
5. The method according to claim 3, further comprising: in response to the given instance of the synthesized voice audio data being audibly rendered via the one or more speakers of the additional client device for presentation to the representative receiving, at the client device and from the additional client device, a given instance of response audio data including the responsive content in response to at least the given instance of the synthesized voice audio data; processing the given instance of the response audio data using an automatic speech recognition (ASR) model to generate a given instance of response text data; and determining, based on the given instance of the response text data, whether the execution of one or more of the tasks was successfully completed during the corresponding conversation.
6. The method according to claim 5, wherein the responsive content is provided for presentation to the user of the client device in response to determining that one or more of the tasks were successfully completed during the corresponding conversation.
7. The method according to claim 5, further comprising: in response to determining that one or more of the tasks were not successfully completed during the corresponding conversation: causing the chatbot to render an indication that is prompting the user to join the corresponding conversation for presentation to the representative of the entity; generating a prompt requesting the user to join the corresponding conversation; and causing the prompt requesting the user to join the corresponding conversation to be provided for presentation to the user at the client device.
8. The method according to claim 7, wherein the prompt further includes a specific reason for why one or more of the tasks were not successfully completed during the corresponding conversation.
9. The method according to claim 3, further comprising: before processing one or more of the features of the unstructured free-form natural language input using the fine-tuned LLM: extracting one or more of the features from the unstructured free-form natural language input.
10. The method according to claim 9, wherein one or more of the features are explicitly included in the unstructured free-form natural language input, and wherein extracting one or more of the features that are explicitly included in the unstructured free-form natural language input from the unstructured free-form natural language input includes: utilizing an input analyzer to extract one or more of the features that are explicitly included in the unstructured free-form natural language input.
11. The method according to claim 9, wherein one or more of the features are implicitly included in the unstructured free-form natural language input, and wherein extracting one or more of the features that are implicitly included in the unstructured free-form natural language input from the unstructured free-form natural language input comprises: using an input analyzer to identify one or more of the features that are implicitly included in the unstructured free-form natural language input; and using a coreference resolver to extract one or more of the features that are implicitly included in the unstructured free-form natural language input.
12. The method according to claim 11, wherein the coreference resolver accesses user data locally generated at the client device to extract one or more of the features that are implicitly included in the unstructured free-form natural language input, and wherein the user data comprises one or more of: historical location data, historical time data, user preference data, user account data, calendar information, or email data.
13. The method according to claim 2, wherein the previously trained LLM is stored in the device storage of the client device, and wherein the previously trained LLM stored in the device storage of the client device is a sparsified version of a global previously trained LLM available at a remote system communicatively coupled to the client device.
14. The method according to claim 13, wherein the fine-tuned LLM is stored in the device storage of the client device.
15. The method according to claim 14, further comprising: after causing the chatbot to perform the one or more tasks associated with the entity on behalf of the user: discarding the fine-tuned LLM from the device storage of the client device; and avoiding discarding the previously trained LLM from the device storage of the client device.
16. The method according to any one of the preceding claims, further comprising: receiving, at the client device, additional unstructured free-form natural language input from the user of the client device, the additional unstructured free-form natural language input comprising one or more additional tasks associated with the entity or an additional entity; in response to receiving the additional unstructured free-form natural language input comprising the one or more additional tasks associated with the entity or the additional entity: locally generating, at the client device and at least based on the additional natural language input, an additional chatbot for performing the one or more additional tasks associated with the entity or the additional entity on behalf of the user; and Cause the additional chatbot to perform, on behalf of the user, the one or more additional tasks associated with the entity or the additional entity, wherein causing the additional chatbot to perform, on behalf of the user, the one or more additional tasks associated with the entity or the additional entity includes: Cause the additional chatbot to participate in an additional corresponding conversation with the entity or the additional entity; During the additional corresponding conversation with the entity or the additional entity: Cause the chatbot to render a plurality of additional instances of the synthesized speech audio data for presentation to the representative of the entity or the additional representative of the additional entity, wherein at least a given additional instance of the synthesized speech audio data among the plurality of additional instances of the synthesized speech audio data conveys additional details of the one or more additional tasks associated with the entity or the additional entity; and Receive additional responsive content in response to at least the given additional instance of the synthesized speech audio data; and Cause the additional responsive content to be provided for presentation to the user of the client device.
17. The method according to any one of the preceding claims, further comprising: Identifying the entity associated with the one or more tasks based on the unstructured free-form natural language input; Determining a corresponding identifier for the entity associated with the one or more tasks; And Using the corresponding identifier for the entity associated with the one or more tasks to cause the chatbot to participate in the corresponding conversation with the entity.
18. The method according to claim 17, wherein the corresponding identifier for the entity associated with the one or more tasks is a corresponding telephone number for the entity, and wherein using the corresponding identifier for the entity associated with the one or more tasks to cause the chatbot to participate in the corresponding conversation with the entity includes: Causing the chatbot to initiate an automated telephone call on behalf of the user using the corresponding telephone number for the entity to perform the one or more tasks associated with the entity.
19. The method according to any one of the preceding claims, wherein causing the chatbot to participate in the corresponding conversation with the entity includes: Causing the chatbot to answer a telephone call received at the client device and from the representative of the entity; And Causing the chatbot to participate in the corresponding conversation with the entity as part of the telephone call.
20. The method according to any one of the preceding claims, wherein the user is not an active participant in the corresponding conversation between the chatbot and the representative.
21. The method according to any one of the preceding claims, during the corresponding conversation with the entity, further comprising: Receiving from the representative a request for the user to join the corresponding conversation; And In response to receiving the request for the user to join the corresponding conversation: Generating a prompt requesting the user to join the corresponding conversation; And The prompt to request the user to join the corresponding conversation is provided to be presented to the user at the client device.
22. The method according to any one of the preceding claims, wherein the entity is explicitly identified in the unstructured free-form natural language input.
23. The method according to any one of claims 1 to 21, wherein the entity is not explicitly identified in the unstructured free-form natural language input, wherein the entity is a specific type of entity, wherein the specific type of entity is explicitly identified in the unstructured free-form natural language input, and wherein the one or more tasks are associated with the specific type of entity.
24. The method according to claim 23, wherein the one or more tasks are further associated with additional entities that are other than the entity and belong to the specific type of entity, and wherein causing the chatbot to perform the one or more tasks associated with the entity on behalf of the user further comprises: causing the chatbot to participate in an additional corresponding conversation with the additional entity; during the additional corresponding conversation with the additional entity: causing the chatbot to render multiple additional instances of the synthesized speech audio data to be presented to an additional representative of the additional entity, wherein at least a given additional instance of the synthesized speech audio data among the multiple additional instances of the synthesized speech audio data conveys the details of the one or more tasks that are also associated with the additional entity, and receiving additional responsive content in response to at least the given additional instance of the synthesized speech audio data; and causing the additional responsive content to be provided to be presented to the user of the client device.
25. The method according to claim 24, wherein the chatbot conducts the corresponding conversation with the entity and the additional corresponding conversation with the additional entity in a parallel manner.
26. The method according to claim 24, wherein the chatbot conducts the additional corresponding conversation with the additional entity after the corresponding conversation with the entity and in response to determining that one or more of the tasks have not been successfully completed during the corresponding conversation.
27. The method according to any one of the preceding claims, wherein the representative of the entity is a human representative.
28. The method according to any one of claims 1 to 26, wherein the representative of the entity is an additional chatbot trained to conduct the corresponding conversation on behalf of the entity.
29. The method according to any one of the preceding claims, wherein the unstructured free-form natural language input is a typed input or an oral input that conveys the details of the one or more tasks associated with the entity without defining a corresponding dialogue state diagram, a dialogue state of the corresponding dialogue state diagram, or a dialogue state transition of the corresponding dialogue state diagram to be utilized in the execution of the one or more tasks associated with the entity.
30. The method according to claim 29, wherein the unstructured free-form natural language input is typed input that conveys the details of the one or more tasks associated with the entity in one or more sentences, and wherein the responsive content to be provided for presentation to the user of the client device is the result of the execution of the one or more tasks associated with the entity.
31. The method according to any one of claims 1 to 29, wherein the unstructured free-form natural language input is typed input that conveys the details of the one or more tasks associated with the entity in one or more paragraphs, and wherein the responsive content to be provided for presentation to the user of the client device is a summary of the corresponding conversation with the entity.
32. A method implemented by one or more processors of a remote system, the method comprising: Receiving, at the remote system and from a client device, an indication of an unstructured free-form natural language input from a user of the client device, the unstructured free-form natural language input including one or more tasks associated with an entity; In response to receiving the indication of the unstructured free-form natural language input including the one or more tasks associated with the entity: Remotely generating, at the remote system and at least based on the indication of the natural language input, a chatbot for performing the one or more tasks associated with the entity on behalf of the user; And Causing the chatbot to perform the one or more tasks associated with the entity on behalf of the user, wherein causing the chatbot to perform the one or more tasks associated with the entity on behalf of the user includes: Causing the chatbot to participate in a corresponding conversation with the entity; During the corresponding conversation with the entity: Causing the chatbot to render multiple instances of synthesized speech audio data for presentation to a representative of the entity, wherein at least a given instance of the synthesized speech audio data among the multiple instances of the synthesized speech audio data conveys the details of the one or more tasks associated with the entity; and Receiving responsive content in response to at least the given instance of the synthesized speech audio data; and Transmitting, from the remote system and to the client device, an indication of the responsive content, wherein transmitting the indication of the responsive content to the client device causes the client device to provide the responsive content for presentation to the user of the client device.
33. The method according to claim 32, wherein generating the chatbot for performing the one or more tasks associated with the entity on behalf of the user includes: Obtaining a previously trained large language model (LLM); Causing the previously trained LLM to be fine-tuned based on the unstructured free-form natural language input to generate a fine-tuned LLM; And Utilizing the fine-tuned LLM as the chatbot.
34. The method according to claim 33, wherein the previously trained LLM is stored in a remote storage of the remote system, and wherein the previously trained LLM stored in the remote storage of the remote system is an unsparsified version of a global previously trained LLM that is available at the remote system communicatively coupled to the client device.
35. The method according to claim 34, wherein the fine-tuned LLM is stored in the remote storage of the remote system.
36. The method according to claim 35, further comprising: after causing the chatbot to perform the one or more tasks associated with the entity on behalf of the user: discarding the fine-tuned LLM from the remote storage of the remote system; and avoiding discarding the previously trained LLM from the remote storage of the remote system.
37. A method implemented by one or more processors of a client device, the method comprising: receiving, at the client device, unstructured free-form natural language input from a user of the client device, the unstructured free-form natural language input including one or more tasks associated with an entity; in response to receiving the indication of the unstructured free-form natural language input including the one or more tasks associated with the entity: remotely generating, at the remote system and at least based on the indication of the natural language input, a chatbot for performing the one or more tasks associated with the entity on behalf of the user; and causing the chatbot to perform the one or more tasks associated with the entity on behalf of the user, wherein causing the chatbot to perform the one or more tasks associated with the entity on behalf of the user includes: causing the chatbot to participate in a corresponding conversation with the entity; during the corresponding conversation with the entity: causing the chatbot to render multiple instances of text data for presentation to a representative of the entity, wherein at least a given instance of the text data among the multiple instances of the text data conveys details of the one or more tasks associated with the entity; and receiving responsive content in response to at least the given instance of the text data; and causing the responsive content to be provided for presentation to the user of the client device.
38. The method according to claim 37, further comprising: identifying the entity associated with the one or more tasks based on the unstructured free-form natural language input; determining a corresponding identifier for the entity associated with the one or more tasks; and utilizing the corresponding identifier for the entity associated with the one or more tasks to cause the chatbot to participate in the corresponding conversation with the entity.
39. The method according to claim 38, wherein the corresponding identifier for the entity associated with the one or more tasks is one or more of the following: a corresponding telephone number for the entity, a corresponding email address for the entity, or a corresponding username for the entity, and wherein causing the chatbot to participate in the corresponding conversation with the entity using the corresponding identifier for the entity associated with the one or more tasks includes: Causing the chatbot to initiate a text-based messaging session automated on behalf of the user using the corresponding identifier for the entity to perform the one or more tasks associated with the entity.
40. The method according to any one of claims 37 to 39, wherein causing the chatbot to participate in the corresponding conversation with the entity includes: Causing the chatbot to respond to a text message, email, or other text-based message received at the client device and from a representative of the entity; And Causing the chatbot to participate in the corresponding conversation with the entity to facilitate the text message, the email, or the other text-based message.
41. A method implemented by one or more processors of a remote system, the method comprising: Receiving, at the remote system and from a client device, an indication of an unstructured free-form natural language input from a user of the client device, the unstructured free-form natural language input including one or more tasks associated with an entity; In response to receiving the indication of the unstructured free-form natural language input including the one or more tasks associated with the entity: Remotely generating, at least based on the indication of the natural language input and at the remote system, a chatbot for performing the one or more tasks associated with the entity on behalf of the user; And Causing the chatbot to perform the one or more tasks associated with the entity on behalf of the user, wherein causing the chatbot to perform the one or more tasks associated with the entity on behalf of the user includes: Causing the chatbot to participate in a corresponding conversation with the entity; During the corresponding conversation with the entity: Causing the chatbot to render multiple instances of the synthesized voice audio data for presentation to a representative of the entity, wherein at least a given instance of the synthesized voice audio data among the multiple instances of the synthesized voice audio data conveys details of the one or more tasks associated with the entity; and Receiving responsive content in response to at least the given instance of the synthesized voice audio data; and Transmitting, from the remote system and to the client device, an indication of the responsive content, wherein transmitting the indication of the responsive content to the client device causes the client device to provide the responsive content for presentation to the user of the client device.
42. A method implemented by one or more processors of a client device, the method comprising: Receiving, at the client device, unstructured free-form natural language input from a user of the client device, the unstructured free-form natural language input including one or more tasks associated with an entity; Identifying, in response to receiving the unstructured free-form natural language input including the natural language description of the corresponding dialogue state diagram, a chatbot for performing, on behalf of the user, the one or more tasks associated with the entity; and Causing the chatbot to perform, on behalf of the user, the one or more tasks associated with the entity, wherein causing the chatbot to perform, on behalf of the user, the one or more tasks associated with the entity includes: Causing the chatbot to participate in a corresponding conversation with the entity; During the corresponding conversation with the entity: Causing the chatbot to render multiple instances of the synthesized speech audio data for presentation to a representative of the entity, wherein at least a given instance of the synthesized speech audio data among the multiple instances of the synthesized speech audio data conveys details of the one or more tasks associated with the entity; and Receiving responsive content in response to at least the given instance of the synthesized speech audio data; and Causing the responsive content to be provided for presentation to the user of the client device.
43. The method of claim 42, wherein identifying the chatbot for performing, on behalf of the user, the one or more tasks associated with the entity includes: Obtaining a previously trained large language model (LLM); And Causing the previously trained LLM to be used as the chatbot.
44. The method of claim 43, wherein the previously trained LLM is stored in the on-device storage of the client device, and wherein the previously trained LLM stored in the on-device storage of the client device is a sparsified version of a global previously trained LLM available at a remote system communicatively coupled to the client device.
45. The method of claim 43, further comprising: Avoiding causing the previously trained LLM to be fine-tuned based on the unstructured free-form natural language input.
46. The method of claim 43, wherein causing the chatbot to render the given instance of the synthesized speech among the multiple instances of the synthesized speech for presentation to the representative of the entity includes: Using the previously trained LLM to process the unstructured free-form natural language input and task data reflecting the details of the one or more tasks associated with the entity to generate an instance of text data reflecting a given behavior of the given implicit dialogue state; Using a text-to-speech (TTS) model to process the given instance of the text data reflecting the given behavior of the given implicit dialogue state to generate the given instance of the synthesized speech; And Transmit the given instance of the synthesized voice from the client device and to an additional client device of the representative of the entity, wherein transmitting the given instance of the synthesized voice to the additional client device causes the additional client device to audibly render the given instance of the synthesized voice via one or more speakers of the additional client device for presentation to the representative of the entity.
47. A system comprising: At least one processor; And A memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the method according to any one of claims 1 to 46.
48. A non-transitory computer-readable storage medium storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the operations of the method according to any one of claims 1 to 46.
Citation Information
Cited By
Virtual robot generation method and device, equipment and storage medium
CN121235107A