Resolving a unique person identifier during a corresponding conversation between a voice bot and a human

By processing the automatic speech recognition results through multiple machine learning layers, candidate personal identifiers are generated and refined, solving the problem of automated assistants misidentifying personal identifiers and achieving more efficient and accurate personal identifier parsing and dialogue processing.

CN115735248BActive Publication Date: 2026-03-24GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-22
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing automated assistants are prone to misidentifying personal identifiers, such as email addresses, physical addresses, and usernames, when parsing spoken language, leading to unnecessary consumption of computing resources and privacy issues.

Method used

The system processes the automatic speech recognition results through multiple machine learning layers, generates candidate personal identifiers, refines them through clarification requests, and finally determines the unique personal identifier. It also utilizes multiple machine learning layers to process the synthetic speech intent and ASR speech hypothesis to quickly and accurately parse the personal identifier.

Benefits of technology

It improves the accuracy and efficiency of parsing personal identifiers, reduces unnecessary consumption of computing resources and privacy risks, and supports faster dialogue processing and task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115735248B_ABST
    Figure CN115735248B_ABST
Patent Text Reader

Abstract

Implementations relate to causing a voice bot to utilize a plurality of ML layers to resolve a unique personal identifier for a human while the voice bot is engaged in a corresponding conversation with the human. The unique personal identifier can include a unique alphanumeric character sequence that is personal to the human. In some implementations, ASR speech hypotheses corresponding to spoken utterances that include the unique personal identifier can be processed to generate candidate unique personal identifiers, a given alphanumeric character of a candidate unique personal identifier can be selected, and the voice bot can prompt the human to clarify the given alphanumeric character with a clarification request until it is predicted to correspond to an actual unique personal identifier for the human. The unique personal identifier can then be used to perform further actions by the voice bot and / or other systems.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Humans can engage in a human-to-computer dialog with interactive software applications referred to herein as "automated assistants" (also referred to as "chatbots," "interactive personal assistants," "intelligent personal assistants," "personal voice assistants," "conversational agents," etc.). For example, humans (which when they interact with automated assistants can be referred to as "users") can provide commands and / or requests to an automated assistant using spoken natural language input (i.e., spoken utterances), which can in some cases be converted into text and then processed. The automated assistant typically responds by providing responsive user interface output (e.g., audible and / or visual user interface output), controlling a smart device, and / or performing other actions.

[0002] Automated assistants typically rely on a pipeline of components when interpreting and responding to user requests. For example, an automatic speech recognition (ASR) engine can be used to process audio data corresponding to a spoken utterance, generating a transcription (i.e., a sequence of terms and / or other tokens) of the user's utterance. However, in performing ASR, certain terms can be misrecognized. Such misrecognition can be exacerbated when the spoken utterance corresponds to a sequence of terms and / or other tokens that are unpredictable and / or out of vocabulary. For example, email addresses, physical addresses, usernames, etc. can include sequences of letters, numbers, and / or symbols that are personal and meaningful to a user, but which are frequently misrecognized by ASR.

[0003] As a result of such misrecognition, the automated assistant can wastefully perform actions that the user did not desire, or prevent further actions from being performed. This can cause the user to repeat the same spoken utterance (which can again be misrecognized) or cause the user to perform some other action, thereby prolonging the human-to-computer dialog and / or causing additional computing resources to be consumed outside of the human-to-computer dialog. Additionally or alternatively, such misrecognition can cause the automated assistant to unnecessarily utilize network resources by incorrectly transmitting emails and / or other electronic communications to misrecognized email addresses, usernames, and / or other personal identifiers. This can raise privacy concerns, as the automated assistant can incorrectly transmit content that is personal to a user to an incorrect user. Additionally or alternatively, such misrecognition can cause the automated assistant to request a human to take over the human-to-computer dialog, thereby prolonging the human-to-computer dialog and / or causing additional computing resources to be consumed in requesting the human to take over the human-to-computer dialog. SUMMARY

[0004] Embodiments disclosed herein relate to causing a voice bot to utilize a plurality of machine learning (ML) layers to resolve a unique personal identifier for a corresponding human while the voice bot is engaged in a corresponding conversation with the corresponding human. The unique personal identifier can include a unique sequence of alphanumeric characters that is personal to the human. The unique personal identifier can be, for example, an email address, a physical address, a username, a password, a name of an entity, a product identifier, a domain name, and / or any other unique personal identifier. In some embodiments, a plurality of ML layers can be used to process one or more automatic speech recognition (ASR) speech hypotheses corresponding to a spoken utterance that includes a unique personal identifier to generate one or more candidate unique personal identifiers. Each of the one or more candidate unique personal identifiers can include one or more corresponding alphanumeric characters, each corresponding alphanumeric character being associated with a corresponding prediction measure. Further, one or more corresponding alphanumeric characters of a candidate unique personal identifier can be selected based on the corresponding prediction measures, and the voice bot can generate one or more prompts having a corresponding clarification request requesting clarification regarding the one or more corresponding alphanumeric characters for the unique personal identifier. Based on a corresponding response from the human, the one or more candidate unique personal identifiers can be refined. The voice bot can generate one or more further prompts and continue to refine the one or more candidate unique personal identifiers until it is predicted that a given unique personal identifier corresponds to an actual unique personal identifier provided by the human. The given unique personal identifier can then be used by the voice bot for one or more further actions, such as facilitating the corresponding conversation with the given unique identifier and / or facilitating another action after the corresponding conversation, such as the given unique identifier.

[0005] As one example, assume that a corresponding conversation between a human and a voice bot occurs during a phone call initiated by the human and that the corresponding conversation is associated with a human call customer service of a utility company (e.g., a water supplier, a gas and electricity supplier, a cable or internet supplier, etc.). In this example, the voice bot can solicit a unique personal identifier corresponding to an email address of the human to verify an identity of the human, look up a service associated with the email address, and / or perform any other action requested by the human during the corresponding conversation. The voice bot can use an ASR model to process audio data that captures a spoken utterance from the human and that includes the email address to generate a plurality of ASR speech hypotheses. Further assume that the spoken utterance in this example that includes the email address of the human is “john and then p@exampleurl.com.” In this example, the plurality of ASR speech hypotheses can include ASR speech hypotheses for the “john p” portion of the email address of “jon and then p,” “john and then p,” “jon and then d,” “john and then d,” and / or other ASR speech hypotheses. In this example, the voice bot can use a plurality of ML layers to process one or more of the ASR speech hypotheses to generate one or more candidate unique personal identifiers.

[0006] Further, the voice robot can generate one or more prompts based on the corresponding prediction measure. For example, the voice robot can generate the prompt "and is that john with an h or no h?". Further assume that the human provides the additional spoken utterance "john with a h". In this example, the voice robot can use the ASR model to process additional audio data that captures the additional spoken utterance to generate a plurality of additional ASR speech hypotheses, and the voice robot can use the plurality of ML layers to process one or more of the additional speech hypotheses to refine the corresponding prediction measure and / or the one or more candidate unique person identifiers. In this example, the voice robot can update at least the corresponding prediction measure associated with the alphanumeric character "h" to indicate a high likelihood (e.g., using a binary value, a probability, a log-likelihood, etc.) that there is an email address of the human that starts with the sequence of alphanumeric characters "j o h n". In doing so, the voice robot can limit the one or more candidate unique person identifiers to a subset that is limited to those that start with the sequence of alphanumeric characters "j o h n", thereby eliminating any candidate unique person identifiers that exclude the alphanumeric character "h". Further, the voice robot can generate an additional prompt (e.g., "j o h n and then was that p as in papa or d as in delta?") and continue to refine the corresponding prediction measure and / or the one or more candidate unique person identifiers until it is predicted that the email address corresponds to the human.

[0007] In some implementations, the voice bot can use the plurality of ML layers to process one or more ASR speech hypotheses in response to predicting that the audio data capturing the spoken utterance includes a unique personal identifier. In some versions of those implementations, the voice bot can predict that the spoken utterance includes a unique personal identifier based on particular synthetic speech audio data that includes synthetic speech previously generated by the voice bot that has been provided for presentation to the human during a corresponding conversation. For example, if the voice bot generates synthetic speech audio data that includes synthetic speech that requests the human to provide a unique personal identifier (e.g., “what is your email address?”), the voice bot can predict that the spoken utterance includes a unique personal identifier. In some additional or alternative implementations, the voice bot can predict whether the spoken utterance includes a personal identifier based on the plurality of ASR speech hypotheses generated using the ASR model. For example, if one or more of the plurality of ASR speech hypotheses includes a given alphanumeric character or alphanumeric string of characters that indicates a unique personal identifier (e.g., a string of numbers or cities and state information for a physical address, a particular symbol or character (e.g., an “@” symbol, an underscore, etc.), and / or any other indicator that the spoken utterance includes a unique personal identifier), the system can predict that the spoken utterance includes a unique personal identifier.

[0008] In some implementations, the plurality of ML layers can correspond to those of a transformer ML model (e.g., an input layer, an encoding layer, a decoding layer, a feedforward layer, an attention layer, an output layer, and / or other ML layers), a unidirectional and / or bidirectional RNN model (e.g., an input layer, a hidden layer, an output layer, and / or other ML layers), and / or other ML layers of other ML models. In some implementations, the plurality of ML layers can be used to process one or more ASR speech hypotheses to generate a likelihood tree for a unique personal identifier. The likelihood tree can include a plurality of nodes and a plurality of edges. Each of the plurality of nodes can be associated with a given alphanumeric character that is predicted for the unique personal identifier (e.g., include the given alphanumeric character or include data that maps the node to the given alphanumeric character). Further, each of the plurality of nodes can include a corresponding prediction measure that is associated with the given alphanumeric character of a corresponding node for the unique personal identifier. The plurality of edges can connect one or more of the plurality of nodes. In implementations in which the likelihood tree is generated by processing one or more ASR speech hypotheses with the plurality of ML layers, one or more candidate unique personal identifiers can be generated based on the likelihood tree, and a given one of the candidate unique personal identifiers can be selected that is associated with a node having a corresponding prediction measure that is predicted to correspond to the unique personal identifier. In additional or alternative implementations, the plurality of ML layers can directly use the plurality of ML layers to generate one or more candidate unique personal identifiers when processing one or more ASR speech hypotheses.

[0009] In various implementations, the voice bot can process a corresponding intent of the voice bot associated with synthetic speech audio data presented for presentation to the human prior to receiving the spoken utterance along with one or more ASR speech hypotheses using a plurality of ML layers. The intent of the voice bot can include, for example, requesting the human to provide a unique personal identifier, requesting the human to spell a unique personal identifier, requesting the human to clarify one or more alphanumeric characters of a unique personal identifier, requesting the human to verify a unique personal identifier, and / or any other intent. By processing the intent of the voice bot along with the one or more ASR speech hypotheses, the plurality of ML layers can be leveraged to resolve the correct personal identifier in a faster and more efficient manner. For example, if the voice bot generates a prompt requesting clarification regarding a given alphanumeric character (e.g., “is that p as in papa or d as in delta” for the email address “johnp@exampleurl.com”), the intent of the voice bot associated with the prompt previously presented to the human can be processed using the ML layers to leverage the intent to refine the unique personal identifier regarding the particular alphanumeric character. In implementations in which a plurality of ML layers are leveraged to generate a likelihood tree for a unique personal identifier, this enables the voice bot to perform a beam search on the likelihood tree regarding the particular alphanumeric character and to quickly and efficiently update the likelihood tree to quickly and efficiently refine one or more candidate unique personal identifiers. Furthermore, in implementations in which the intent of the voice bot is leveraged, the plurality of ML layers can achieve a higher level of robustness and / or accuracy with fewer training instances as compared to not leveraging the intent.

[0010] In some implementations, the voice bot can determine whether to generate one or more prompts based on the corresponding predictive measures associated with the one or more candidate unique person identifiers. In some versions of those implementations, the voice bot can generate a prompt requesting a human to spell if the corresponding predictive measure associated with one or more of the alphanumeric characters fails to satisfy a threshold, or provide the unique person identifier on a character-by-character basis (e.g., “can you spell that for me?” and the like). In some additional or alternative versions of those implementations, the voice bot can generate a prompt requesting a human to clarify one or more particular alphanumeric characters for the unique person identifier (e.g., “so it begins with t as in tango?” “was that f as in foxtrot or s as in sierra?” and the like). In some additional or alternative versions of those implementations, the voice bot can generate a prompt requesting a human to verify the unique person identifier (e.g., “so the email address is j o h n p@exampleurl.com” and the like). In some implementations, the voice bot can predict that one or more of the candidate unique person identifiers corresponds to the actual unique person identifier based on each of the corresponding predictive measures associated with each of the alphanumeric characters for a given one of the candidate unique person identifiers satisfying a threshold.

[0011] In some implementations, the corresponding conversation between the human and the voice bot can be conducted during a telephone call performed using various voice communication protocols. In additional or alternative implementations, the corresponding conversation between the human and the voice bot can be conducted during a human-to-bot conversation session between the human and the voice bot. As described above, in these and other implementations, the voice bot can utilize a given unique personal identifier to facilitate the corresponding conversation in response to determining that the given one of the candidate unique personal identifiers does correspond to an actual unique personal identifier. In implementations in which the corresponding conversation is conducted during a telephone call between the voice bot and the human, the voice bot can utilize the unique personal identifier to continue performing a task requested by the human (e.g., for customer service, for a query related to a user account, and / or any other task that can be performed during a telephone call). For example, the voice bot can utilize the unique personal identifier to verify or authenticate the identity of the human, search for information related to the unique personal identifier, and / or the voice bot can utilize the unique personal identifier to continue the telephone call in any other manner. In implementations in which the corresponding conversation is conducted during a conversation session between the voice bot and the human, the voice bot can utilize the unique personal identifier to incorporate the unique personal identifier into a transcription (e.g., while the human is instructing the voice bot to dictate an email, a text message, an SMS message, a note, a calendar entry, and / or otherwise instruct the voice bot), perform an action on behalf of the user (e.g., make a purchase on behalf of the user, log into an account of the user, and / or any other action on behalf of the user), deliver content to the user (e.g., electronic content when the unique personal identifier corresponds to an email address and / or physical content when the unique personal identifier corresponds to a physical address), and / or the voice bot can utilize the unique personal identifier to facilitate the corresponding conversation in any other manner.

[0012] In various implementations, and prior to the voice bot utilizing the plurality of ML layers to determine a unique person identifier during a corresponding conversation, the plurality of ML layers can be trained based on a plurality of training instances. Each of the plurality of training instances can include a training instance input and a training instance output. The training instance input can include one or more of: one or more ASR speech hypotheses for a unique person identifier, audio data corresponding to the one or more ASR speech hypotheses, or an intent of the voice bot associated with synthetic speech audio data that was presented for presentation to a human prior to receiving the audio data. The training instance output can include a ground truth output corresponding to the unique person identifier, the ground truth output including a ground truth alphanumeric character for the unique person identifier and / or a corresponding ground truth predictive measure of the ground truth alphanumeric character for the unique person identifier. In some implementations, and with appropriate permissions from the participants, one or more of the plurality of training instances can be generated based on an actual conversation including at least one human participant (e.g., an actual conversation between a human and the voice bot or an actual conversation between multiple humans). In additional or alternative implementations, one or more of the plurality of training instances can be synthetically generated based on actual unique person identifiers stored in one or more databases.

[0013] In some versions of those implementations, the plurality of ML layers can be trained for utilization in a single round of a corresponding conversation. For example, for a given training instance, the plurality of ML layers can be used to process a training instance input to generate a corresponding predicted measurement for each alphanumeric character included in a unique personal identifier. Further, the corresponding predicted measurement for each alphanumeric character included in the unique personal identifier can be compared to a training instance output (e.g., on a character-by-character basis). Based on this comparison, one or more losses can be generated, and one or more of the plurality of ML layers can be updated based on one or more of the losses. In some additional or alternative versions of those implementations, the plurality of ML layers can be further trained for utilization in n additional rounds of a corresponding conversation via a simulator (e.g., where n is a positive integer greater than 1). For example, one or more candidate unique personal identifiers can be generated based on the corresponding predicted measurements (and optionally using a likelihood tree), a simulated voice bot portion of the simulator can process the one or more candidate unique personal identifiers to generate a simulated prompt, a simulated human portion of the simulator can generate a simulated response based on the simulated prompt and based on a ground truth output for the given training instance, and the plurality of ML layers can process the simulated response to refine the one or more candidate unique personal identifiers. For example, assume that the unique personal identifier is an email address “johnp@exampleurl.com” and that the corresponding predicted measurements for the predicted alphanumeric characters of the email address indicate that it can end with either a “p” or a “d”. In this example, a simulated prompt generated using the simulated voice bot portion of the simulator can be “was that p as in papa or das in delta” and a simulated response generated using the simulated human portion of the simulator can be “p as in papa” based on the prompt and based on the ground truth output including the corresponding alphanumeric character “p”. The simulation can continue in this iterative manner for n rounds until the one or more candidate unique personal identifiers are predicted to correspond to the actual unique personal identifier associated with the given training instance. In various implementations, the plurality of ML layers can be trained based on a plurality of training instances for a single round of a corresponding conversation until one or more conditions are satisfied prior to further training for n rounds of the corresponding conversation.

[0014] In some implementations, the plurality of training instances used to train the plurality of ML layers can be obtained based on actual conversations and / or synthetically generated to reflect actual distributions of unique person identifiers. This allows the plurality of ML layers to obtain a high level of accuracy and / or traceability of unique person identifiers in actual use when utilized by a voice bot. Moreover, by obtaining a high accuracy and / or traceability of unique person identifiers, corresponding conversations including unique person identifiers can be ended more quickly and efficiently because the plurality of ML layers utilized by the voice bot and trained using the techniques described herein are more capable of understanding subtle differences in human speech and responding accordingly to resolve unique person identifiers. Furthermore, the voice bot utilizing the plurality of ML layers described herein is more scalable and reduces memory consumption because the plurality of ML layers can be shared among a plurality of different voice bots. For example, a plurality of third parties can develop respective voice bots for specific tasks without having to train the respective voice bots to determine unique person identifiers. Instead, the respective voice bots can each simply use the plurality of ML layers (or respective instances thereof).

[0015] The above description is provided only as an overview of some implementations disclosed herein. Those implementations and other implementations are described in more detail herein. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 A block diagram depicting an example environment that illustrates various aspects of the disclosure and in which implementations disclosed herein can be implemented is depicted.

[0017] Figure 2A An example process flow for training a plurality of machine learning layers utilized by a voice bot in determining unique person identifiers according to various implementations is depicted.

[0018] Figure 2B An example process flow for utilizing a plurality of machine learning layers by a voice bot in determining unique person identifiers according to various implementations is depicted.

[0019] Figure 3 A flowchart illustrating an example method of obtaining training instances for training a plurality of machine learning layers utilized by a voice bot in determining unique person identifiers according to various implementations is depicted.

[0020] Figure 4 A flowchart illustrating an example method of training a plurality of machine learning layers utilized by a voice bot in determining unique person identifiers according to various implementations is depicted.

[0021] Figure 5A flow diagram illustrating an example method of utilizing multiple machine learning layers by a voice bot in determining a unique person identifier, in accordance with various embodiments, is depicted.

[0022] Figure 6A FIGS. 6B and 6C depict various non-limiting examples of corresponding conversations between a voice bot and a human including determining a unique person identifier, in accordance with various embodiments.

[0023] Figure 7 An example architecture of a computing device, in accordance with various embodiments, is depicted. DETAILED DESCRIPTION

[0024] Turning now to Figure 1 A block diagram of an example environment in which various aspects of the present disclosure can be implemented and in which embodiments disclosed herein can be implemented is depicted. A client device 110 is illustrated in Figure 1 and includes, in various embodiments, a user input engine 111, a rendering engine 112, and a voice bot client 113. The client device 110 can be, for example, a standalone auxiliary device (e.g., having a microphone, a speaker, and / or a display), a smartphone, a laptop computer, a desktop computer, a tablet computer, a wearable computing device, a vehicle computing device, and / or any other client device capable of implementing the voice bot development system client 113.

[0025] The user input engine 111 can detect various types of user input at the client device 110. User input detected at the client device 110 can include verbal input detected via a microphone of the client device 110, touch input detected via a user interface input device (e.g., a touchscreen) of the client device 110, and / or typed input detected via a user interface input device of the client device 110 (e.g., via a virtual keyboard on a touchscreen, a physical keyboard, a mouse, a stylus, and / or any other user interface input device of the client device 110). The rendering engine 112 can cause content to be visually and / or aurally rendered at the client device 110 for presentation to a user (or human) via a user interface output device. The output can include, for example, various types of user interfaces associated with a voice bot and / or notifications associated with a voice bot that can be visually rendered via a display of the client device 110 and / or aurally rendered via a speaker of the client device 110, and / or any other output described herein that can be visually and / or aurally rendered via the client device 110.

[0026] In various implementations, the voice bot client 113 can include an automatic speech recognition (ASR) engine 130A1, a natural language understanding (NLU) engine 140A1, and a text-to-speech (TTS) engine 150A1. Further, the voice bot development client 113 can communicate with the voice bot system 120 over one or more networks 1991 (e.g., any combination of Wi-Fi, Bluetooth, near field communication (NFC), local area network (LAN), wide area network (WAN), Ethernet, the Internet, and / or other networks). From the perspective of a user interacting with the client device 110, the voice bot client 113 and the voice bot system 120 form a logical instance of a voice bot. Although the voice bot system 120 is described in Figure 1 the background as being implemented remotely from the client device 110 (e.g., via one or more servers), it should be appreciated that this is for purposes of example, and not limitation. For example, one or more aspects of the voice bot development 120 can alternatively be implemented locally at the client device 110 and / or at one or more additional client devices 195 over one or more networks 1992.

[0027] A developer (e.g., a user of the client device 110) can interact with the voice bot system 120 (e.g., via the client device 110) to train a plurality of machine learning (ML) layers of one or more ML models stored in the ML layer database 170A1. The plurality of ML layers can correspond to those of a transformer ML model (e.g., an input layer, an encoding layer, a decoding layer, a feed-forward layer, an attention layer, an output layer, and / or other ML layers), a unidirectional and / or bidirectional RNN model (e.g., an input layer, a hidden layer, an output layer, and / or other ML layers), and / or other ML layers of other ML models. Further, the plurality of ML layers can be subsequently used by the voice bot while in conversation with a corresponding human to determine a corresponding unique personal identifier provided by the corresponding human during the conversation. The corresponding unique personal identifier can include any sequence of alphanumeric characters that is personal to a given human, and can be, for example, one or more of an email address, a physical address, a username, a password, a product identifier, a name or domain name of an entity.

[0028] In some implementations, a single instance of multiple ML layers can be utilized to resolve a unique person identifier. For example, a single instance of multiple ML layers can be trained and utilized to resolve email addresses, physical addresses, and the like. In some versions of those implementations, the type of unique person identifier (e.g., email address, physical address, username, and the like) can optionally be processed along with one or more ASR speech hypotheses as described herein to resolve the unique person identifier. In additional or alternative implementations, multiple instances of multiple ML layers can be trained and utilized to resolve unique person identifiers. For example, a first plurality of ML layers can be trained and utilized to resolve email addresses, a second plurality of ML layers can be trained and utilized to resolve physical addresses, a third plurality of ML layers can be trained and utilized to resolve usernames, and the like. In these implementations, different pluralities of ML layers can be trained based only on one or more training instances corresponding to the type of unique person identifier that the plurality of ML layers is trained to resolve, and the type of unique person identifier encountered during a corresponding conversation can be used to select the appropriate plurality of ML layers to process one or more ASR speech hypotheses.

[0029] The voice bot can correspond to one or more processors utilizing a plurality of additional ML layers of one or more ML models stored in the voice bot database 170A2, and can be a first party voice bot or a third party voice bot. As used herein, the term first party refers to an entity that publishes the voice bot system, while the term third party refers to an entity that is different from the entity associated with the first party and that does not publish the voice bot system. The developer that trains the plurality of ML layers can be a first party developer associated with the first party entity or a third party developer associated with the third party entity.

[0030] In some implementations, the voice bot can be an example-based voice bot. For example, an example-based voice bot can be trained based on a plurality of training instances obtained based on a corresponding conversation. The corresponding conversation can be, for example, an illustrative conversation defined by a developer for the purpose of training the voice bot and / or a previously-inferred conversation, which can or can not include an instance of the voice bot as a participant in the previously-inferred conversation. In additional or alternative implementations, the voice bot can be a rule-based voice bot. For example, a rule-based voice bot can be associated with one or more intent patterns defined at least in part by a developer. The intent patterns can define, for example, one or more responses that the voice bot should provide in response to determining an intent of a human included in spoken utterances of a human participating in a corresponding conversation with the voice bot. In contrast to a rule-based voice, an example-based voice bot can be trained to process audio data capturing spoken utterances (or speech hypotheses corresponding thereto) to directly predict one or more responses that the voice bot should provide without directly determining an intent of the human.

[0031] A corresponding conversation in which a voice bot described herein can utilize a plurality of ML layers can include any conversation between the voice bot and a human associated with the client device 110 and / or one or more additional client devices 195. In some implementations, a given corresponding conversation can occur during a corresponding telephone call in which the voice bot participates in the corresponding conversation with a human. The corresponding telephone call described herein can be performed using various voice communication protocols (e.g., Voice over Internet Protocol (VoIP), Public Switched Telephone Network (PSTN), and / or other telephone communication protocols). As described herein, synthetic speech is rendered by the voice bot as part of the corresponding telephone call, which can include injecting the synthetic speech into the corresponding telephone call such that the human participating in the corresponding conversation can perceive the synthetic speech. The synthetic speech can be generated by the client device and / or injected into one of the endpoints of the corresponding telephone call (e.g., the voice bot system 120, the client device 110, and / or one additional client device 195). In additional or alternative implementations, a given corresponding conversation can occur during a corresponding conversation session in which the human associated with the client device 110 and / or one additional client device 195 invokes the voice bot to perform an action on behalf of the human. Corresponding conversations are described below (e.g., with reference to Figure 6A to 6C ).

[0032] In various implementations, the voice bot system 120 includes an ASR engine 130A2, an NLU engine 140A2, a TTS engine 150A2, an ML training engine 160, and a voice bot engine 170. The ML training engine 160 can be used to train a plurality of ML layers that are later used by the voice bot during a corresponding conversation to determine a corresponding unique person identifier, and in various implementations can include a training instance engine 161, a training engine 162, and a simulation engine 163. Further, the voice bot engine 170 can later utilize the voice bot to conduct the corresponding conversation, and in various implementations can include a response engine 171 and a unique person identifier engine 172.

[0033] In some implementations, the ASR engine 130A1 of the client device 110 (or of a further client device 195) can use the ASR model 130A to process audio data that captures a spoken utterance. In further or alternative implementations, the client device 110 can transmit the audio data to the voice bot system 120 over the networks 1991 and / or 1992, and the ASR engine 130A2 can use the ASR model 130A to process the audio data that captures the spoken utterance. The ASR engine 130A1 and / or 130A2 can generate a plurality of ASR speech hypotheses for the spoken utterance based on the processing of the audio data, and can optionally select a particular speech hypothesis as recognized text for the spoken utterance based on corresponding values (e.g., binary values, probability values, log-likelihood values, and / or other values) associated with the plurality of ASR speech hypotheses. In various implementations, the ASR model 130A is an end-to-end speech recognition model, such that the ASR engine 130A1 and / or 130A2 can directly use the model to generate the plurality of ASR speech hypotheses. For example, the ASR model 130A can be an end-to-end model for generating each of the plurality of ASR speech hypotheses on a character-by-character basis (or other token-by-token basis). Some non-limiting examples of such end-to-end models for generating recognized text on a character-by-character basis are recurrent neural network transducer (RNN-T) models and transformer models. RNN-T models are in the form of sequence-to-sequence models that do not employ an attention mechanism, while transformer models do employ an attention mechanism. In other implementations, the ASR model 130A is not an end-to-end speech recognition model, such that the ASR engine 130A1 and / or 130A2 can instead generate predicted phonemes (and / or other representations). For example, the predicted phonemes (and / or other representations) can then be used by the ASR engine 130A1 and / or 130A2 to determine a plurality of ASR speech hypotheses that conform to the predicted phonemes. In doing so, the ASR engine 130A1 and / or 130A2 can optionally employ a decoding graph, a lexicon, and / or other resources. In various implementations, any detected transcriptions of spoken utterances can be visually rendered via a display of the client device 110 (or of a further client device 195) using the rendering engine 112.

[0034] In some versions of those implementations, the NLU engine 140A1 of the client device 110 and / or the NLU engine 140A2 of the voice bot system 120 can use the NLU model 140A to process the recognized text generated by the ASR engine 130A1 and / or 130A2 to determine an intent included in the spoken utterance. For example, if the client device 110 detects the spoken utterance “tell Branden to send the prize money to my quick cash account and I’ll see him later tonight,” the client device 110 and / or the voice bot system 120 can use the ASR model 130A1 and / or 130A2 to process the audio data capturing the spoken utterance to generate recognized text corresponding to the spoken input, and can use the NLU model 140A to process the recognized text to determine, at least, an intent to add a message with message content parameters having values of “send the prize money to my quick cash account and I’ll you later tonight.” In some versions of those implementations, the TTS engine 150A1 of the client device 110 and / or the TTS engine 150A2 of the voice bot system 120 can generate synthesized speech audio data capturing the synthesized speech. The synthesized speech can be audibly rendered via one or more speakers of the client device 110 (or one other client device 195) using the rendering engine 112. The synthesized speech can capture any output generated by the voice bot or voice bot system 120 described herein.

[0035] The training instance engine 161 can obtain a plurality of training instances for training the plurality of ML layers based on user input provided by a developer and detected at the client device 110 via the user input engine 111. The plurality of training instances can be stored in a training instance database 161A. Each of the plurality of training instances can include a training instance input and a training instance output. The training instance input can include one or more of: at least one ASR speech hypothesis for a unique person identifier, audio data for which the at least one ASR speech hypothesis was generated, or an intent of a voice bot associated with synthesized speech to which the audio data was responsive. The training instance output can include a ground truth output corresponding to the unique person identifier.

[0036] In some implementations, one or more of the plurality of training instances can be obtained based on a corresponding prior conducted conversation that includes a corresponding unique person identifier for a corresponding human (e.g., as described with respect to method 300A). Figure 3 In implementations in which a given prior conducted conversation is between at least the corresponding human and another human, one or more of the plurality of training instances can be generated based on one or more portions of the given prior conducted conversation.

[0037] For example, training instance engine 161 can identify a portion of a given prior conducted conversation that includes a unique person identifier and utilize the given ASR speech hypothesis that includes the unique person identifier as a training instance input. Training instance engine 161 can identify the portion of the given prior conducted conversation that includes the unique person identifier by causing ASR engine 130A1 and / or 130A2 to process a corresponding portion of audio data for the given prior conducted conversation using ASR model 130A to generate a plurality of ASR speech hypotheses. Further, training instance engine 161 can analyze one or more of the plurality of ASR speech hypotheses for the given prior conducted conversation to identify the portion that includes the unique person identifier. For example, assume that the given ASR speech hypothesis includes tokens that correspond to “my email address is…,” “my home address is…,” “my username is…,” etc. In this example, training instance engine 161 can identify the given ASR speech hypothesis as including a unique person identifier (e.g., where the ellipses indicate a unique person identifier that follows). As another example, assume that the given ASR speech hypothesis includes tokens that correspond to “what is your email address?” “can you please provide your home address?” “what’s the username?” etc. In this example, training instance engine 161 can identify the given ASR speech hypothesis for subsequent audio data that responds to the question as including a unique person identifier. In these examples, the given ASR speech hypothesis that includes at least the ASR speech hypothesis for the unique person identifier can be utilized as a training instance input for a given training instance.

[0038] Further, the training instance engine 161 can identify a ground truth output for a given training instance based on one or more supervisory signals associated with a given prior conducted conversation. The one or more supervisory signals can include, for example, entry of a unique person identifier into the system (and optionally successful finding of a match to the unique person identifier in the system) during the given prior conducted conversation, manual entry of the unique person identifier into the given prior conducted conversation by at least one human (e.g., via typed input), and / or editing of a predicted unique person identifier by at least one human (e.g., via touch input or typed input) to the unique person identifier during the given prior conducted conversation, subsequent correction of the unique person identifier by a human reviewer after the given prior conducted conversation, and / or any other supervisory signal. The training instance input and the training instance output for the given training instance can be stored in the training instance database 161 A and subsequently used to train the plurality of ML layers (e.g., as described with respect to the training engine 162).

[0039] In implementations in which the given prior conducted conversation is between instances of at least a corresponding human and a voice bot, one or more of the plurality of training instances can be generated based on one or more portions of the given prior conducted conversation. Notably, in these implementations, a plurality of ASR speech hypotheses can have been generated by the instance of the voice bot during the given prior conducted conversation. Accordingly, the training instance engine 161 can analyze the plurality of ASR speech hypotheses in the same or similar manner as described above, and a given ASR speech hypothesis (which includes at least an ASR speech hypothesis for a unique person identifier) can be used as the training instance input for a given training instance. Further, a ground truth output corresponding to the unique person identifier can be identified based on one or more of the supervisory signals described above, and can be used as the training instance output for the given training instance. The training instance input and the training instance output for the given training instance can be stored in the training instance database 161 A and subsequently used to train the plurality of ML layers (e.g., as described with respect to the training engine 162).

[0040] In further or alternative implementations, one or more of the plurality of training instances can be obtained based on unique person identifiers stored in one or more databases (e.g., as described with respect to the training engine 162). For example, the training instance engine 161 can identify a plurality of unique person identifiers stored in one or more databases, and can generate a plurality of training instances based on the plurality of unique person identifiers. In these implementations, the training instance engine 161 can generate a training instance input for each of the plurality of training instances based on the plurality of unique person identifiers. Further, the training instance engine 161 can identify a ground truth output for each of the plurality of training instances based on one or more supervisory signals associated with a given prior conducted conversation. The one or more supervisory signals can include, for example, entry of a unique person identifier into the system (and optionally successful finding of a match to the unique person identifier in the system) during the given prior conducted conversation, manual entry of the unique person identifier into the given prior conducted conversation by at least one human (e.g., via typed input), and / or editing of a predicted unique person identifier by at least one human (e.g., via touch input or typed input) to the unique person identifier during the given prior conducted conversation, subsequent correction of the unique person identifier by a human reviewer after the given prior conducted conversation, and / or any other supervisory signal. The training instance input and the training instance output for the given training instance can be stored in the training instance database 161 A and subsequently used to train the plurality of ML layers (e.g., as described with respect to the training engine 162). Figure 3The training instance engine 161 can generate a plurality of tokens based on the given unique personal identifier (e.g., based on the n-grams of the given unique personal identifier). For example, assume that the given unique personal identifier is the email address “tatortator13@exampleurl.com”. In this example, the email address can be composed of at least the tokens “tator”, “tator”, “13”, and “exampleurl.com”. The training instance engine 161 can replace one or more of the tokens based on one or more distributions of n-grams to modify the given unique personal identifier. For example, the one or more distributions can indicate that, for unique personal identifiers, tokens of “james”, “john”, “robert”, “susan”, and “karen” are more common than “tator” and, thus, more likely to be encountered when a voice bot utilizes the plurality of ML layers. Accordingly, the training instance engine 161 can generate a modified unique personal identifier that includes at least the tokens “john”, “tator”, “13”, and “exampleurl.com”.

[0041] In additional or alternative versions of those implementations, the training instance engine 161 can synthesize one or more training instances based on the given unique personal identifier retrieved from the one or more databases (e.g., to preserve the privacy of the human associated with the given unique personal identifier). The training instance engine 161 can generate a plurality of tokens based on the given unique personal identifier based on one or more distributions of n-grams. For example, assume that the given unique personal identifier retrieved from the one or more databases is the email address “tatortator13@exampleurl.com”. In this example, the email address can be composed of at least the tokens “tator”, “tator”, “13”, and “exampleurl.com”. The training instance engine 161 can replace one or more of the tokens based on one or more distributions of n-grams to modify the given unique personal identifier. For example, the one or more distributions can indicate that, for unique personal identifiers, tokens of “james”, “john”, “robert”, “susan”, and “karen” are more common than “tator” and, thus, more likely to be encountered when a voice bot utilizes the plurality of ML layers. Accordingly, the training instance engine 161 can generate a modified unique personal identifier that includes at least the tokens “john”, “tator”, “13”, and “exampleurl.com”.

[0042] While only the token "tator" is depicted as being replaced relative to the modified unique person identifier, it should be understood that this is for the purposes of example and not limitation, and that the training instance engine 161 can replace any number of tokens for a given unique person identifier. Moreover, one or more of the n-grams can be replaced based on respective distributions to generate a plurality of tokens corresponding to the unique person identifier. For example, one or more of the n-grams can be replaced based on a distribution of first names, a distribution of last names, a distribution of nicknames, and / or a distribution of usernames, a distribution of non-name characters (e.g., special characters or symbols), a distribution of numbers, and / or any other distribution. For example, where the unique person identifier corresponds to a street address, a less popular street name (e.g., "Keewood Court") can be replaced with a more popular street name (e.g., "Main Street"). Thus, the tokens generated when synthesizing one or more training instances can more accurately reflect the rate or distribution of tokens encountered when the plurality of ML layers trained on the plurality of training instances are subsequently utilized by a voice bot.

[0043] Additionally, the training instance engine 161 can process the plurality of tokens to generate the synthetic text based on the plurality of tokens for the modified unique personal identifier. In some implementations, the training instance engine 161 can inject one or more fillers into the plurality of tokens. Continuing the example above, the system can prepend one or more fillers to one or more of the plurality of tokens or append one or more fillers to one or more of the plurality of tokens. For example, the training instance engine 161 can pre-pend the filler “oh sure it’s uhh” to the email address, resulting in the synthetic text “oh sure it’s uhh johntatorl 3@exampleurl.com,” or append the filler “then” or “as in” between one or more of the tokens, resulting in the synthetic text “john then tator then 13 as in thenumber and@exampleurl.com,” and so on. In some additional or alternative implementations, the training instance engine 161 can inject a phonetic spelling for one or more alphanumeric characters included in a given unique personal identifier. Continuing the example above, the system can inject the phonetic spelling “t as in tango” for “tator,” resulting in the synthetic text “john then tator then 13 as in the number and@exampleurl.com, that’s t as in tango.” The system can inject a phonetic spelling for one or more alphanumeric characters based on a probability that the alphanumeric character can be spelled out. For example, the alphanumeric characters “t” and “d” are often confused in speech recognition, so the probability that a human will provide the phonetic spelling “t as in tango” or “d as in dog” can be greater than the probability that a human will provide the phonetic spelling “z as in zulu.”In these examples, the resulting synthesized text can include "my email is uhh john and then t as in tango a t o r 13 and then@exampleurl.com." By injecting fillers and / or voice spellings into the plurality of tokens to generate the synthesized text, the system can learn to ignore fillers and utilize voice spellings when interpreting the modified unique personal identifier.

[0044] Further, the training instance engine 161 can process the synthesized text to generate a given ASR phonetic hypothesis based on the synthesized text by injecting one or more ASR errors into the given ASR phonetic hypothesis. In some implementations, the one or more ASR errors injected into the given ASR phonetic hypothesis can include replacing one or more tokens included in the synthesized text with one or more corresponding homophonic tokens. Continuing the above example, the token "13" in the email address can be replaced with "thirdteen," the token "tator" can be replaced with "date her" or "tate her," and / or other tokens can be replaced with corresponding homophonic tokens. In some additional or alternative implementations, the one or more ASR errors injected into the given ASR phonetic hypothesis can include replacing one or more alphanumeric characters of the synthesized text with one or more corresponding homophonic alphanumeric characters. Continuing the above example, the alphanumeric character "t" of the token "tator" can be replaced with "d," resulting in the token "dator," and / or other alphanumeric characters can be replaced with corresponding homophonic alphanumeric characters. By injecting ASR errors into the given ASR phonetic hypothesis, the plurality of ML layers can be subsequently trained to handle these ASR errors that can be encountered when the plurality of ML layers are deployed for use by a voice bot. The given ASR phonetic hypothesis can be used as a training instance input for a given training instance, and the generated tokens (e.g., generated based on the given unique personal identifier and prior to generating the synthesized text) can be used as a ground truth output for the given training instance. The given training instance can be stored in the training instance database 161 A, and can be subsequently used to train the plurality of ML layers (e.g., as described with respect to the training engine 162).

[0045] The training engine 162 can train the plurality of ML layers stored in the ML layer database 170A1 using the plurality of training instances obtained by the training instance engine 161 (e.g., stored in the training instance database 161A). The plurality of ML layers can correspond to those of a transformer ML model (e.g., input layers, encoding layers, decoding layers, feed-forward layers, attention layers, output layers, and / or other ML layers), a unidirectional and / or bidirectional RNN model (e.g., input layers, hidden layers, output layers, and / or other ML layers), and / or other ML layers of other ML models.

[0046] For example, with specific reference to Figure 2A An example process flow 200A for training a plurality of ML layers utilized by a voice bot in determining a unique person identifier is depicted. In some implementations, the training instance engine 161 can obtain a given training instance from among the plurality of training instances stored in the training instance database 161A. In some implementations, the training instance input for the given training instance can include one or more ASR speech hypotheses 202A corresponding to a unique person identifier or a respective modified unique person identifier (referred to simply as a unique person identifier). In some versions of those implementations, the one or more ASR speech hypotheses 202A can be stored in the training instance database 161A prior to training the plurality of ML layers (e.g., as described above with respect to the training engine 161). In additional or alternative versions of those implementations, the one or more ASR speech hypotheses 202A can be generated during training. For example, audio data 201A capturing a spoken utterance including at least a unique person identifier can be used as the training instance input. The training instance engine can cause the ASR engine 130A1 and / or 130A3 to process the audio data 201A using the ASR model 130A to generate one or more of the ASR speech hypotheses 202A. Further, the training instance output for the given training instance can include a ground truth output 203A corresponding to the unique person identifier corresponding to the one or more speech hypotheses 202A.

[0047] The training engine 162 can cause the unique person identifier engine 172 to process the one or more ASR speech hypotheses 202A using the plurality of ML layers stored in the ML layer database 170A1 to generate corresponding predicted measures 204A associated with one or more alphanumeric characters of a unique person identifier corresponding to the one or more ASR speech hypotheses 202A. The corresponding predicted measures 204A can include binary values, probabilities, log-likelihoods, or other measures corresponding to likelihoods of one or more alphanumeric characters corresponding to actual alphanumeric characters of the unique person identifier. Further, the training engine 162 can cause the loss engine 162a1 to compare the corresponding predicted measures 204A associated with the one or more alphanumeric characters of the unique person identifier to ground truth measures corresponding to each alphanumeric character of the unique person identifier corresponding to the one or more ASR speech hypotheses 202A included in the ground truth output 203A. The loss engine 162a1 can generate one or more losses 205A based on the comparison. Further, the training engine 162 can cause the update engine 162A2 to update the plurality of ML layers stored in the ML layer database 170A1 based on the one or more losses 205A.

[0048] For example, assume that the unique person identifier for a given training instance is the physical address “1234 Keewood Court.” Further assume that one or more candidate unique person identifiers are generated for the “Keewood” portion of the physical address, and based on corresponding predicted measures 204A for each of the alphanumeric characters included in “Key wood,” “Leewood,” “Key wraw,” “Qi wood,” and / or other candidate unique person identifiers. Each alphanumeric character in each candidate unique person identifier can be associated with a given one of the corresponding predicted measures indicating a likelihood of the alphanumeric character(s) corresponding to the unique person identifier. In this example, each alphanumeric character of the “Keewood” portion of the physical address and / or the corresponding ground truth measures associated with each alphanumeric character of the “Keewood” portion of the physical address can be compared to each of the corresponding alphanumeric characters and / or corresponding predicted measures 204A of the one or more candidate unique person identifiers on a character-by-character basis. Further, one or more losses can be generated based on comparing the alphanumeric characters (e.g., cross-entropy loss), and these losses can be utilized to update the plurality of ML layers. For example, the one or more losses can be backpropagated across one or more of the plurality of ML layers to update respective weights of one or more of the plurality of ML layers.

[0049] The training engine 162 can continue training the plurality of ML layers in this manner based on one or more additional training instances stored in the training instance database 161A. In some implementations, the training engine 162 continues training the plurality of ML layers in this manner until one or more conditions are met. The one or more conditions can include, for example, validation of the updated plurality of ML layers, convergence of the updated plurality of ML layers (e.g., zero loss or within a threshold range of zero loss), a determination that the plurality of ML layers performs better (e.g., in terms of precision and / or traceability) than an instance of the plurality of ML layers currently utilized, if any, by the voice bot. Based on occurrence of at least a threshold amount of training of the plurality of training instances, and / or based on a duration of training of the plurality of training instances. Notably, by training the plurality of ML layers in this manner, the plurality of ML layers can process the one or more ASR speech hypotheses 202A and generate a predicted unique person identifier. However, even with unlimited training and training resources, the plurality of ML layers can not accurately predict every unique person identifier that can be encountered from this initial round of the corresponding conversation from which the one or more candidate unique person identifiers are generated.

[0050] Accordingly, in various implementations, the simulation engine 163 can configure the simulated environment to simulate additional rounds of the corresponding conversation utilizing the simulator 163A. In some implementations, and as discussed above with respect to the simulation engine 163, the simulator 163A can include a simulated voice bot portion 163A1 (e.g., an instance of the voice bot utilizing the plurality of ML layers) and a simulated human portion 163A2 (e.g., another instance of the voice bot utilizing the plurality of ML layers). In additional or alternative implementations, the developer can replace the simulated voice bot portion 163A1 and / or the simulated human portion 163A2 of the simulator 163A and can provide input via one or more user interface input devices of the client device 110 during a given simulation. Moreover, the simulated voice bot portion 163A1 and the simulated human portion 163A2 can have access to one or more databases including candidate speech that can be utilized in generating prompts and / or responses described herein. Figure 2A

[0051] ​For example, assume that the unique person identifier for a given training instance is the physical address "1234 Keewood Court." Further assume that the simulation engine 163 causes the unique person identifier engine 172 to use a plurality of ML layers to process one or more ASR speech hypotheses corresponding to the physical address to generate one or more of the candidate unique person identifiers for at least the "Keewood" portion of the physical address and based on corresponding prediction measures generated. Further assume that the unique person identifier engine 172 selects one or more alphanumeric characters 206A included in one or more of the candidate unique person identifiers generated based on corresponding prediction measures for one or more alphanumeric characters predicted to correspond to the physical address. Further assume that the one or more alphanumeric characters 206A correspond to "K e y" for the "Kee" portion of "Keewood" for the unique person identifier. The unique person identifier engine 172 can determine an intent that the voice bot should utilize when subsequently generating one or more simulated prompts 207A based on the corresponding prediction measures associated with the one or more given alphanumeric characters 206A. For example, the unique person identifier engine 172 can determine an intent associated with requesting clarification regarding the one or more given alphanumeric characters 206A, requesting the person to spell the unique person identifier, and / or any other intent described herein.

[0052] Further, the simulation engine 163 can use the simulated voice bot portion 163A1 of the simulator 163A to process the one or more alphanumeric characters 206A (and optionally the intent) to generate one or more simulated prompts 207A that include a corresponding clarification request requesting clarification regarding the one or more alphanumeric characters of the unique person identifier. In some implementations, the one or more simulated prompts 207A can be processed by the TTS engine 150A1 and / or 150A2 using the TTS model 150A to generate synthetic speech audio data that can be rendered for presentation to the developer via one or more speakers of the client device 110. For example, assume that the simulated voice bot portion 163A1 generates the first prompt "does that start with k as in kilo" based on a corresponding prediction measure associated with "k" indicating that the voice bot is not highly confident in the selected alphanumeric character "k."

[0053] Moreover, the simulation engine 163 can use the simulated human portion 163A2 of the simulator 163 A to process one or more of the simulation prompts 207A to generate one or more simulated responses 208A. The simulated human portion 163A2 of the simulator 163A can utilize the ground truth output corresponding to the unique identifier in generating the one or more simulated responses 208A. In some implementations, the one or more simulated responses 208A can correspond to additional synthetic speech audio data that captures the one or more simulated responses 208A generated using the TTS engine 150A1 and / or 150A2. Continuing the example above, assume that the first response “yes, k as in kilo and then e e” is provided in response to the first prompt. Further, the simulation engine 162 can cause the unique person identifier engine 172 to use the plurality of ML layers to process one or more speech hypotheses corresponding to the first response to refine the one or more alphanumeric characters 206A. For example, the unique person identifier engine 172 can update the corresponding prediction measure to indicate that the alphanumeric characters of “k e e” are very likely to be correct, resulting in a refined unique person identifier of “Kee wood” for the “Keewood” portion of the unique person identifier.

[0054] The simulation engine 163 can cause the simulated voice bot portion 163A1 and the simulated human portion 163A2 to perform n additional rounds of simulated conversation (e.g., where n is a positive integer) until the voice bot until one or more alphanumeric characters 206A correspond to a ground truth output for a unique person identifier. For example, the next prompt can be “is Kee wood one word or two words?” and the next response can be “yes,” resulting in the refined unique person identifier “Keewood.” Further, the next prompt can be “so its K e e w o o d, all one word?” and the next response can be “yes,” resulting in the refined unique person identifier “Keewood,” thereby resolving the unique person identifier for the training instance. In some implementations, one or more alphanumeric characters 206A can be considered to correspond to a ground truth output for a unique person identifier when the corresponding predicted measure for each alphanumeric character satisfies a threshold. In further or alternative implementations, such as the example above, one or more alphanumeric characters 206A can be considered to correspond to a ground truth output for a unique person identifier when a human (or simulated human portion 163A2) verifies that the predicted unique person identifier is the actual predicted unique person identifier. In some implementations, the simulation engine 163 continues to train the plurality of ML layers in this manner until one or more conditions are satisfied (e.g., as described above).

[0055] In various implementations, the unique person identifier engine 172 can use multiple ML layers to process the intent of the voice bot along with other inputs (e.g., one or more ASR speech hypotheses 202A and / or one or more simulated responses 208A). For example, the intent associated with the voice bot for an initial round to predict a unique person identifier (e.g., associated with one or more of the ASR speech hypotheses 202A) can include an intent to request to provide a unique person identifier, an intent to spell the unique person identifier, and / or other intents. As another example, the intent associated with the voice bot for a subsequent round to refine the unique person identifier (e.g., associated with one or more simulated responses 208A) can include an intent to request to verify one or more alphanumeric characters, an intent to spell the unique person identifier, an intent to verify the unique person identifier, and / or other intents. By processing the intent along with other inputs, the unique person identifier engine 172, i.e., the voice bot, can better process the inputs processed by the multiple ML layers, which allows for more rapid and efficient identification and refinement of the one or more alphanumeric characters 206A while maintaining the same level of accuracy and traceability.

[0056] In additional or alternative implementations, the unique person identifier engine 172 can use multiple ML layers to process a type of unique person identifier along with other inputs (e.g., one or more ASR speech hypotheses 202A and / or one or more simulated responses 208A). The type of unique person identifier can correspond to whether it is an email address, a physical address, a username, etc. In additional or alternative implementations, the type of unique person identifier can be utilized by the unique person identifier engine 172 after processing at least one or more of the ASR speech hypotheses. By processing the type of unique person identifier along with another input, the voice bot can better process the inputs processed by the multiple ML layers, which allows for more rapid and efficient identification and refinement of the one or more alphanumeric characters 206A while maintaining the same level of accuracy and traceability. Although the multiple ML layers are described as being trained in a particular manner and using a particular architecture, it should be appreciated that this is for purposes of example and is not meant to be limiting.

[0057] After training the voice bot, the voice bot engine 170 can then utilize the trained multiple ML layers stored in the ML layer database 170A1 and the voice bot stored in the voice bot database 170A2 to determine a unique person identifier while the voice bot is in the middle of a conversation. For example, with specific reference to Figure 2BFig. 2B depicts an example process flow 200B that illustrates the utilization of multiple ML layers by a voice bot to determine a unique person identifier. For purposes of example, assume that the voice bot is associated with a supermarket self-checkout system that is implemented at one of the additional client devices 195 that is in communication with the voice bot system 120. Further assume that a human has completed checkout at the supermarket self-checkout system and that the voice bot has prompted the human to provide an email address to which the supermarket self-checkout system can send receipts and / or other content (e.g., coupons, advertisements, and / or other promotional materials). In this example, audio data 201B that captures a spoken utterance that includes the email address can be processed by the ASR engine 130A1 and / or 130A2 to generate one or more speech hypotheses 202B.

[0058] In some implementations, at block 203B, the voice bot can determine whether one or more of the speech hypotheses 202B corresponds to a unique person identifier. In some versions of those implementations, the voice bot can determine whether one or more of the speech hypotheses 202B corresponds to a unique person identifier based on one or more of the speech hypotheses including a token for one or more alphanumeric characters and / or one or more strings of tokens that indicate alphanumeric characters of a unique person identifier. For example, if one or more of the speech hypotheses 202B includes “dot com,” it can be determined that it includes an email address, if one or more of the speech hypotheses 202B includes a city and a state, it can be determined that it includes a physical address, and any other linguistic cues for various unique person identifiers ad infinitum. In some additional or alternative implementations, the voice bot can determine whether one or more of the speech hypotheses 202B is generated based on a spoken utterance that is received in response to soliciting a unique person identifier. If the voice bot determines at block 203B that the audio data 201B does not include a unique person identifier, the voice bot can continue to analyze additional audio to determine whether it includes a unique person identifier. If the voice bot determines at block 203B that the audio data 201B includes a unique person identifier, the voice bot can cause the unique person identifier engine 172 to process one or more of the ASR speech hypotheses 202B using the multiple ML layers stored in the ML layer database 170A1.

[0059] In some implementations, the likelihood tree 172A can be generated based on the outputs generated across the plurality of ML layers. The likelihood tree can include a plurality of nodes 204B, 205B, 206B1, 206B1, 207B, 208B1, 208B2, and / or other nodes, and a plurality of edges connecting one or more nodes. Each of the plurality of nodes can be associated with an alphanumeric character and a corresponding prediction measure associated with the alphanumeric character. Further, one or more candidate unique person identifiers can be generated based on the likelihood tree 172, and a given candidate unique person identifier can be selected from among the one or more candidate unique person identifiers based on the corresponding prediction measure. In additional or alternative implementations, the unique person identifier engine 172 can directly generate one or more candidate unique person identifiers without utilizing the likelihood tree 172A.

[0060] Continuing the above example, assume that the email address captured in the audio data 201B is "johnp@exampleurl.com." In this example, and focusing on the "johnp" portion of the email address, the likelihood tree 172A can include a first node 204B for the alphanumeric character "j" associated with a corresponding prediction measure of 0.87, a second node 205B for the alphanumeric character "o" associated with a corresponding prediction measure of 0.90, a third node 206B1 for the alphanumeric character "h" associated with a corresponding prediction measure of 0.5, a third alternative node 206B2 for the alphanumeric character "[null]" associated with a corresponding prediction measure of 0.5, a fourth node 207B for the alphanumeric character "n" associated with a corresponding prediction measure of 0.95, a fifth node 208B1 for the alphanumeric character "p" associated with a corresponding prediction measure of 0.6, and a fifth alternative node 208B2 for the alphanumeric character "d" associated with a corresponding prediction measure of 0.4. In this example, the unique person identifier engine can select the given unique person identifier "jonp" instead of the user-desired "johnp." Although Figure 2BEach of the plurality of nodes shown in the middle corresponds to a single alphanumeric character, although it should be appreciated that this is for purposes of example and is not limiting. For example, assume that the corresponding prediction measures for two or more adjacent alphanumeric characters indicate that the corresponding to a single person identifier. In this example, the nodes for the two or more adjacent alphanumeric characters can be combined into a single node. As yet another example, assume that the email address captured in the audio data 201B is “johnp25@exampleurl.com”. In this example, the numeric portion of the email address, “25”, can be represented as one or more of the following nodes: “2”, “5”, “25”, “twenty”, “five”, “twenty-five”, “two”, “five”, and / or any other combination.

[0061] The response engine 171 can process the given unique person identifier selected by the unique person identifier engine 172 based on the likelihood tree 172A. For example, even though the unique person identifier selected the alphanumeric character “[null]” at the alphanumeric character “h”, the response engine 171 can determine the corresponding prediction measures of the nodes in the likelihood tree within the threshold range (e.g., the third node 206B1 and the alternative third node 206B2) and generate the prompt 209B “is that john with an h or jon without an h”. Further, the TTS engine 150A1 and / or 150A2 can process the prompt 209B using the TTS model 150A to generate the synthesized speech audio data 210B including the prompt 209B.

[0062] Continuing the above example, assume the human provides an additional spoken utterance indicating that the email address includes "john with an h." In this example, additional audio data captures the additional spoken utterance in the same or similar manner as described above, and the unique person identifier engine 172 can process one or more additional ASR speech hypotheses corresponding to "john with an h." Based on the additional spoken utterance, the unique person identifier engine 172 can update the likelihood tree 172A. In this example, the likelihood tree 172A can be updated to remove the third replacement node 206B2 and limit any candidate unique identifiers to those that include "john with an h." The response engine 171 can generate an additional prompt to determine whether the email address includes a "p" or a "d" in the same or similar manner. This process can be repeated until a unique person identifier for the human is determined. After a unique person identifier for the human is determined in the email, the receipt and / or other content can be transmitted to the email address of the human, and a notification that the receipt was successfully sent can be rendered for presentation to the human at the self-checkout system.

[0063] In some implementations, the plurality of training instances used to train the plurality of ML layers can be obtained based on actual conversations and / or synthetically generated to reflect actual distributions of unique person identifiers. This allows the plurality of ML layers to obtain a high level of accuracy and / or traceability for unique person identifiers in actual use when utilized by a voice bot. Moreover, by obtaining a high level of accuracy and / or traceability for unique person identifiers, corresponding conversations that include unique person identifiers can be ended more quickly and efficiently because the plurality of ML layers, which are trained by a voice bot and using the techniques described herein, are more capable of understanding subtle differences in human speech and responding accordingly to resolve unique person identifiers. Furthermore, voice bots that utilize the plurality of ML layers described herein are more scalable, and memory consumption is reduced because the plurality of ML layers can be shared among a plurality of different voice bots. For example, a plurality of third parties can develop respective voice bots for particular tasks without having to train the respective voice bots to determine unique person identifiers. Instead, the respective voice bots can each simply use the plurality of ML layers (or respective instances thereof).

[0064] Turning now to Figure 3 , flowcharts illustrating example methods 300A and 300B of obtaining training instances for training the plurality of ML layers utilized by a voice bot in determining unique person identifiers are depicted. For convenience, the operations of methods 300A and 300B are described with reference to a system that performs the operations. This system of methods 300A and 300B includes a computing device (e.g., client device 110 of Figure 1 ,Figure 1 the voice bot system 120, Figure 7 of the computing device 710, server, and / or other computing device). Also, while the operations of the methods 300A and 300B are shown in a particular order, this is not meant to be limiting. One or more operations can be reordered, omitted, and / or added.

[0065] In some implementations, the one or more training instances can be obtained based on processing a previously conducted conversation that includes at least a human participant. For example, with particular reference to Figure 3 At block 352A of the method 300A, the system obtains a corresponding previously conducted conversation that includes a unique personal identifier for the at least one human participant. The unique personal identifier can be included in a spoken utterance provided by the at least one human. In some implementations, the corresponding previously conducted conversation can be between multiple humans (e.g., the at least one human and another human). For example, the corresponding previously conducted conversation can be conducted during a phone call in which the multiple humans are located in different environments, during a conversation in which the multiple humans are co-located in the same environment, and / or during other conversations between the multiple humans.

[0066] In additional or alternative implementations, the corresponding previously conducted conversation can be between the at least one human and an instance of the voice bot. For example, the corresponding previously conducted conversation can be conducted during a phone call between the at least one human and the instance of the voice bot (e.g., in cases in which the instance of the voice bot is implemented by a remote computing device), during a conversation session between the at least one human and the instance of the voice bot (e.g., in cases in which the instance of the voice bot is implemented by a local computing device), and / or during other conversations between the at least one human and the instance of the voice bot.

[0067] At block 354A, the system obtains one or more ASR speech hypotheses for the unique personal identifier provided by the at least one human participant. In implementations in which the corresponding previously conducted conversation is between multiple humans, the system can use an ASR model to process spoken utterances provided by the at least one human during the corresponding previously conducted conversation that include the unique personal identifier to generate one or more ASR speech hypotheses for the unique personal identifier. For example, the spoken utterances provided by the at least one human during the corresponding previously conducted conversation that include the unique personal identifier for the at least one human can be identified based on one or more items of the spoken utterances provided by the at least one human that indicate the unique personal identifier for the at least one human and / or based on one or more items of the previous spoken utterances that request the at least one human to provide spoken utterances that include the unique personal identifier. For example,

[0068] In implementations in which the corresponding prior ongoing conversation is between instances of at least one human and a voice bot, the system can identify one or more ASR speech hypotheses that correspond to spoken utterances that include a unique personal identifier generated during the corresponding prior ongoing conversation. The system can identify the one or more ASR speech hypotheses that include the unique personal identifier based on one or more items of spoken utterance provided by the at least one human that are indicative of the unique personal identifier for the at least one human. For example, the one or more ASR speech hypotheses can include a particular alphanumeric character (e.g., an "@" sign or symbol for an email address) or a particular sequence of alphanumeric characters (e.g., a string of numbers followed by a string of letters followed by another string of numbers for a physical address). These ASR speech hypotheses can be identified as being indicative of the corresponding unique personal identifier.

[0069] At block 356A, the system utilizes the one or more ASR speech hypotheses of the unique personal identifier obtained at block 354A as a training instance input for the given training instance and utilizes a ground truth output corresponding to the unique personal identifier as a training instance output for the given training instance. The ground truth output corresponding to the unique personal identifier can be identified based on one or more supervisory signals. The one or more supervisory signals can include, for example, the unique personal identifier being input into the system (and optionally successfully finding a match to the unique personal identifier in the system) during the corresponding prior ongoing conversation, the at least one human manually inputting the unique personal identifier into the corresponding prior ongoing conversation, and / or the at least one human editing the unique personal identifier as predicted during the corresponding prior ongoing conversation (e.g., via touch input or key-in input) to the unique personal identifier, a human reviewer subsequently correcting the unique personal identifier after the corresponding prior ongoing conversation, and / or any other supervisory signal.

[0070] The given training instance can be stored in one or more databases (e.g., a training instance database 161A of the system) and the method 300A can be repeated to obtain additional training instances. In various implementations, one or more training instances can be obtained in a synchronous manner in accordance with the method 300A such that the training instances are generated during the corresponding prior ongoing conversation. In additional or alternative implementations, one or more training instances can be obtained in an asynchronous manner in accordance with the method 300A such that the training instances are generated after the corresponding prior ongoing conversation. Figure 1

[0071] In additional or alternative implementations, one or more training instances can be obtained based on processing one or more unique personal identifiers stored in one or more databases accessible by the system. For example, with particular reference to the system 100 of FIG. 1, the system 100 can obtain one or more training instances based on processing one or more unique personal identifiers stored in the training instance database 161A of the system 100. Figure 3 ​At block 352B of method 300B, the system accesses one or more databases to identify the unique personal identifier. The one or more databases can include, for example, an email address database, a physical address database, a username database, a product identifier database, a named entity database, a domain name database, and / or any other database that includes unique personal identifiers.

[0072] At block 354B, the system generates a plurality of tokens based on the unique personal identifier. In some implementations, the system can identify a plurality of n-grams included in the unique personal identifier and can generate a plurality of tokens corresponding to the unique personal identifier based on the distribution of the plurality of n-grams. For example, assume that the unique personal identifier identified from the one or more databases corresponds to the email address “tatortatorl3@exampleurl.com”. In this example, assume that at least the n-grams “tator”, “tator”, “13”, and “exampleurl.com” are identified. In this example, “tator” can not frequently appear in one or more distributions of n-grams for first names, last names, nicknames, or usernames, and thus the first instance or the second instance of “tator” can be replaced with another n-gram, such as “john” that appears more frequently in the distribution, resulting in the tokens “john”, “tator”, “13”, and “exampleurl.com” in the plurality of tokens. Notably, the originally identified tokens (e.g., at least the n-grams “tator”, “tator”, “13”, and “exampleurl.com”) can be used as tokens. However, by generating the tokens using these techniques while also preserving aspects of the unique personal identifier that can be encountered in the real world, the privacy of the user associated with the unique personal identifier can be preserved. Moreover, the resulting tokens can more accurately reflect the distribution of the unique personal identifier by replacing less frequently occurring n-grams with those that occur more frequently. Although only one n-gram is described as being replaced with respect to the unique personal identifier, it should be understood that any number of n-grams used for the unique personal identifier can be replaced in generating the tokens for purposes of example and not limitation. Also, one or more of the n-grams can be replaced based on respective distributions to generate the plurality of tokens corresponding to the unique personal identifier. For example, one or more of the n-grams can be replaced based on distributions of first names, last names, nicknames, and / or usernames, one or more distributions of non-name characters (e.g., special characters or symbols), one or more distributions of numbers, and / or any other distribution.

[0073] At block 356B, the system generates a synthesized text based on the plurality of tokens. In some implementations, the system can inject one or more fillers into the plurality of tokens. Continuing the example above, the system can pre-pend one or more fillers to one or more of the plurality of tokens or append one or more fillers to one or more of the plurality of tokens. For example, the system can pre-pend the filler "oh sure it's" before the email address or append the filler "then" or "as in" between one or more n-grams. In some additional or alternative implementations, the system can inject a phonetic spelling for one or more alphanumeric characters included in the personal identifier. Continuing the example above, the system can inject the phonetic spelling "t as in tango" for "tator." The system can inject a phonetic spelling for one or more alphanumeric characters based on a probability that the alphanumeric character can be spelled out. For example, the alphanumeric characters "t" and "d" are often confused in speech recognition, so the probability that a human provides the phonetic spelling "t as in tango" or "d as in dog" can be greater than the probability that a human provides the phonetic spelling "z as in zulu." In these examples, the resulting synthesized text can include "my email is uhh john and then t as in tango a t or 13and then@exampleurl.com." By injecting fillers and / or phonetic spellings into the plurality of tokens to generate the synthesized text, the system can learn to ignore fillers and utilize phonetic spellings when interpreting the unique personal identifier.

[0074] At block 358B, the system generates one or more ASR speech hypotheses based on the synthesized text. The system can generate the one or more ASR speech hypotheses by injecting one or more ASR errors into the one or more ASR speech hypotheses. In some implementations, the one or more ASR errors injected into the one or more ASR speech hypotheses can include substituting one or more n-grams included in the synthesized text with one or more corresponding homophonic n-grams. Continuing the example described above, the n-gram "13" in the email address can be substituted with "third teen," the n-gram "john" can be substituted with "jon," the n-gram "tator" can be substituted with "date her" or "tate her," and / or other n-grams can be substituted with corresponding homophonic tokens. In some additional or alternative implementations, the one or more ASR errors injected into the one or more ASR speech hypotheses can include substituting one or more alphanumeric characters of the synthesized text with one or more corresponding homophonic alphanumeric characters. Continuing the example described above, the alphanumeric character "t" of the n-gram "tator" can be substituted with "d," resulting in the n-gram "dator," and / or other alphanumeric characters can be substituted with corresponding homophonic alphanumeric characters. By injecting ASR errors into the one or more ASR speech hypotheses, the plurality of ML layers can be subsequently trained to handle these ASR errors that can be encountered when the plurality of ML layers is deployed for use by the voice bot.

[0075] At block 360B, the system utilizes the one or more ASR speech hypotheses generated at block 358B as training instance inputs for the given training instance and utilizes the ground truth output as a training instance output for the given training instance. The ground truth output corresponding to the unique person identifier can be identified based on one or more supervision signals. The one or more supervision signals can include, for example, the unique person identifier corresponding to the plurality of tokens generated by the system, the human reviewer subsequently providing the unique person identifier, and / or any other supervision signal.

[0076] The given training instance can be stored in one or more databases (e.g., the training instance database 161A of FIG. 1) and the method 300B can be repeated to obtain additional training instances. Notably, multiple instances of the method 300A and / or the method 300B can be executed in serial or in parallel to generate a plurality of training instances that are subsequently used to train the plurality of ML layers utilized by the voice bot in determining unique person identifiers encountered during corresponding conversations. Figure 1 The given training instance can be stored in one or more databases (e.g., the training instance database 161A of FIG. 1) and the method 300B can be repeated to obtain additional training instances. Notably, multiple instances of the method 300A and / or the method 300B can be executed in serial or in parallel to generate a plurality of training instances that are subsequently used to train the plurality of ML layers utilized by the voice bot in determining unique person identifiers encountered during corresponding conversations.

[0077] Turning now to FIG. 3B, a method 300B for training a plurality of ML layers utilized by a voice bot in determining unique person identifiers encountered during corresponding conversations is shown. The method 300B can be performed by one or more computing devices, such as the computing device 102 of FIG. 1, the computing device 202 of FIG. 2, and / or any other computing device. Figure 4depicts a flow diagram of an example method 400 that illustrates training a plurality of machine learning layers utilized by a voice bot in determining a unique person identifier. The operations of method 400 are described with reference to the system that performs the operations. This system of method 400 includes a computing device (e.g., client device 110, Figure 1 Figure 1 voice bot system 120, Figure 7 computing device 710, server, and / or other computing device) of the system of method 400 includes at least one processor, at least one memory, and / or other components. Moreover, while operations of method 400 are shown in a particular order, this is not meant to be limiting. One or more operations can be reordered, omitted, and / or added.

[0078] At block 452, the system obtains a plurality of training instances, each of the plurality of training instances including a training instance input and a training instance output, the training instance input including at least one or more ASR hypotheses for a unique person identifier, and the training instance output including a ground truth output corresponding to the unique person identifier. The plurality of training instances can be obtained in the same or similar manner as described with respect to Figure 3 method 300A of FIG. 3A and / or Figure 3 method 300B of FIG. 3B. The unique person identifier can include an alphanumeric character sequence that is personal to a given human. Moreover, the unique person identifier can be, for example, an email address, a physical address, a username, a password, a product identifier, a name of an entity, and / or a domain name

[0079] ​At block 454, the system processes one or more ASR speech hypotheses for the unique person identifier using the plurality of ML layers of the one or more ML models and for the given training instance to generate corresponding predicted measures associated with one or more alphanumeric characters of the unique person identifier. The corresponding predicted measures can include binary values, probabilities, log-likelihoods, and / or any other measures corresponding to the one or more alphanumeric characters. For example, assume that the unique person identifier associated with the given training instance corresponds to the username “examplel,” and assume that the one or more ASR speech hypotheses for the username “examplel” include at least ASR speech hypotheses corresponding to the tokens “x,” “a,” “m,” “p,” “l,” “e,” and “one.” Further, each of the one or more alphanumeric characters can be associated with a probability as the corresponding predicted measure (e.g., a first probability associated with “x,” a second probability associated with “a,” etc.). Notably, the one or more ASR speech hypotheses can include the tokens and respective measures for each token. However, in processing one or more of the plurality of ASR speech hypotheses, the plurality of ML layers can reweight the respective measures, resulting in respective predicted measures, and the corresponding predicted measures can be analyzed to remove or filter out those tokens that are not predicted to correspond to the unique person identifier. In some implementations, the system can process an intent of the voice bot along with the one or more ASR speech hypotheses using the plurality of ML layers. In the example of block 454, the intent of the voice bot can be an intent associated with requesting the unique person identifier. By processing the intent along with the one or more ASR speech hypotheses, the system can account for ASR errors associated with initially requesting the unique person identifier, and, when subsequently updating the plurality of ML layers, the system can learn from these ASR errors and refine predictions of the unique person identifier based on these ASR errors.

[0080] At block 456, the system compares the corresponding predicted measures associated with the one or more alphanumeric characters of the unique person identifier to the ground truth output corresponding to the unique person identifier to generate one or more losses. Continuing the example described above, further assume that the ground truth output corresponding to the unique person identifier includes the ground truth tokens "e", "x", "a", "m", "p", "l", "e", and "l", where each of these tokens is associated with a corresponding ground truth probability. In this example, the system can compare the tokens of "x", "a", "m", "p", "l", "e", and "one" of the ASR speech hypothesis and the corresponding predicted measures to the ground truth tokens "e", "x", "a", "m", "p", "l", "e", and "l" and the ground truth probabilities of the ground truth output. For example, the system can compare the token of "[null]" to the ground truth character "e" because the system did not predict the silent "e" character, and compare the token of "one" to the ground truth character "1" because the system predicted the word "one" instead of the letter "1". Further, the system can generate one or more losses (e.g., one or more cross-entropy losses) on a character-by-character basis.

[0081] At block 458, the system updates the plurality of ML layers based on the one or more losses. For example, the system can cause the one or more losses to backpropagate across one or more of the plurality of ML layers to update respective weights of one or more of the plurality of ML layers.

[0082] At block 460, the system determines whether one or more conditions are satisfied. The one or more conditions can include, for example, validation of the updated plurality of ML layers, convergence (e.g., zero loss or within a threshold range of zero loss) of the updated plurality of ML layers, a determination that the plurality of ML layers performs better (e.g., in terms of accuracy and / or traceability) than an instance of the plurality of ML layers currently utilized (if any) by the voice bot, an occurrence of training of at least a threshold amount of the plurality of training instances, and / or a duration of training of the plurality of training instances. If, in an iteration of block 460, the system determines that the one or more conditions are not satisfied, the system can return to block 454 to update the plurality of ML layers based on another training instance. In other words, the system can repeat the process of blocks 454, 456, 458, and 460 until the plurality of ML layers is sufficiently trained for the corresponding single-turn portion of the conversation during which the plurality of ML layers is utilized to generate corresponding predicted measures for the unique person identifier. If, in an iteration of block 460, the system determines that the one or more conditions are satisfied, the system can proceed to block 462.

[0083] At block 462, the system processes one or more ASR speech hypotheses for the given training instance using the plurality of updated ML layers and for the given training instance to generate a predicted unique person identifier. The system can process the one or more ASR speech hypotheses in the same or similar manner as described above with respect to block 454, but with the plurality of updated ML layers. In some implementations, the system can process the intent of the voice bot along with the one or more ASR speeches using the plurality of layers. Similar to block 454, the intent of the voice bot at block 462 can include an intent to request a unique person identifier. In some implementations, in generating the predicted unique person identifier, the system can use the plurality of updated ML layers to generate a likelihood tree for the unique person identifier. The likelihood tree can include, for example, a plurality of nodes and a plurality of edges. Each of the plurality of nodes can correspond to a given alphanumeric character of the predicted unique person identifier and can be associated with a corresponding prediction measure for the given alphanumeric character, and the plurality of edges can connect one or more of the plurality of nodes. In some versions of those implementations, the system can traverse the likelihood tree and / or perform a beam search in the likelihood tree to generate the predicted unique person identifier based on the likelihood tree.

[0084] At block 464, the system processes the predicted unique person identifier using a voice bot-human simulator to generate a simulated prompt for the predicted unique person identifier and a simulated response to the simulated prompt. The voice bot-human simulator can correspond to one or more processors implementing a plurality of additional ML layers of one or more ML models, and can include a voice bot simulator portion and a human simulator portion. The system can process the predicted unique person identifier and the corresponding predicted measure for the predicted unique person identifier using a simulated voice portion of the voice bot-human simulator to generate a simulated prompt for a simulated human. In generating the simulated prompt, the system identifies, from among the one or more alphanumeric characters for the predicted unique person identifier, a given alphanumeric character associated with a respective predicted measure that fails to satisfy the threshold, and generates a clarification request requesting clarification regarding the given alphanumeric character. For example, assume that the unique person identifier associated with a given training instance corresponds to the username “examplel,” and assume that the one or more ASR speech hypotheses for the username “examplel” include at least ASR speech hypotheses corresponding to the tokens “x,” “a,” “m,” “p,” “l,” “e,” and “one.” Further assume that the system is not highly confident that the username begins with “x.” Thus, in this example, the system can generate the simulated prompt “so it starts with x as in x-ray.” In addition, the system can process the simulated prompt using a simulated human portion of the voice bot-human simulator to generate a simulated response from the simulated human. Continuing the above example, the system can generate the simulated response “no, it starts with e and then x.”

[0085] At block 466, the system processes the simulated response using the plurality of updated ML layers to refine the predicted unique person identifier. In refining the predicted unique person identifier, the system can update one or more alphanumeric characters and / or corresponding prediction measures for one or more alphanumeric characters. Continuing the example above, the system can process the simulated response "no, it starts with e and then x" resulting in the tokens "e," "x," "a," "m," "p," "l," "e," and "one." In this example, the token for the alphanumeric character "e" is added to the one or more alphanumeric characters along with a corresponding prediction measure indicating that the system is highly confident that "e" corresponds to the first alphanumeric character of the unique person identifier. Further, the corresponding prediction measure associated with the token for the alphanumeric character "x" can also be updated to indicate that the system is highly confident that "x" corresponds to the second alphanumeric character of the unique person identifier. In some implementations, similar to block 454 and block 462, the system can process the intent of the voice bot along with the one or more ASR speech hypotheses using the plurality of ML updated layers. However, at the instance of block 466, the intent of the voice bot can be an intent associated with a request for clarification for the unique person identifier. By processing the intent along with the one or more ASR speech hypotheses, the system can refine the prediction of the unique person identifier based on these ASR errors by modifying the likelihood tree (e.g., adding new nodes, removing existing nodes, adjusting corresponding measures of one or more existing nodes, etc.).

[0086] At block 468, the system determines whether to generate a further simulated prompt including a further clarification request requesting further clarification for one or more corresponding alphanumeric characters of the predicted unique person identifier. The system can determine to generate the further simulated prompt in response to determining that the corresponding prediction measure associated with one or more alphanumeric characters of the refined predicted unique person identifier fails to satisfy the threshold. Although block 468 is depicted as occurring after block 466, it should be understood that this is for the purpose of example and not limitation. For example, the instance of block 468 can occur after block 462, after block 464, and / or after block 466, as depicted. Thus, if the system is sufficiently confident in each of the one or more alphanumeric characters of the predicted unique person identifier for a given training instance, then blocks 464 and 466 can be skipped for the given training instance.

[0087] If the system determines at the iteration of box 468 that it generates an additional simulated prompt, the system can return to box 464 to generate an additional simulated prompt (e.g., "is that the number 1 or is it spelled out"), and this process can be repeated until the system is highly confident in each alphanumeric character of the unique personal identifier used for prediction. If, at the iteration of box 468, the system determines that it does not generate an additional simulated prompt, the system can return to box 462 to perform additional simulation based on additional training instances. The system can repeat the processes of boxes 462, 464, 466, and 468 until one or more conditions are met (e.g., the same or similar conditions described above with respect to box 460). In other words, after initial training of multiple ML layers through the processes of boxes 454, 456, 458, and 460 described above, multiple ML layers can be further trained for multi-turn portions of the corresponding dialogue, during which multiple ML layers are used to generate clarification requests regarding alphanumeric characters, and the corresponding predictive measurement of the unique personal identifier used for prediction is refined based on simulated responses to the clarification requests.

[0088] In response to determining that one or more conditions are met, the system can proceed to box 470. In box 470, the system enables the voice robot to utilize multiple ML layers. The voice robot can be a previously trained example-based or rule-based voice robot, and the multiple ML layers can be used to determine unique personal identifiers encountered while the voice robot is engaging in a conversation with a corresponding human.

[0089] Now go to Figure 5 A flowchart illustrating an example method 500 for determining a unique personal identifier by a voice robot using multiple machine learning layers is provided. For convenience, the operation of method 500 is described with reference to a system performing the operation. This system of method 500 includes a computing device (e.g., Figure 1 Client device 110 Figure 1 Voice robot system 120 Figure 7 The computing device 710, server, and / or other computing device includes at least one processor, at least one memory, and / or other components. Furthermore, although the operations of method 500 are shown in a specific order, this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.

[0090] At block 552, the system receives audio data that captures a spoken utterance of the human during a corresponding conversation between the human and the voice bot. The audio data can be generated by one or more microphones of a client device of the human. In some implementations, the corresponding conversation can be conducted during a telephone call between the human and the voice bot via various voice communication protocols (e.g., VoIP, PSTN, and / or other telephone communication protocols). In some additional or alternative implementations, the corresponding conversation can be conducted during a conversational session between the human and the voice bot.

[0091] At block 554, the system processes the audio data using an ASR model to generate a plurality of ASR speech hypotheses. In some implementations, the ASR is an end-to-end speech recognition model that can directly use the model to generate the plurality of speech hypotheses (e.g., on a character-by-character basis or on other token-by-token basis). In other implementations, the ASR model is not an end-to-end speech recognition model and can instead generate predicted phonemes (and / or other representations) and can optionally utilize a decoding graph, a lexicon, and / or other resources.

[0092] At block 556, the system predicts whether the spoken utterance includes a unique personal identifier. In some implementations, the system can predict whether the spoken utterance includes a personal identifier based on synthetic speech audio data that includes synthetic speech that was previously provided for presentation to the human by the voice bot during the corresponding conversation. For example, if the voice bot previously requested the human to provide a unique personal identifier, the system can predict that the spoken utterance includes a unique personal identifier. In some additional or alternative implementations, the system can predict whether the spoken utterance includes a personal identifier based on the plurality of speech hypotheses generated using the ASR model. For example, if one or more of the plurality of ASR speech hypotheses includes a given alphanumeric character that indicates a unique personal identifier (e.g., a string of numbers, a particular symbol or character (e.g., an “@” symbol, an underscore, etc.), and / or any other indicator that the spoken utterance includes a unique personal identifier), the system can predict that the spoken utterance includes a unique personal identifier.

[0093] If at an iteration of block 556 the system predicts that the spoken utterance does not include a unique personal identifier, the system can return to block 552 to receive additional audio data and repeat the process of blocks 552, 554, and 556 to determine whether the additional audio data captures an additional spoken utterance that includes a unique personal identifier. This process can be repeated for any further additional audio data received during the corresponding conversation. If at an iteration of block 556 the system predicts that the spoken utterance includes a unique personal identifier, the system can proceed to block 558.

[0094] At block 558, the system processes one or more of the plurality of ASR speech hypotheses using a plurality of ML layers of one or more ML models to generate one or more candidate unique person identifiers and corresponding predicted measures associated with one or more corresponding alphanumeric characters of the one or more candidate unique person identifiers. The plurality of ML layers can be trained in the same or similar manner as described with respect to the method 400 of Figure 4 For example, assume that the unique person identifier included in the spoken utterance is the email address "johnp@exampleurl.com." Further assume that the system generates a first candidate unique person identifier having at least tokens for "j" associated with a first probability, "o" associated with a second probability, "n" associated with a third probability, and "p" associated with a fourth probability (e.g., corresponding to the alphanumeric sequence "jonp"), generates a second candidate unique person identifier having at least tokens for "j" associated with a first probability, "o" associated with a second probability, "h" associated with a third probability, "n" associated with a fourth probability, and "d" associated with a fifth probability (e.g., corresponding to the alphanumeric sequence "johnd"), and so on. In some implementations, the system generates a likelihood tree and generates the one or more candidate unique person identifiers based on the likelihood tree (e.g., as described with respect to the method 400 of Figure 2B In additional or alternative implementations, the system generating the one or more candidate unique person identifiers can be generated directly across the plurality of ML layers.

[0095] At block 560, the system selects one or more given alphanumeric characters for the unique person identifier based on the corresponding predicted measures associated with the one or more corresponding alphanumeric characters for the one or more candidate unique person identifiers. The system can select the one or more given alphanumeric characters for the unique person identifier, for example, on a token-by-token (or other character-by-character) basis, on a candidate-by-candidate basis, and / or on any other basis. Continuing the above example, the system can identify at least the tokens "j," "o," "h," "n," "p," and "d" as corresponding to a given unique person identifier on a token-by-token basis and select one or more of the given alphanumeric characters, such as "j," "o," "n," and "p," based on the corresponding measures associated with each of these tokens, resulting in the candidate unique person identifier of "jonp" as corresponding to the given unique person identifier based on the corresponding measures associated with each of these tokens of the unique person identifier or the corresponding measure associated with the candidate unique person identifier in its entirety.

[0096] At block 562, the system determines whether to generate a prompt including a clarification request requesting clarification of one or more given alphanumeric characters for the given unique personal identifier. The system can determine whether to generate the prompt in response to determining that the corresponding prediction measure associated with the one or more given alphanumeric characters of the given unique personal identifier fails to satisfy the threshold. If at the iteration of block 562 the system determines not to generate the prompt, the system can refrain from generating the prompt and proceed to block 574. Block 574 is described below. If at the iteration of block 562 the system determines to generate the prompt, the system can generate the prompt and can proceed to block 564. In generating the prompt, the system identifies one or more particular alphanumeric characters from among the one or more given alphanumeric characters that are associated with the corresponding prediction measure that fails to satisfy the threshold, and generates a clarification request requesting clarification regarding the given alphanumeric character. Continuing the example described above, assume that the corresponding prediction measure associated with at least the token "h" fails to satisfy the threshold. In this example, the system can identify the token "h" and generate a prompt requesting clarification. For example, the prompt can include the clarification request "is there an h or no h," "is that jon with an h," "is it j on or j o h n," "so its j o n," and so forth.

[0097] At block 564, the system causes a prompt to be presented to the human clarifying the one or more given alphanumeric characters. The prompt can be visually rendered via a display of the human's client device or another client device and / or can be audibly rendered via one or more speakers of the client device or another client device. At block 566, the system receives additional audio data capturing an additional spoken utterance of the human in response to the prompt. At block 568, the system processes the additional audio data using the ASR model to generate a plurality of additional ASR speech hypotheses. At block 570, the system processes the additional ASR speech hypotheses using the plurality of ML layers to refine the one or more given unique personal identifiers. Continuing the example above, assume that the prompt includes the clarifying request "is that jon with an 'h'" and the prompt is audibly rendered for presentation to the human. Further assume that the human provides an additional spoken utterance indicating that the unique personal identifier includes "john with an h." Thus, in processing the one or more additional ASR speech hypotheses, the system can refine the given unique personal identifier to include the token "h" and update the corresponding prediction measure associated with the token "h" to indicate that the system is highly confident in the given unique personal identifier including the token "h." In implementations in which the given unique personal identifier is generated based on a tree of possibilities, any replacement nodes for the token "h," such as a "null" node, can be removed from the tree of possibilities.

[0098] At block 572, the system determines whether to generate a further prompt that includes a further clarification request that requests further clarification of one or more given alphanumeric characters. The system can determine to generate the further prompt in response to determining that the corresponding prediction measure associated with the one or more given alphanumeric characters of the refined given unique personal identifier fails to satisfy the threshold. If at the iteration of block 572 the system determines not to generate the further prompt, the system can refrain from generating the further prompt and proceed to block 574. Block 574 is described below. If at the iteration of block 562 the system determines to generate the further prompt, the system can generate the further prompt and can proceed to block 564. In generating the further prompt, the system identifies one or more further particular alphanumeric characters from among the one or more corresponding alphanumeric characters that are associated with the corresponding prediction measure that fails to satisfy the threshold, and generates the further clarification request that requests further clarification with respect to the one or more further particular alphanumeric characters. Continuing the example above, further assume that the system is not confident whether the given unique personal identifier is corresponding to "johnp" or "johnd." In this example, the system can identify the token "p" and / or the token "d," and generate the further prompt "and was that p as in papa or d as in delta," "so it's j o h n and then p," "so it's j o h n and then d," etc. Further assume that the prompt includes the clarification request "and was that p as in papa or d as in delta," assume that the prompt is audibly rendered for presentation to the human, and assume that the human provides a further spoken utterance that indicates that the unique personal identifier includes "p as in papa." Thus, in processing the one or more further ASR speech hypotheses, the system can refine the given unique personal identifier to include the token "p," and update the corresponding prediction measure associated with the token "p" to indicate that the system is highly confident in the given unique personal identifier that includes the token "p." In implementations in which the given unique personal identifier is generated based on a likelihood tree, any replacement nodes for the token "p" can be removed from the likelihood tree, e.g., the "d" node.

[0099] The system can repeat the processes of blocks 564, 566, 568, 570, and 572 until the system determines that the corresponding predicted measure associated with each of the one or more given alphanumeric characters of the given unique personal identifier for refinement satisfies the threshold. In other words, the system can continue to prompt the human to clarify one or more corresponding alphanumeric characters until the system is sufficiently confident that the given unique personal identifier for refinement is in fact the unique personal identifier of the human that was initially provided by the human.

[0100] At block 574, the system causes the voice bot to utilize the unique personal identifier to facilitate the corresponding conversation. In implementations in which the corresponding conversation is during a telephone call between the voice bot and the human, the voice bot can utilize the unique personal identifier to continue performing a task requested by the human (e.g., for customer service, for a query related to a user account, and / or any other task that can be performed during a telephone call). For example, the voice bot can utilize the unique personal identifier to verify or authenticate the identity of the human, search for information related to the unique personal identifier, and / or any other way in which the voice bot can utilize the unique personal identifier to continue performing the telephone call. In implementations in which the corresponding conversation is during a conversation session between the voice bot and the human, the voice bot can utilize the unique personal identifier to incorporate the unique personal identifier into a transcription (e.g., when the human is dictating an email, a text message, an SMS message, a note, a calendar entry, and / or otherwise dictating to the voice bot), perform an action on behalf of the user (e.g., make a purchase on behalf of the user, log into an account of the user, and / or any other action on behalf of the user), and / or any other way in which the voice bot can utilize the unique personal identifier to continue performing the conversation.

[0101] Turning now to Figure 6A to 6C various non-limiting examples of corresponding conversations between a voice bot and a human that include determining a unique personal identifier are depicted. Figure 6A and 6B each depict a client device 610A having a graphical user interface 680A and can include one or more of Figure 1 the client device 110 of Figure 1 another client device 195 of Figure 7 components of the computing device 710 of Figure 6A and 6B operations of Figure 6A and 6BThe client device 610A is described as a smartphone, but it should be understood that this is not intended to be restrictive. For example, Figure 6C Client device 610C is described as a stand-alone auxiliary device without any display. As other examples, client devices 610A and / or 610C may be stand-alone auxiliary devices with displays, laptop computers, desktop computers, vehicle computing devices, and / or any other client devices capable of making telephone calls and / or engaging in human-computer dialogue.

[0102] Figure 6A and 6B The graphical user interface 680A also includes a text response interface element 684, which the user can choose to generate user input via a virtual keyboard or other touch and / or typing input, and a voice response interface element 685, which the user can choose to generate user input via the microphone of the client device 610A. In some embodiments, the user can generate user input via the microphone without selecting the voice response interface element 685. For example, active monitoring of audible user input via the microphone may occur to avoid the need for the user to select the voice response interface element 685. For example, active monitoring of audible user input during a telephone call and / or active monitoring of audible user input including specific words or phrases may occur to avoid the need for the user to select the voice response interface element 685. In some and / or other embodiments of those embodiments, the voice response interface element 685 may be omitted. Moreover, in some embodiments, the text response interface element 684 may be additionally and / or alternatively omitted (e.g., the user can provide only audible user input). Figure 6A and 6B The graphical user interface 680A also includes system interface elements 681, 682, and 683, which can be interacted with by the user to cause the client device 610A to perform one or more actions.

[0103] For example, see specific references Figure 6A Suppose that a human associated with client device 610A initiates a corresponding telephone call with the customer service of a hypothetical technology company, referred to as the example widget. Further assume that the example widget utilizes a voice robot trained to handle inbound telephone calls for customer service requests and has access to multiple ML layers of one or more ML models, which are configured according to the techniques described herein (e.g., regarding...). Figure 2A and 4) is trained and used during the corresponding conversation to determine a unique personal identifier. In this example, the corresponding telephone call can be performed using various voice communication protocols including, for example, VoIP, PSTN, and / or other telephone communication protocols. As described herein, synthetic speech can be rendered by the voice bot and on behalf of the example widget as part of the corresponding telephone call, which can include injecting the synthetic speech into the corresponding telephone call such that it is perceptible by a human associated with the client device 610A. The synthetic speech can be generated and / or injected by the client device 610A that is one of the endpoints of the corresponding telephone call, and / or can be generated and / or injected by a server that is in communication with the client device 610A and also connected to the corresponding telephone call.

[0104] It is further assumed that the human associated with the client device 610A navigates the corresponding telephone call to a point where the voice bot causes synthetic speech 652A1“Thanks for contacting customer support, may I please have your username?” to be rendered for presentation to the human associated with the client device 610A. For example, the human can select a“customer service” option from among a plurality of options presented to the user via an interactive voice response (IVR) system, the human can provide a free input of“customer service” when prompted to provide a reason for initiating the corresponding telephone call, and / or any other manner of navigating the corresponding telephone call to a point where the voice bot causes synthetic speech 652A1 to be rendered for presentation to the human associated with the client device 610A. It is further assumed that the human provides spoken utterance 654A1“yes, it’s tatortator13” in response to rendering the synthetic speech 652A1. In this example, the username“tatortator13” provided by the human associated with the client device 610A is a unique personal identifier of an account of the human associated with the example widget.

[0105] The voice bot can use the ASR model to process audio data corresponding to the spoken utterance 654A1 to generate a plurality of ASR speech hypotheses for the spoken utterance 654A1. In addition, in response to predicting that the spoken utterance 654A1 includes a unique personal identifier, the voice bot can use the techniques described herein (e.g., with respect to FIGS. 6A-6C) to generate a plurality of ASR speech hypotheses for the spoken utterance 654A1 that include a unique personal identifier. Figure 1 、 Figure 2A and / or Figure 4) the one or more ML models trained using the techniques described herein to process one or more of the plurality of ASR speech hypotheses to generate one or more candidate unique personal identifiers and a corresponding predicted measure associated with each of the one or more candidate unique personal identifiers. In some implementations, the voice bot can predict that the spoken utterance 654A1 includes a unique personal identifier based on the synthesized speech 652A1 including a request for a human to provide a unique personal identifier (e.g., “may I please have your username” included in the synthesized speech 652A1 to find your account”). In additional or alternative implementations, the voice bot can predict that the spoken utterance 654A1 includes a unique personal identifier based on one or more of the plurality of ASR speech hypotheses.

[0106] In some implementations, in generating the one or more candidate unique personal identifiers and the corresponding predicted measures, the voice bot can generate a likelihood tree for the unique personal identifier included in the spoken utterance 654A1 based on the outputs generated across the plurality of ML layers. For example, assume that a given ASR speech hypothesis for a unique personal identifier corresponds to the tokens “tatordater 13” (and optionally associated with respective predicted measures for one or more of the tokens). In this example, the voice bot can use the plurality of ML layers to process the tokens for the given ASR speech hypothesis to generate outputs. Further, the voice bot can generate a likelihood tree based on the outputs. The likelihood tree can include a plurality of nodes corresponding to the alphanumeric characters for the unique personal identifier and the corresponding predicted measures associated with each of the alphanumeric characters, and can include a plurality of edges connecting the plurality of nodes (e.g., as described with respect to FIG. 6B). In some implementations, the voice bot can generate the likelihood tree based on the outputs generated by the plurality of ML layers. In additional or alternative implementations, the voice bot can generate the likelihood tree based on the outputs generated by the plurality of ML layers and one or more additional models (e.g., a language model, a grammar model, a lexicon model, etc.). Figure 2BThe voice bot can generate one or more of the candidate unique person identifiers based on the likelihood tree. For example, the voice bot can identify a first node corresponding to the token "t," a second node corresponding to the token "a" connected to the first node by a first edge, a third node corresponding to the token "t" connected to the second node by a second edge, a fourth node corresponding to the token "o" connected to the third node by a third edge, a fifth node corresponding to the token "r" connected to the fourth node by a fourth edge, and so on, to generate a first candidate unique person identifier including at least the alphanumeric characters "t a t o r." In this example, the voice bot can identify a first replacement node corresponding to the token "d" connected to the second node by a replacement of the first edge, and can utilize the second through fifth nodes to generate a second candidate unique person identifier including at least the alphanumeric characters "d a t o r." It will be appreciated that the nodes and edges are scalable to include a plurality of additional nodes and a plurality of additional edges corresponding to alphanumeric characters (and including corresponding predictive measures), and that this example is provided for purposes of illustration. Notably, the one or more candidate unique person identifiers generated based on the likelihood tree can include every combination of alphanumeric characters, a subset thereof including only alphanumeric characters associated with corresponding predictive measures satisfying a threshold, and / or those identified across the likelihood tree responsive to the beam search.

[0107] In further or alternative implementations, in generating the one or more candidate unique person identifiers and corresponding predictive measures, the voice bot can generate the one or more candidate unique person identifiers based on outputs generated across the plurality of ML layers. For example, assume that a given ASR speech hypothesis for a unique person identifier corresponds to the tokens "tator dater 13" (and optionally associated with respective measures for one or more of the tokens). In this example, the voice bot can use the plurality of ML layers to process the tokens for the given ASR speech hypothesis to generate outputs. In this example, the outputs generated across the model can be an n-dimensional vector including one or more candidate unique person identifiers (e.g., where n is a positive integer). For example, the outputs can include at least: a first candidate unique person identifier including at least the alphanumeric characters "t a t o r," a second candidate unique person identifier including at least the alphanumeric characters "d a t o r," and / or other strings of alphanumeric characters predicted to correspond to one or more portions of the unique person identifier captured in the spoken utterance 654A1.

[0108] The voice bot can generate one or more prompts that include a corresponding clarification request that requests clarification of the unique person identifier. For example, the voice bot can generate the following prompt: request that the human spell the unique person identifier or a portion thereof, request that the human clarify a particular alphanumeric character included in the unique person identifier, request that the human verify that the unique person identifier as perceived by the voice bot is in fact the unique person identifier, and / or any other request relevant to determining the unique person identifier. For example, in response to the human associated with the computing device 610A providing the spoken utterance 654A1 that includes the unique person identifier, the voice bot assumes generates and renders the synthesized speech 652A2 "Okay, looking for that username, can you spell it for me?" Notably, the synthesized speech 652A2 includes a prompt that requests the human to spell the username. In some implementations, the voice bot can prompt the human to spell the unique person identifier in response to determining that a corresponding prediction measure for a first alphanumeric character of one or more candidate unique person identifiers fails to satisfy a threshold. In further or alternative implementations, the voice bot can prompt the human to spell the unique person identifier in response to determining that a corresponding prediction measure for each alphanumeric character of one or more candidate unique person identifiers fails to satisfy a threshold.

[0109] Further assume, in response to rendering synthesized speech 652A2, the human provides spoken utterance 654A2 “Yes, t a t or t a t or 13” (Good, t a t or t a t or 13). The voice bot can use an ASR model to process audio data corresponding to spoken utterance 654A2 to generate a plurality of additional ASR speech hypotheses for spoken utterance 654A2. Further, the voice bot can use a plurality of ML layers to process one or more of the plurality of additional ASR speech hypotheses to refine one or more candidate unique person identifiers and / or a corresponding predicted measure associated with each of the one or more candidate unique person identifiers. For example, the voice bot can update the corresponding predicted measure to indicate that the alphanumeric character of the unique person identifier corresponding to the letter “t”. For example, the voice bot can generate synthesized speech 652A3 “Alright, so t a t or and then d as in delta” (Good, so t a t or and then d as in delta)?”. This indicates that the voice bot is highly confident in the alphanumeric character of the unique person identifier corresponding to the letter “t a t or”, but is not confident after the letter “r”. Notably, synthesized speech 652A3 includes a prompt requesting the human to clarify the alphanumeric character after the letter “r” (e.g., “and then das in delta”).

[0110] It is further assumed that the human provides a spoken utterance 654A3 "No, it's t a t or and then t as in tango, a t or, and then 13" in response to the rendered synthesized speech 652A3. The voice bot can use an ASR model to process audio data corresponding to the spoken utterance 654A3 to generate a plurality of further additional ASR speech hypotheses for the spoken utterance 654A3. In addition, the voice bot can use a plurality of ML layers to process one or more of the plurality of further additional ASR speech hypotheses to further refine one or more candidate unique person identifiers and / or a corresponding predicted measure associated with each of the one or more candidate unique person identifiers. For example, in implementations in which the voice bot generates the one or more candidate unique person identifiers based on a likelihood tree, the voice bot can remove or ignore replacement nodes for at least the first six alphanumeric characters (e.g., t a t o r t). As another example, in implementations in which the voice bot generates the one or more candidate unique person identifiers based on outputs generated across models, the voice bot can prune any candidate unique person identifiers that do not confirm with at least the first six alphanumeric characters (e.g., t a t o r t). The process of prompting the human for a corresponding clarification request and refining the one or more candidate unique person identifiers can be repeated until the voice bot is sufficiently confident that a given one of the one or more candidate unique person identifiers is the unique person identifier provided by the human. For example, the voice bot can generate synthesized speech 652A4 "Alright, so t a t o r t a t o r, and then the number 13?" This indicates that the voice bot is highly confident that a given one of the one or more candidate unique person identifiers is in fact the unique person identifier provided by the human in the spoken utterance 654A1 (and as verified by the human through the spoken utterance 654A4 "Yes"). The voice bot can then utilize the determined unique person identifier to facilitate the corresponding telephone call.

[0111] In various implementations, and in conjunction with the above Figure 6AOne or more of the plurality of ASR speech hypotheses described in the above process, the voice bot can use a plurality of ML layers to process an intent of the voice bot. The intent of the voice bot can include, for example, a request for the human to provide a unique personal identifier (e.g., associated with the above synthesized speech 652A1), a request for the human to spell a unique personal identifier (e.g., associated with the above synthesized speech 652A2), a request for the human to clarify one or more alphanumeric characters included in a unique personal identifier (e.g., associated with the above synthesized speech 652A2), a request for the human to verify a given one of one or more candidate unique personal identifiers as the human’s unique personal identifier (e.g., associated with the above synthesized speech 652A4), and / or any other intent related to determining a unique personal identifier of the human associated with the client device 610A. By processing the intent of the voice bot along with one or more of the plurality of ASR speech hypotheses, the voice bot can more quickly and efficiently determine the unique personal identifier of the human associated with the client device 610A by limiting the one or more candidate unique personal identifiers based on the intent of the voice bot.

[0112] Although described with respect to determining a unique personal identifier of the human associated with the client device 610A during a corresponding phone call, Figure 6A it should be understood that this is for purposes of example and is not limiting. For example, with specific reference to Figure 6B , assume that the human associated with the client device 610A initiates a corresponding dialog session with the voice bot by providing a spoken utterance 654B1, “Voice Bot, email Branden the tickets for the concert tonight.” In this example, the voice bot can use an ASR model to process audio data corresponding to the spoken utterance 654B1 to determine that the spoken utterance corresponds to a request for the voice bot to send an email to Branden that includes tickets for the concert tonight. However, assume that Branden’s email is not readily ascertainable by the voice bot, and assume that the voice bot generates and renders a synthesized speech 652B1, “Okay, what’s Branden’s email?” at the client device 610A. In this example, the unique personal identifier is an email address of another human (e.g., Branden) that can be provided in a spoken utterance 654B2, and the voice bot can prompt the human associated with the client device 610A via further synthesized speech (e.g., 652B2 and 652B3), and based on the spoken utterance 654B2, determine the unique personal identifier of the human associated with the client device 610A as the email address of Branden. Figure 6AThe same or similar manner is used to refine one or more candidate unique personal identifiers in response to further verbal utterances from a human corresponding to the prompt (e.g., 654B3 and 654B4). Furthermore, when performing an action (e.g., successfully sending a ticket email to Branden), the voice robot can cause the generation of synthetic speech 652B4 and render it for presentation to the human associated with client device 610A.

[0113] Moreover, despite the description of... Figure 6A The corresponding telephone call and about Figure 6B The transcription of the corresponding dialogue session is described, but it should be understood that this is for illustrative purposes and does not imply limitation. For example, see specific references. Figure 6C Client device 610C is described as lacking a display. Furthermore, it is assumed that the human associated with client device 610C (e.g., Figure 6C The human 601 shown initiates a corresponding dialogue session with the voicebot by providing the synthesized voice 654C1, “Voice Bot, tell Branden to send the prize money to my quick cash account and I'll see him later tonight.” However, it is assumed that the human's quick cash account is not easily ascertainable by the voicebot, and that the voicebot generates and renders the synthesized voice 652C1, “Okay, what is your quickcash account?”, at the client device 610C. In this example, the unique personal identifier is the human's account associated with the client device 610C, which can be provided in the spoken utterance 654C2, and the voicebot can prompt the human associated with the client device 610A via further synthesized voice (e.g., 652C2), and based on the information provided… Figure 6A The same or similar manner is used to refine one or more candidate unique personal identifiers in response to further verbal utterances from a human corresponding to the prompt (e.g., 654C3). Furthermore, when an action is performed (e.g., a message is successfully sent to Branden), the voice robot can cause the generation of synthetic speech 652C3 and render it for use with a human associated with client device 610A.

[0114] It should be understood that, regarding Figure 6A to 6CThe described example corresponding conversation unique personal identifiers, prompts, responses, and / or any other aspects are provided for purposes of illustration only and are not meant to be limiting. For example, a unique personal identifier can include any other sequence of alphanumeric characters that is personal to a human being. As another example, a prompt generated by a voice bot and rendered for presentation to a human being can depend on corresponding predicted measures associated with one or more candidate unique personal identifiers, and a response provided by a human being can be a free-form typed input or a spoken input. As yet another example, a corresponding conversation (and any tasks performed by a voice bot during a corresponding conversation) can depend on needs of a human being participating in a corresponding conversation. Nonetheless, the technology described herein can be utilized by a voice bot in any of these scenarios to determine a unique personal identifier provided by a human being.

[0115] Figure 7 is a block diagram of an example computing device 710 that can optionally be used to perform one or more aspects of the technology described herein. In some implementations, one or more of a client device, cloud-based automated assistant components, and / or other components can include one or more components of the example computing device 710.

[0116] The computing device 710 typically includes at least one processor 714 which communicates with a number of peripheral devices via bus subsystem 712. These peripheral devices can include a storage subsystem 724, including, for example, a memory subsystem 725 and a file storage subsystem 726, user interface input devices 722, user interface output devices 720, and a network interface subsystem 716. The input and output devices allow user interaction with the computing device 710. Network interface subsystem 716 provides an interface to an external network (e.g., the Internet) and is coupled to corresponding interface devices in other computing devices.

[0117] User interface input devices 722 can include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen incorporated into the display, audio input devices such as voice recognition systems, microphones, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information to the computing device 710 or to a communication network.

[0118] User interface output devices 720 can include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices. The display subsystem can include a cathode ray tube (CRT), a flat-panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may, for example, provide non-visual display via audio output devices. In general, use of the term "output device" is intended to include all

[0119] Storage subsystem 724 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 724 can include the logic to perform selected aspects of the methods disclosed herein, as well as to implement various components depicted in the Figure 1 flowcharts of FIGS. 6-9.

[0120] These software modules are generally executed by processor 714 alone or in combination with other processors. Memory 725 used in the storage subsystem 724 can include a number of memories including a main random access memory (RAM) 730 for storage of instructions and data during program execution and a read only memory (ROM) 732 in which fixed instructions are stored. A file storage subsystem 726 can provide persistent (nonvolatile) storage for program and data files, and can include a hard disk drive, a floppy disk drive along with associated removable media, a CD-ROM drive, an optical drive, or removable media cartridges. The modules implementing the functionality of certain implementations can be stored by file storage subsystem 726 in the storage subsystem 724, or in other machines accessible by the processor(s) 714.

[0121] Bus subsystem 712 provides a mechanism for letting the various components and subsystems of computer 710 communicate with each other as intended. Although bus subsystem 712 is shown schematically as a single bus, alternative implementations of the bus subsystem 712 can use multiple busses.

[0122] Computer 710 can be of various types, including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computer 710 Figure 7 The description of computer 710 depicted in FIG. 6 is intended only as a specific example for purposes of illustrating some implementations. Many other configurations of computer 710 are possible having more or fewer components than the computer depicted in FIG. 6, each configured to perform the functions of the various other implementations. Figure 7 The description of computer 710 depicted in FIG. 6 is intended only as a specific example for purposes of illustrating some implementations. Many other configurations of computer 710 are possible having more or fewer components than the computer depicted in FIG. 6, each configured to perform the functions of the various other implementations.

[0123] In situations in which the systems described herein collect or otherwise monitor personal information about users, or can make use of personal and / or monitored information, the users can be provided with an opportunity to control whether programs or features collect user information (e.g., information about a user's social network, social actions or activities, profession, a user's preferences, or a user's current geographic location), or to control whether and / or how to receive content from the content server that can be more relevant to the user. Also, certain data can be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user's identity can be treated so that no personally identifiable information can be determined for the user, or a user's geographic location can be generalized where location information is obtained (such as to a city, postal code, or state level), so that a particular geographic location of a user cannot be determined. Thus, the user can have control over how information is collected about the user and / or used.

[0124] In some implementations, a method implemented by one or more processors is provided and the method includes receiving audio data that captures a spoken utterance of a human, the spoken utterance being received by a voice bot during a corresponding conversation between the human and the voice bot; processing the audio data with an automatic speech recognition (ASR) model to generate a plurality of ASR speech hypotheses; and in response to predicting that the spoken utterance includes a unique personal identifier, the unique personal identifier including a sequence of unique alphanumeric characters that is personal to the human: processing one or more of the plurality of ASR speech hypotheses using a plurality of machine learning (ML) layers of one or more ML models to generate one or more candidate unique personal identifiers, each of the one or more candidate unique personal identifiers including a corresponding predicted measure associated with one or more corresponding alphanumeric characters for each of the one or more candidate unique personal identifiers; selecting one or more given alphanumeric characters from among the one or more corresponding alphanumeric characters based on the corresponding predicted measure associated with the one or more corresponding alphanumeric characters for each of the one or more candidate unique personal identifiers; generating a prompt for a clarification request that includes a request for clarification of the one or more given alphanumeric characters based on the corresponding predicted measure associated with the one or more given alphanumeric characters; and causing the prompt to be provided for presentation to the human.

[0125] These and other implementations of the technology disclosed herein can optionally include one or more of the following features.

[0126] In some implementations, the method can further include, in response to the prompt being provided for presentation to the human: receiving additional audio data that captures an additional spoken utterance of the human, the additional spoken utterance being received by the voice bot during the corresponding conversation; processing the additional audio data using the ASR model to generate a plurality of additional ASR speech hypotheses; and processing one or more of the plurality of additional speech hypotheses using the plurality of ML layers to refine one or more of the given alphanumeric characters.

[0127] In some versions of those implementations, refining one or more of the given alphanumeric characters can include updating corresponding prediction measures for one or more of the given alphanumeric characters that are predicted to correspond to the unique personal identifier of the human based on the clarification received in response to the clarification request.

[0128] In some versions of those implementations, the method can further include, prior to the one or more given alphanumeric characters being predicted to correspond to the unique personal identifier of the human: generating one or more corresponding additional prompts based on corresponding prediction measures associated with one or more of the given alphanumeric characters, each of the one or more corresponding additional prompts including a corresponding additional clarification request that requests an additional clarification of the one or more of the given alphanumeric characters; and causing one or more of the corresponding additional prompts to be provided for presentation to the human. In some further versions of those implementations, predicting that the one or more given alphanumeric characters correspond to the unique personal identifier of the human can include determining that the corresponding prediction measure associated with each of the given alphanumeric characters satisfies a threshold. In still further versions of those implementations, the method can further include, in response to predicting that the one or more given alphanumeric characters correspond to the unique personal identifier of the human: facilitating the corresponding conversation between the voice bot and the human with the given unique personal identifier that includes the one or more given alphanumeric characters.

[0129] In some implementations, generating the prompt that includes the clarification request can be in response to determining that the corresponding prediction measure associated with one or more of the given alphanumeric characters fails to satisfy a threshold. In some versions of those implementations, generating the prompt that includes the clarification request can include identifying one or more particular alphanumeric characters from among the one or more of the given alphanumeric characters that are associated with the corresponding prediction measure that fails to satisfy the threshold; and generating a clarification request that requests a clarification regarding the one or more of the particular alphanumeric characters.

[0130] In some implementations, the method can further include, in response to predicting that the one or more given alphanumeric characters correspond to a unique personal identifier of a human: facilitating a corresponding conversation between the voice bot and the human using the given unique personal identifier that includes the one or more given alphanumeric characters. In some versions of those implementations, predicting that the one or more given alphanumeric characters correspond to a unique personal identifier of a human can include determining that a corresponding prediction measure associated with each of the one or more given alphanumeric characters for a given candidate unique personal identifier satisfies a threshold.

[0131] In some implementations, predicting that the spoken utterance includes a unique personal identifier can include predicting that the audio data will include a unique personal identifier based on the synthetic speech audio data including synthetic speech that was previously provided for presentation to the human by the voice bot during a corresponding conversation.

[0132] In some implementations, predicting that the spoken utterance includes a unique personal identifier can include predicting that the spoken utterance includes a unique personal identifier based on one or more of the plurality of ASR speech hypotheses generated using the ASR model.

[0133] In some implementations, processing one or more of the plurality of speech hypotheses to generate one or more candidate unique personal identifiers can include iteratively processing each of the plurality of ASR speech hypotheses using a plurality of ML layers to iteratively generate a likelihood tree for a unique personal identifier, the likelihood tree including a plurality of nodes and a plurality of edges, each of the plurality of nodes corresponding to one or more of a corresponding alphanumeric character for each of the one or more corresponding alphanumeric characters, each of the plurality of nodes being associated with a corresponding prediction measure for each of the one or more corresponding alphanumeric characters, and each of the plurality of nodes being connected by one or more of the plurality of edges. Selecting a given candidate unique personal identifier can be based on the likelihood tree. In some versions of those implementations, the likelihood tree can be constrained by the plurality of ASR speech hypotheses. In some further versions of those implementations, the likelihood tree can be constrained by the plurality of unique personal identifiers stored in one or more databases.

[0134] In some implementations, processing one or more of the plurality of speech hypotheses to generate one or more candidate unique personal identifiers can include processing each of the plurality of ASR speech hypotheses using a plurality of ML layers to generate a likelihood tree for a unique personal identifier, the likelihood tree including a plurality of nodes and a plurality of edges, each of the plurality of nodes corresponding to one or more of a corresponding alphanumeric character for each of one or more corresponding alphanumeric characters, each of the plurality of nodes being associated with a corresponding predicted measure for each of the one or more corresponding alphanumeric characters, and each of the plurality of nodes being connected by one or more of the plurality of edges. Selecting a given candidate unique personal identifier can be based on the likelihood tree.

[0135] In some implementations, the unique personal identifier can be one or more of: an email address, a physical address, a username, a password, a name of an entity, or a domain name.

[0136] In some implementations, the method can further include obtaining an intent of the voice bot for a portion of a corresponding conversation between the voice bot and the human. Processing one or more of the plurality of ASR speech hypotheses using a plurality of ML layers to generate one or more of the candidate unique personal identifiers can further include processing the intent of the voice bot using the plurality of ML layers to generate one or more of the candidate unique personal identifiers. In some versions of those implementations, the intent of the voice bot includes one or more of: requesting the human to provide a unique personal identifier; requesting the human to spell a unique personal identifier; or requesting the human to provide clarification of one or more of a given alphanumeric character.

[0137] In some implementations, a method implemented by one or more processors is provided and the method includes: receiving audio data that captures a spoken utterance of a human, the spoken utterance being received by a voice bot during a corresponding conversation between the human and the voice bot; processing the audio data with an automatic speech recognition (ASR) model to generate a plurality of ASR speech hypotheses; and in response to predicting that the spoken utterance includes a unique personal identifier, the unique personal identifier including a sequence of unique alphanumeric characters that is personal to a given user: processing one or more of the plurality of ASR speech hypotheses using a plurality of machine learning (ML) layers of one or more ML models to generate one or more candidate unique personal identifiers, each of the one or more candidate unique personal identifiers including a corresponding predicted measure associated with one or more corresponding alphanumeric characters; and based on the corresponding predicted measure associated with one or more of the corresponding alphanumeric characters for each of the one or more candidate unique personal identifiers, selecting one or more given alphanumeric characters from among the one or more of the corresponding alphanumeric characters. The method further includes, prior to a given candidate unique personal identifier being predicted to correspond to the unique personal identifier: generating a corresponding prompt based on the corresponding predicted measure associated with one or more of the corresponding alphanumeric characters for the given candidate unique personal identifier, the corresponding prompt including a corresponding clarification request that requests clarification for the one or more of the corresponding alphanumeric characters for the given candidate unique personal identifier; causing the corresponding prompt to be provided for presentation to the human; and refining the given candidate unique personal identifier based on corresponding additional audio data that processes a corresponding additional spoken utterance of the human and that is responsive to the corresponding prompt. The method further includes, in response to predicting that the given candidate unique personal identifier corresponds to the unique personal identifier: facilitating the corresponding conversation between the voice bot and the human with the given unique personal identifier that includes the one or more given alphanumeric characters.

[0138] In some implementations, a method implemented by one or more processors is provided and the method includes: obtaining a plurality of training instances, each of the plurality of training instances including: a training instance input, the training instance input including at least one automatic speech recognition (ASR) speech hypothesis for a unique personal identifier, the unique personal identifier including a sequence of unique alphanumeric characters that is personal to a given human, and a training instance output, the training instance output including a corresponding ground truth output that corresponds to the unique personal identifier. The method further includes: training a plurality of machine learning (ML) layers of one or more ML models based on the plurality of training instances; and subsequent to training the plurality of ML layers based on the plurality of training instances: causing a voice bot to utilize the plurality of ML layers to process one or more ASR speech hypotheses that are generated while the voice bot is conducting a corresponding conversation to determine the unique personal identifier.

[0139] These and other implementations of the technology disclosed herein can each optionally include one or more of the following features.

[0140] In some implementations, training the plurality of ML layers based on a given training instance of the plurality of training instances can include processing, using the plurality of ML layers, at least one ASR speech hypothesis for the unique person identifier to generate a corresponding predicted measure for each of the alphanumeric characters included in the at least one ASR speech hypothesis for the unique person identifier; comparing the corresponding predicted measure for each of the alphanumeric characters included in the unique person identifier to a corresponding ground truth measure for each of the alphanumeric characters included in a corresponding ground truth output corresponding to the unique person identifier to generate one or more losses; and causing respective weights of one or more of the plurality of ML layers to be updated based on one or more of the losses.

[0141] In some implementations, the method can further include, after training the plurality of ML layers based on the plurality of training instances and before causing the voice bot to utilize the plurality of ML layers to process a unique person identifier encountered while the voice bot is conducting a corresponding conversation, further training the plurality of ML layers utilizing the voice bot-human simulator.

[0142] In some versions of those implementations, further training the plurality of ML layers utilizing the voice bot-human simulator can include, for a given training instance of the plurality of training instances and before the predicted unique person identifier corresponds to a corresponding ground truth output corresponding to the unique person identifier: accessing a voice bot-human simulator of a corresponding conversation between a simulated voice bot and the given human; processing, using the plurality of ML layers, at least one ASR speech hypothesis for the unique person identifier to generate a predicted unique person identifier and a corresponding predicted measure for each of the alphanumeric characters; processing, using the simulated voice bot of the voice bot-human simulator, the predicted unique person identifier and the corresponding predicted measure for the predicted unique person identifier to generate a simulated prompt of a simulated human; processing, using the simulated human of the voice bot-human simulator, the simulated prompt to generate a simulated response from the simulated human responsive to the simulated prompt; and processing, using the plurality of ML layers, the simulated response from the given human to refine the predicted unique person identifier.

[0143] In some further versions of those implementations, the simulated prompt can include a clarification request requesting clarification of one or more of the alphanumeric characters of the predicted unique person identifier. In still further versions of those implementations, the simulated response responsive to the clarification request can include a clarification of one or more of the alphanumeric characters of the predicted unique person identifier. In yet still further versions of those implementations, refining the predicted unique person identifier can include updating the corresponding prediction measure for one or more of the alphanumeric characters based on the clarification received responsive to the clarification request.

[0144] In some further versions of those implementations, causing the voice bot to utilize the plurality of ML layers to process one or more of the ASR speech hypotheses generated while the voice bot is conducting the corresponding conversation in determining the unique person identifier can also be after further training of the plurality of ML layers.

[0145] In some further versions of those implementations, the method can also include obtaining an intent of the simulated voice bot associated with the simulated prompt. Processing the simulated response from the given human using the plurality of ML layers to refine the predicted unique person identifier can also include processing the intent of the simulated voice bot along with the simulated response to refine the predicted unique person identifier.

[0146] In some implementations, obtaining the plurality of training instances can include generating one or more of the plurality of training instances based on a corresponding previously conducted conversation including at least one human. Generating one or more of the plurality of training instances based on a corresponding previously conducted conversation including at least one human can include determining audio data capturing a human-provided utterance of the at least one human during the previously conducted conversation, obtaining at least one ASR speech hypothesis for the unique person identifier, the at least one ASR speech hypothesis generated during the previously conducted conversation based on processing the audio data capturing the unique person identifier using an ASR model, utilizing the at least one ASR speech hypothesis generated based on the audio data capturing the unique person identifier for the at least one human as a training instance input, and utilizing a corresponding ground truth output corresponding to the unique person identifier as a training instance output.

[0147] In some versions of those implementations, the corresponding previously conducted conversation including the at least one human can be between the at least one human and another human.

[0148] In some versions of those implementations, the corresponding previously conducted conversation including the at least one human can be between the at least one human and an instance of the voice bot.

[0149] In some implementations, obtaining the plurality of training instances can include synthesizing one or more of the plurality of training instances based on a corresponding unique person identifier stored in one or more databases. Synthesizing one or more of the plurality of training instances based on a corresponding unique person identifier stored in one or more of the databases can include: accessing one or more of the databases to identify a unique person identifier; generating a plurality of tokens based on the unique person identifier; generating synthetic text based on the plurality of tokens corresponding to the unique person identifier; generating at least one ASR speech hypothesis for the unique person identifier based on the synthetic text; utilizing the at least one ASR speech hypothesis generated based on the synthetic text as a training instance input; and utilizing the unique person identifier as a training instance output.

[0150] In some versions of those implementations, generating the plurality of tokens corresponding to the unique person identifier can include identifying a plurality of n-grams included in the unique person identifier identified from one or more of the databases; and generating the plurality of tokens based on a distribution for the plurality of n-grams. In some further versions of those implementations, generating the synthetic text corresponding to at least the unique person identifier can include one or more of: injecting one or more fillers into the plurality of tokens; or injecting a phonetic spelling for one or more of the alphanumeric characters included in the unique person identifier. In some further versions of those implementations, generating the at least one ASR speech hypothesis for the unique person identifier can include replacing one or more n-grams of the synthetic text with one or more corresponding homophonic n-grams. In some further versions of those implementations, generating the at least one ASR speech hypothesis for the unique person identifier can include replacing one or more alphanumeric characters of the synthetic text with one or more corresponding homophonic alphanumeric characters.

[0151] In some implementations, the unique person identifier can be one or more of: an email address, a physical address, a username, a password, a name of an entity, or a domain name.

[0152] Further, some implementations include one or more processors (e.g., central processing units (CPUs), graphics processing units (GPUs), and / or tensor processing units (TPUs)) of one or more computing devices, where the one or more processors are operable to execute instructions stored in associated memory, and where the instructions are configured to cause performance of any of the preceding methods. Some implementations also include one or more non-transitory computer- readable storage media storing computer instructions executable by one or more processors to perform any of the preceding methods. Some implementations also include a computer program product comprising instructions executable by one or more processors to perform any of the preceding methods.

[0153] It should be appreciated that all combinations of the above- described concepts and additional concepts described in greater detail herein are contemplated, to the extent not already described. For instance, a state of the art implementation of a claimed subject matter at the end of this disclosure is also contemplated.

Claims

1. A method implemented by one or more processors, the method comprising: Receive audio data capturing human spoken words, which are received by a voice robot during a corresponding dialogue between the human and the voice robot; The audio data is processed using an automatic speech recognition model to generate multiple automatic speech recognition speech hypotheses; as well as In response to the prediction that the spoken utterance includes a unique personal identifier, the unique personal identifier comprising a unique alphanumeric character sequence that is personal to the human being: Multiple machine learning layers of one or more machine learning models are used to process one or more of the multiple automatic speech recognition speech hypotheses to generate one or more candidate unique personal identifiers, each of the one or more candidate unique personal identifiers including a corresponding predictive measurement associated with one or more corresponding alphanumeric characters for each of the one or more candidate unique personal identifiers. One or more given alphanumeric characters are selected from one or more of the corresponding alphanumeric characters based on a corresponding prediction measurement associated with one or more corresponding alphanumeric characters used for each of the one or more candidate unique personal identifiers; Whether to generate a prompt is determined by comparing the corresponding predictive measure associated with a subset of one or more of the given alphanumeric characters with a predictive measure threshold generated based on the corresponding predictive measure generated by the plurality of machine learning layers of the plurality of automatic speech recognition speech hypotheses and using one or more of the machine learning models; the prompt includes a clarification request requesting clarification of the subset of one or more of the given alphanumeric characters. In response to the determination that the corresponding predicted measurement associated with the subset of one or more of the given alphanumeric characters fails to meet the predicted measurement threshold: Generate the prompt, which includes the clarification request for clarification of a subset of one or more of the given alphanumeric characters; as well as This allows the prompt to be provided for presentation to the human. as well as In response to predicting that one or more given alphanumeric characters correspond to the unique personal identifier of the human: The unique personal identifier, comprising one or more given alphanumeric characters, is used to facilitate corresponding dialogue between the voice robot and the human.

2. The method according to claim 1, further comprising: In response to the prompt being provided for presentation to the human: Receive additional audio data capturing additional spoken words of the human, which are received by the voice robot during the corresponding dialogue; The automatic speech recognition model is used to process the additional audio data to generate multiple additional automatic speech recognition speech hypotheses. as well as The plurality of machine learning layers are used to process one or more of the plurality of additional automatic speech recognition speech hypotheses to refine one or more of the given alphanumeric characters.

3. The method according to claim 2, wherein, Refining one or more of the given alphanumeric characters includes: Based on the clarification received in response to the clarification request, the corresponding prediction measurement for the one or more given alphanumeric characters predicted to correspond to the unique personal identifier is updated.

4. The method according to claim 2, further comprising: Before the one or more given alphanumeric characters are predicted to correspond to the unique personal identifier of the human: One or more additional corresponding prompts are generated based on the corresponding prediction measurement associated with one or more additional corresponding subsets of the given alphanumeric characters, each of the one or more additional corresponding prompts including a corresponding additional clarification request requesting additional clarification of the additional corresponding subsets of the given alphanumeric characters; as well as This causes one or more of the corresponding additional prompts to be provided for presentation to the human.

5. The method according to claim 4, wherein, Predicting that one or more given alphanumeric characters correspond to the unique personal identifier of the human includes: Determine that the corresponding predicted measurement associated with each of the given alphanumeric characters satisfies a threshold.

6. The method according to claim 1, wherein, Generating the prompt that includes the clarification request includes: Identify the subset of one or more of the given alphanumeric characters associated with the corresponding predicted measurement that failed to meet the threshold; and Generate a clarification request for a subset of one or more of the given alphanumeric characters.

7. The method according to claim 1, wherein, Predicting that one or more given alphanumeric characters correspond to the unique personal identifier of the human includes: Determine that the corresponding prediction measurement associated with each of the one or more given alphanumeric characters used for each of the one or more candidate unique personal identifiers satisfies a threshold.

8. The method according to claim 1, wherein, The predicted spoken utterances include the unique personal identifier, which includes: Based on the synthetic speech audio data, including synthetic speech previously provided for presentation to the human by the voice robot during the corresponding dialogue, it is predicted that the audio data will include the unique personal identifier.

9. The method according to claim 1, wherein, The predicted spoken utterances include the unique personal identifier, which includes: The spoken utterance is predicted to include the unique personal identifier based on one or more of the plurality of automatic speech recognition speech hypotheses generated using the automatic speech recognition model.

10. The method according to claim 1, wherein, Processing one or more of the plurality of automatic speech recognition speech hypotheses to generate the one or more candidate unique personal identifiers includes: The plurality of machine learning layers iteratively process each of the plurality of automatic speech recognition speech hypotheses to iteratively generate a probability tree for the unique personal identifier. The probability tree includes a plurality of nodes and a plurality of edges, each of the plurality of nodes corresponding to one or more of the corresponding alphanumeric characters for each of the one or more corresponding alphanumeric characters, each of the plurality of nodes being associated with a corresponding predictive measurement for each of the one or more corresponding alphanumeric characters, and each of the plurality of nodes being connected through one or more of the plurality of edges. The generation of the one or more candidate unique personal identifiers is based on the probability tree.

11. The method according to claim 10, wherein, The probability tree is constrained by the plurality of automatic speech recognition speech hypotheses.

12. The method according to claim 11, wherein, The probability tree is constrained by multiple unique personal identifiers stored in one or more databases.

13. The method according to claim 1, wherein, The unique personal identifier is one or more of the following: email address, physical address, username, password, entity name, and domain name.

14. The method according to any one of claims 1-13, further comprising: Obtain the intention of the voice robot regarding a portion of the corresponding dialogue between the voice robot and the human; as well as The process of using the plurality of machine learning layers to process one or more of the plurality of automatic speech recognition speech hypotheses to generate one or more of the candidate unique personal identifiers also includes using the plurality of machine learning layers to process the intention of the voice robot to generate one or more of the candidate unique personal identifiers.

15. The method according to claim 14, wherein, The intent of the voice robot includes one or more of the following: The human being is requested to provide the unique personal identifier. The human is requested to spell the unique personal identifier; and The human is requested to provide clarification on one or more of the given alphanumeric characters.

16. The method according to claim 1, wherein, The corresponding dialogue between the voice robot and the human is a corresponding telephone conversation between the voice robot and the human, wherein the corresponding telephone conversation between the voice robot and the human is facilitated by the unique personal identifier comprising one or more given alphanumeric characters.

17. A method implemented by one or more processors, the method comprising: Receive audio data capturing human spoken words, which are received by a voice robot during a corresponding dialogue between the human and the voice robot; The audio data is processed using an automatic speech recognition model to generate multiple automatic speech recognition speech hypotheses; In response to the prediction that the spoken utterance includes a unique personal identifier, the unique personal identifier comprising a unique alphanumeric character sequence that is personal to a given user: Multiple machine learning layers of one or more machine learning models are used to process one or more of the multiple automatic speech recognition speech hypotheses to generate one or more candidate unique personal identifiers, each of the one or more candidate unique personal identifiers including a corresponding predictive measurement associated with one or more corresponding alphanumeric characters; as well as Based on the corresponding prediction measurement associated with one or more of the corresponding alphanumeric characters for each of the one or more candidate unique personal identifiers, select one or more given alphanumeric characters from the one or more of the corresponding alphanumeric characters; Before a given candidate unique personal identifier is predicted to correspond to the unique personal identifier: Whether to generate a corresponding prompt is determined by comparing a corresponding prediction measurement associated with a corresponding subset of one or more of the given alphanumeric characters with a corresponding prediction measurement threshold generated based on the corresponding prediction measurement generated by the plurality of machine learning layers of one or more of the plurality of automatic speech recognition speech hypotheses and using one or more of the machine learning models; the corresponding prompt includes a corresponding clarification request requesting clarification of the corresponding subset of one or more of the given alphanumeric characters. In response to the determination that the corresponding prediction measurement associated with one or more of the given alphanumeric characters fails to meet the corresponding prediction measurement threshold: Generate the corresponding prompt, the corresponding prompt including the corresponding clarification request for clarification of the corresponding subset of one or more of the given alphanumeric characters for the given candidate unique personal identifier; This allows the corresponding prompts to be provided for presentation to the human. as well as The given candidate unique personal identifier is refined based on the processing of additional audio data that captures the corresponding additional spoken words of the human and responds to the corresponding cues. as well as In response to predicting that the given candidate unique personal identifier corresponds to the unique personal identifier: The given candidate unique personal identifier, which includes one or more given alphanumeric characters, is used to facilitate the corresponding dialogue between the voice robot and the human.

18. A system comprising: One or more processors; as well as A memory for storing instructions, which, when executed, cause the one or more processors to perform the method according to any one of claims 1 to 17.

19. A non-transitory computer-readable storage medium storing instructions, which, when executed, cause one or more processors to perform the method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Network credential provisioning using audible commands

    US9730073B1