Automated assistant adapted for multiple age groups and / or vocabulary levels

By detecting the user's age and vocabulary level, the automated assistant system adjusts its behavior patterns, solving the problem of poor adaptability of existing systems, improving interaction efficiency and resource utilization, and providing responses that better meet user needs.

CN118471216BActive Publication Date: 2026-06-02GOOGLE LLC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GOOGLE LLC
Filing Date
2019-04-16
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing automated assistant systems struggle to adapt to users of different age groups and vocabulary levels, resulting in low interaction efficiency, being particularly unfriendly to children, and consuming excessive computing resources.

Method used

By detecting the user's age range and vocabulary level, the automated assistant system can switch to the corresponding mode and adjust the intent recognition, parsing, and response methods, including speech recognition, grammar tolerance, and output methods, to adapt to the needs of different user groups.

Benefits of technology

It improves the efficiency of interaction with users of different age groups and vocabulary levels, reduces the consumption of computing resources, provides a more user-friendly response, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118471216B_ABST
    Figure CN118471216B_ABST
Patent Text Reader

Abstract

This application relates to automated assistants that accommodate multiple age groups and / or vocabulary levels. Described herein are techniques for enabling an automated assistant to adjust its behavior depending on a detected age range and / or "vocabulary level" of a user that is interfacing with the automated assistant. Data indicative of utterances of the user can be used to estimate one or more of an age range and / or vocabulary level of the user. The estimated age range / vocabulary level can be used to influence aspects of a data processing pipeline employed by the automated assistant. Aspects of the data processing pipeline that can be influenced by the age range / vocabulary level of the user can include one or more of automated assistant invocation, speech-to-text processing, intent matching, intent resolution, natural language generation, and / or text-to-speech processing. In some implementations, one or more tolerance thresholds associated with one or more of these aspects such as grammatical tolerance, vocabulary tolerance, etc. can be adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Case Analysis

[0002] This application is a divisional application of Chinese invention patent application 201980032199.7, filed on April 16, 2019. Technical Field

[0003] This application relates to an automated assistant that adapts to multiple age groups and / or vocabulary levels. Background Technology

[0004] People can engage in human-computer dialogue with interactive software applications referred to herein as “automatic assistants” (also known as “chatbots,” “interactive personal assistants,” “intelligent personal assistants,” “personal voice assistants,” “conversational agents,” etc.). For example, humans (who may be referred to as “users” when interacting with an automatic assistant) can use free-form natural language input to provide commands, queries, and / or requests (collectively referred to herein as “queries”), which may include spoken utterances that have been converted into text and then processed, and / or typed free-form natural language input.

[0005] Users can interact with automated assistants using various types of computing devices such as smartphones, laptops, and tablets. An "assistant device" is a computing device primarily or even exclusively designed to facilitate human-computer dialogue between the user and the automated assistant. A common example of an assistant device is a standalone interactive speaker that allows the user to verbally interact with the automated assistant, for example, by uttering a pre-defined phrase to activate the assistant, enabling it to respond to the user's subsequent words.

[0006] The focus on verbal interaction in assistant devices makes them particularly suitable for children. However, many features built into or otherwise accessible through commercially available automated assistants may not be suitable for children. For example, if a child asks if the Tooth Fairy is real, a typical automated assistant might reply based on an online document, "No, the Tooth Fairy is an imaginary character evoked by parents to encourage children to pull loose teeth." As another example, automated assistants can be configured to interact with independent agents, such as third-party apps, that allow users to order goods / services like pizza, movies, toys, etc.—a capability that could be used by a child who might not be able to judge all the consequences of their actions. Additionally, typical automated assistants are designed to interact with people who have a fully developed vocabulary. Rather than attempting to interpret a user's request based on a "best guess" about their intent, an automated assistant can request clarification and / or disambiguation if the user's input is unclear. Such long round trips can lead to excessive consumption of various computer and / or network resources (e.g., as a result of generating and rendering requests for clarification and / or processing the resulting input) and / or can be frustrating for children with limited vocabulary. Summary of the Invention

[0007] This paper describes techniques for enabling automated assistants to adjust their behavior based on the detected age range and / or “vocabulary level” of the user interacting with them. Thus, an automated assistant can use one mode, such as a “child mode,” when interacting with a child, and another mode, such as a “normal” or “adult” mode, when interacting with users not considered children (e.g., teenagers and older adults). In some implementations, the automated assistant may be able to switch between a series of modes, each associated with a specific age range (or, alternatively, a series of modes associated with multiple vocabulary levels). Various aspects of the automated assistant’s behavior can be influenced by the mode selected based on the user’s age range (or vocabulary level), such as (i) the recognition of the user’s intent, (ii) the parsing of the user’s intent, and (iii) how the results of parsing the user’s intent are output.

[0008] As described in more detail herein, various implementations enable automated assistants to generate responses to a variety of inputs that would not be parsable without the techniques described herein. As further described herein, various implementations alleviate the need for automated assistants to request clarification of various inputs, thereby conserving various computer and / or network resources that would otherwise be used to generate and render such requests for clarification and / or process further input in response to such requests. Furthermore, automated assistants configured with selected aspects of this disclosure can facilitate more effective interaction with users in situations where users might typically have difficulty interacting with the assistant device. This may occur, for example, when a user's voice is less clear than that of a typical user of such a device (e.g., when a subsequent user is a young child, has a disability affecting the clarity of their speech, and / or is a non-native speaker).

[0009] While the examples in this article primarily concern determining a user's age range and acting accordingly, this is not intended to be restrictive. Various user characteristics, such as gender and location, can be detected and used to influence the behavior of automated assistants. For example, in some implementations, a user's vocabulary level, rather than their age range, can be estimated so that younger users with more advanced vocabulary will be appropriately engaged with by the automated assistant (ensuring that children with advanced vocabulary are not "rejected" by the automated assistant). Similarly, older users with adult-like speech patterns but limited vocabulary (e.g., due to learning disabilities, non-native languages, etc.) can be engaged with by the automated assistant in a way that helps their limited vocabulary. Furthermore, when the automated assistant communicates with the user at the same vocabulary level used by the user, this allows the user to know that he or she has "made" his or her language "intelligible" when communicating with the automated assistant. This may encourage the user to speak more naturally.

[0010] An automated assistant equipped with selected aspects of this disclosure can be configured to enter various age-related modes based on various signals, such as "standard" (e.g., suitable for adults) and "child mode" (e.g., suitable for young children). In some implementations, a parent or other adult (e.g., guardian, teacher) can manually switch the automated assistant to child mode, for example, on demand and / or during predetermined time intervals when the child is likely to interact with the automated assistant.

[0011] Additionally or alternatively, in some embodiments, the automated assistant may automatically detect (e.g., predict, estimate) the user's age range, for example, based on characteristics of the user's voice such as rhythm, pitch, phonemes, vocabulary, grammar, pronunciation, etc. In some implementations, a machine learning model may be trained to generate an output indicating the user's predicted age based on data indicative of the user's spoken words (e.g., audio recordings, feature vectors, embeddings). For example, a feedforward neural network may be trained, for instance, using training examples of audio samples labeled by the speaker's age (or age range), to generate multiple probabilities associated with multiple age ranges (e.g., 2-4 years: 25%; 5-6 years: 38%; 7-9 years: 22%; 10-14 years: 10%; under 14 years: 5%). In various implementations, the age range with the highest probability may be selected as the user's actual age range. In other implementations, other types of artificial intelligence models and / or algorithms may be employed to estimate the user's age.

[0012] Additionally or alternatively, in some implementations, for example, in response to configuration by one or more users, the automated assistant may employ voice recognition to distinguish and identify individual speakers. For instance, a household may configure one or more assistant devices in its home to recognize the voices of all family members, such that each member's profile can be active when that member interacts with the automated assistant. In some such implementations, each family member's profile may include their age (e.g., date of birth), enabling the automated assistant to determine the speaker's age when identifying them.

[0013] Once the user / speaker's age and / or age range (or vocabulary level) is determined, it can influence various aspects of how the automated assistant operates. For example, in some implementations, when the speaker is determined to be a child (or another user with limited language skills), the automated assistant may be less stringent about which utterances would qualify as calling phrases compared to when the speaker is determined to be an adult or other skilled speaker. In some implementations, one or more on-device models (e.g., trained AI models) can be used locally on the client device, for example, to detect the pre-defined calling phrases. If the speaker is detected to be a child, in some implementations, a calling model specifically designed for children can be employed. Additionally or alternatively, if a single calling model is used for all users, one or more thresholds that must be met to classify a user's utterance as an appropriate calling phrase can be lowered, for example, making it possible to still classify an attempt by a child to mispronounce a phrase as an appropriate calling phrase.

[0014] As another example, the user's estimated age range and / or vocabulary level can be used when detecting the user's intent. In various implementations, one or more candidate "query understanding models," each associated with a specific age range, may be available for use by the automated assistant. Each query understanding model may be used to determine the user's intent, but may operate differently from the others. A "standard" query understanding model designed for adults may have a specific "grammatical tolerance," for example, lower than the grammatical tolerance associated with "children's" query understanding models. For example, a children's query understanding model may have such a grammatical tolerance (e.g., a minimum confidence threshold) that it agrees to give the automated assistant considerable leeway to "guess" the user's intent, even when the user's grammar / vocabulary is imperfect, as is often the case with young children. Conversely, when the automated assistant chooses a "standard" query understanding model, it may have a lower grammatical tolerance, thus allowing it to seek disambiguation and / or clarification from the user more quickly, rather than "guessing" or selecting a relatively low-confidence candidate intent as the user's actual intent.

[0015] Additionally or alternatively, in some implementations, the query understanding model can influence speech-to-text (“STT”) processing adopted by or on behalf of an automated assistant. For example, a conventional STT processor might fail to process a child’s “giggy-like meowing” utterance. Instead of generating a speech recognition output (e.g., text) that tracks the speech made by the user, the STT process could simply reject the utterance (e.g., because the word is not identified in the dictionary), and the automated assistant could say something like, “I’m sorry, I didn’t catch that.” However, an STT processor configured with selected aspects of this disclosure can be more tolerant of mispronunciations and / or below-average grammar / vocabulary in response to determining that the speaker is a child. When a child speaker is detected, the STT processor generates a speech recognition output (e.g., a textual explanation of the utterance, embeddings, etc.) that tracks the speech made by the user, even if the speech recognition output includes some words / phrases not found in the dictionary. Similarly, the natural language understanding module can use a child-centered query understanding model to interpret the text "giggy" as "kitty," but if an adult-centered query understanding model is used, the word "giggy" may not be interpreted.

[0016] In some implementations, the STT processor can use a dedicated parsing module when processing simplified grammar used by children. In some such implementations, the dedicated parsing module can use STT models trained using existing techniques, but with noisier (e.g., incomplete, grammatically ambiguous, etc.) training data. For example, in some implementations, the training data may be at least partially generated from data originally generated from adults, but transformed according to a model of the child's vocal range and / or by adding noise to the text input to represent incomplete grammar.

[0017] Generally speaking, an automated assistant configured with selected aspects of this disclosure can be more proactive than a conventional automated assistant when interacting with a child. For example, and as previously described, it may be more willing to "guess" what the child's intentions are. Additionally, the automated assistant may be less strict about requiring the use of a phrase when it detects a child as the speaker. For example, in some implementations, if the child calls an animal's name, the automated assistant may, upon determining that the speaker is a child, abandon the requirement for the child to utter the phrase and may instead mimic the animal's sound. Additionally or alternatively, the automated assistant may, for example, attempt to "teach" the child proper grammar, pronunciation, and / or vocabulary in response to grammatically incorrect and / or mispronounced speech.

[0018] Regarding the parsing of user intent, various actions and / or information may not be suitable for children. Therefore, in various embodiments, the automated assistant can determine whether the user's intent is parsable based on the user's predicted age range. For example, if the user is determined to be a child, the automated assistant can limit the online corpus of data it can use to retrieve information in response to the user's request, for example, by limiting it to a "whitelist" of child-friendly websites and / or a "blacklist" of non-child-friendly websites. The same might apply to music, for example. If a child says "play music!", the automated assistant can limit the music it plays to a library of child-friendly music, rather than an adult-centric library that typically includes music for older individuals. The automated assistant may also not require the child user to specify a playlist or artist and can simply play music appropriate for the user's detected age. Conversely, an adult's request to "play music" might prompt the automated assistant to seek additional information about what music to play. In some such implementations, the volume can also be adjusted when the automated assistant is in child mode, for example, to have a lower limit than would otherwise be available in an adult setting. As another example, various actions, such as ordering goods / services through third-party apps, may not be suitable for children. Therefore, when an automated assistant determines that it is engaging with a child, it can refuse to perform actions that might involve spending money or facilitating online interactions with strangers.

[0019] Regarding the output of the parsed user intent, in various embodiments, the automated assistant may select a given voice synthesis model from a plurality of candidate voice synthesis models that is associated with a predetermined age group predicted for the user. For example, the default voice synthesis model adopted by the automated assistant may be an adult voice speaking at a relatively fast pace (e.g., similar to a real-life conversation between adults). Conversely, the voice synthesis model adopted by the automated assistant when interacting with a child user may be a cartoon character's voice, and / or may speak at a relatively slow pace.

[0020] Additionally or alternatively, in some implementations, the automated assistant may output more or less detail depending on the user's predicted age. This can be achieved, for example, by providing multiple natural language generation models, each tailored to a specific age group. As an example, when interacting with a user determined to be between two and four, the automated assistant may employ a suitable natural language generation model to enable it to use simple words and short sentences. As the detected age of the speaker increases, the vocabulary used by the automated assistant may grow (e.g., depending on the natural language generation model it chooses), causing it to use longer sentences and more complex or advanced words. In some embodiments, the automated assistant may provide output encouraging child users to speak in more complete sentences, suggesting alternative words, etc. In some implementations, it is not typically required that words and / or phrases explained to adults be explained more fully by the automated assistant when interacting with children.

[0021] Additionally or alternatively, in some implementations, the natural language generation (“NLG”) template may include logic that specifies providing one natural language output when the user is estimated to be in a first age range, and another natural language output when the user is estimated to be in a second age range. Thus, children hear output tailored to them (e.g., using child-appropriate words and / or slang), while adults hear different outputs.

[0022] As another example, a common use case for automated assistants is summarizing web pages (or portions thereof). For instance, users often ask automated assistants random questions that might be answered using entries from online encyclopedias. Automated assistants can employ various techniques to summarize relevant portions of a web page into coherent answers. Using the techniques described herein, automated assistants can consider the user's estimated age range and / or vocabulary level when summarizing relevant portions of a web page. For example, a child might only receive advanced and / or easily understood concepts described on a web page, while an adult might receive more detail and / or description. In some embodiments, if the summary is not only abstract but also generative, it can even be powered and / or combined with an automated machine translation system (e.g., an "adult English to simple English" translation system).

[0023] In some implementations, the automated assistant can be configured to report a child's grammatical and / or vocabulary progress. For example, when the automated assistant determines that it is engaging with an adult, or especially when it recognizes the parent's voice, the adult / parent user can ask the automated assistant about the progress of one or more children interacting with it. In various implementations, the automated assistant can provide a variety of data in response to such inquiries, such as which words or syllables the child tends to mispronounce or struggles with, whether stuttering tendencies are detected in the child, what questions the child has asked, how the child is progressing in interactive games, and so on.

[0024] In some implementations, a method is provided executed by one or more processors, the method comprising: receiving spoken utterance from a user at one or more input components of one or more client devices; applying data instructing the spoken utterance across a trained machine learning model to generate output; determining, based on the output, that the user falls into a predetermined age group; selecting, from a plurality of candidate query understanding models, a given query understanding model associated with the predetermined age group; using the given query understanding model to determine the user's intent; determining, based on the predetermined age group, that the user's intent is parsable; parsing the user's intent to generate response data; and outputting the response data at one or more output components of one or more client devices.

[0025] In various implementations, the multiple candidate query understanding models may include at least one candidate query understanding model with a syntax tolerance different from that of a given query understanding model. In various implementations, the data indicating spoken utterance may include audio recordings of the user's utterance, and a machine learning model is trained to generate an output indicating the user's age based on one or more phonemes contained in the audio recording.

[0026] In various implementations, the method may further include: selecting a given natural language generation model associated with a predetermined age group from a plurality of candidate natural language generation models, wherein the selected given natural language output model is used to generate response data. In various implementations, the plurality of candidate natural language generation models may include at least one candidate natural language generation model that uses a more complex vocabulary than that used by the given natural language output model.

[0027] In various implementations, the method may further include: selecting a given speech synthesis model associated with a predetermined age group from a plurality of candidate speech synthesis models, wherein output response data is processed using the given speech synthesis model. In various implementations, a given query understanding model may be applied to perform speech-to-text processing of spoken utterances. In various implementations, a given query understanding model may be applied to perform natural language understanding of speech recognition output generated from spoken utterances.

[0028] In another aspect, a method executed by one or more processors is provided, the method comprising: receiving spoken utterance from a user at one or more input components of one or more client devices; applying data indicative of the spoken utterance across a trained machine learning model to generate output; determining, based on the output, a given lexical level in which the user falls among a plurality of predetermined lexical levels; selecting, from a plurality of candidate query understanding models, a given query understanding model associated with the given lexical level; using the given query understanding model to determine the user's intent; determining, based on the predetermined lexical level, that the user's intent is parsable; parsing the user's intent to generate response data; and outputting the response data at one or more output components of one or more client devices.

[0029] Additionally, some implementations include one or more processors of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in associated memory, and wherein said instructions are configured to cause any of the aforementioned methods to be performed. Some implementations also include one or more non-transitory computer-readable storage media storing computer instructions executable by one or more processors to perform any of the aforementioned methods.

[0030] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are considered part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered part of the subject matter disclosed herein. Attached Figure Description

[0031] Figure 1 This is a block diagram of an example environment in which the implementation methods disclosed herein can be carried out.

[0032] Figure 2 Exemplary processing flows illustrating various aspects of this disclosure according to various embodiments are described.

[0033] Figure 3A and 3B Describe example dialogues between a user and an automated assistant according to various implementation methods.

[0034] Figure 4A and Figure 4B Describe example dialogues between a user and an automated assistant according to various implementation methods.

[0035] Figure 5A and 5B Describe example dialogues between a user and an automated assistant according to various implementation methods.

[0036] Figure 6 A flowchart illustrating an example method according to an embodiment disclosed herein.

[0037] Figure 7 The diagram illustrates an example architecture for a computing device. Detailed Implementation

[0038] Now go to Figure 1 The illustration shows an example environment in which the techniques disclosed herein can be implemented. The example environment includes one or more client computing devices 106. Each client device 106 can execute a corresponding instance of an automated assistant client 118. Communication can be coupled to the client device 106 via one or more local area networks and / or wide area networks (e.g., the Internet), typically indicated at 114. 1-N One or more cloud-based automated assistant components 119, such as a natural language understanding engine 135, are implemented on one or more computing systems (collectively referred to as “cloud” computing systems).

[0039] In various implementations, an instance of the automated assistant client 108, through its interaction with one or more cloud-based automated assistant components 119, can form what appears from the user's perspective as a logical instance of the automated assistant 120, with which the user can engage in human-computer dialogue. Figure 1 An instance of such an automated assistant 120 is depicted with a dashed line. Therefore, it should be understood that each user interacting with the automated assistant client 108 running on client device 106 can actually interact with his or her own logical instance of automated assistant 120. For brevity and simplicity, the term "automated assistant" used herein to "serve" a particular user will refer to a combination of the automated assistant client 108 running on the user-operated client device 106 and one or more cloud-based automated assistant components 119 (which may be shared among multiple automated assistant clients 108). It should also be understood that in some implementations, automated assistant 120 can respond to requests from any user, regardless of whether that particular instance of automated assistant 120 is actually "serving" that user.

[0040] One or more client devices 106 may include one or more of the following: desktop computing devices, laptop computing devices, tablet computing devices, mobile phone computing devices, computing devices of a user's vehicle (e.g., in-vehicle communication systems, in-vehicle entertainment systems, in-vehicle navigation systems), independent interactive speakers, smart appliances such as smart TVs (and / or independent TVs equipped with networked electronic dog features with automatic assistant capabilities), and / or wearable devices of the user including computing devices (e.g., watches of users with computing devices, glasses of users with computing devices, virtual or augmented reality computing devices). Additional and / or alternative client computing devices may be provided.

[0041] As described in more detail herein, the automated assistant 120 engages in a human-computer dialogue session with one or more users via user interface input and output devices of one or more client devices 106. In some embodiments, the automated assistant 120 may engage in a human-computer dialogue session with a user in response to user interface input provided by the user via one or more user interface input devices of one of the client devices 106. In some of those embodiments, the user interface input is explicitly directed to the automated assistant 120. For example, the user may utter a predetermined invocation phrase, such as “OK, Assistant” or “Hey, Assistant”, to cause the automated assistant 120 to begin actively listening.

[0042] In some implementations, the automated assistant 120 may participate in a human-computer dialogue session in response to user interface input, even when the user interface input is not explicitly directed to it. For example, the automated assistant 120 may respond to certain terms present in the user interface input and / or examine the content of the user interface input based on other prompts and participate in the dialogue session. In many implementations, the automated assistant 120 may utilize speech recognition to convert speech from the user into text and respond to the text, for example, by providing search results, general information, and / or taking one or more responsive actions (e.g., playing media, starting a game, ordering food, etc.). In some implementations, the automated assistant 120 may additionally or alternatively respond to speech without converting the speech into text. For example, the automated assistant 120 may convert the speech input into an embedding, into an entity representation (indicating one or more entities present in the speech input), and / or other “non-textual” representations, and operate on such non-textual representations. Therefore, the implementations described herein that operate based on text converted from speech input may additionally and / or alternatively operate directly on the speech input and / or other non-textual representations of the speech input.

[0043] Each of the client computing device 106 and the computing device operating the cloud-based automated assistant component 119 may include one or more memories for storing data and software applications, one or more processors for accessing data and executing applications, and other components that facilitate communication over a network. The operations performed by the client computing device 106 and / or by the automated assistant 120 may be distributed across multiple computer systems. The automated assistant 120 may be implemented, for example, as a computer program running on one or more computers in one or more locations coupled to each other via a network.

[0044] As described above, in various embodiments, the client computing device 106 can operate the automated assistant client 108. In various embodiments, the automated assistant client 108 may include a voice capture module 110, a proficiency detector 111, and / or an invocation module 112. In other embodiments, one or more aspects of the voice capture module 110, the proficiency detector 111, and / or the invocation module 112 may be implemented separately from the automated assistant client 108, for example, through one or more cloud-based automated assistant components 119. In various embodiments, the voice capture module 110 may interface with hardware such as a microphone (not described) to capture audio recordings of user speech. Various types of processing can be performed on the audio recording for various purposes, as described below.

[0045] In various implementations, the proficiency detector 111, which can be implemented using any combination of hardware or software, can be configured to analyze audio recordings captured by the voice capture module 110 to make one or more determinations based on the user's apparent speech proficiency. In some implementations, these determinations may include predicting or estimating the user's age, such as classifying the user into one of several age ranges. Additionally or alternatively, these determinations may include predicting or estimating the user's vocabulary level, or classifying the user into one of several vocabulary levels. As will be described below, the determinations made by the proficiency detector 111 can be used, for example, by various components of the automated assistant 120 to accommodate users of multiple different age groups and / or vocabulary levels.

[0046] The proficiency detector 111 can employ various techniques to determine a user's speaking proficiency. For example, in Figure 1 In this implementation, the proficiency detector 111 is communicatively coupled to a proficiency model database 113 (which may be integrated with and / or remotely hosted from the client device 106, such as in the cloud). The proficiency model database 113 may include one or more artificial intelligence (or machine learning) models trained to generate outputs indicative of a user's speech proficiency. In various implementations, the artificial intelligence (or machine learning) models may be trained to generate outputs indicative of a user's age based on one or more phonemes contained in an audio recording, the pitch of the user's speech detected in the audio recording, articulation, etc.

[0047] As a non-limiting example, a neural network can be trained (and stored in database 113) such that an audio recording of a user's utterance, or a feature vector extracted from that audio recording, can be applied as input across the neural network. In various implementations, the neural network can generate multiple outputs, each corresponding to an age range and an associated probability. In some such implementations, the age range with the highest probability can be used as a prediction of the user's age range. Such a neural network can be trained using various forms of training examples, such as audio recordings (or feature vectors generated from them) labeled according to the user's corresponding age (or age range). When applying the training examples across the network, the difference between the generated outputs and the labels associated with the training examples can be used, for example, to minimize the loss function. The various weights of the neural network can then be tuned, for example, using standard techniques such as gradient descent and / or backpropagation.

[0048] Voice capture module 110 may be configured to capture a user's speech, for example, via a microphone (not depicted), as previously described. Additionally or alternatively, in some embodiments, voice capture module 110 may be further configured to convert the captured audio into text and / or other representations or embeddings, for example, using speech-to-text (“STT”) processing techniques. Additionally or alternatively, in some embodiments, voice capture module 110 may be configured to convert text into computer-synthesized speech, for example, using one or more speech synthesizers. However, because client device 106 may be relatively constrained in terms of computing resources (e.g., processor cycles, memory, battery, etc.), the local voice capture module 110 for client device 106 may be configured to convert a limited number of different spoken phrases—particularly phrases invoking the automation assistant 120—into text (or other forms, such as lower-dimensional embeddings). Additional voice input may be sent to the cloud-based automation assistant component 119, which may include a cloud-based TTS module 116 and / or a cloud-based STT module 117.

[0049] In some implementations, client device 106 may include invocation module 112, configured to determine whether a user's utterance qualifies as an invocation phrase to initiate a human-computer dialogue session with automated assistant 120. Once the user's age (or age range) is estimated by proficiency detector 111, in some implementations, invocation module 112 may analyze data indicative of the user's utterance, such as audio recordings or feature vectors (e.g., embeddings) extracted from the audio recordings in conjunction with the user's estimated age range. In some implementations, as the user's estimated age falls into one or more age ranges, such as those associated with young children, the threshold used by invocation module 112 to determine whether to invoke automated assistant 120 may be lowered. Therefore, when interacting with a young child, a phrase such as "OK assisa," which differs from the appropriate invocation phrase "OK assistant" but is slightly similar in pronunciation, may still be accepted as an invocation.

[0050] Additionally or alternatively, an on-device invocation model can be used by invocation module 112 to determine whether a utterance qualifies as an invocation. Such an on-device invocation model can be trained to detect common variations in invocation phrases. For example, in some implementations, training examples (e.g., one or more neural networks) can be used to train the on-device invocation model, each of which includes an audio recording (or extracted feature vector) of the child's utterance. Some training examples may include attempts by the child to utter an appropriate invocation phrase; such training examples can be positive training examples. Additionally, in some implementations, some training examples may include other utterances from the child that are not attempts to utter the invocation phrase. These other training examples can be used as negative training examples. The (positive and / or negative) training examples can be applied as input to the on-device invocation model to generate outputs. The outputs can be compared with labels associated with (e.g., positive or negative) training examples, and the on-device invocation model can be trained using the difference (error function), for example, using techniques such as gradient descent and / or backpropagation. In some implementations, if the on-device calling model indicates that a utterance is eligible as a call, but the confidence score associated with the output is relatively low (e.g., because the utterance was generated by a child who is prone to mispronouncing words), the low confidence score itself can be used, for example, by the proficiency detector 111 to estimate the child's age.

[0051] The cloud-based TTS module 116 can be configured to utilize the potentially greater computing resources of the cloud to convert text data (e.g., natural language responses formulated by the automated assistant 120) into computer-generated speech output. In some embodiments, the TTS module 116 can provide the computer-generated speech output to the client device 106, for example, by direct output using one or more speakers. In other embodiments, the text data (e.g., natural language responses) generated by the automated assistant 120 can be provided to the speech capture module 110, which can then convert the text data into locally output computer-generated speech. In some embodiments, the cloud-based TTS module 116 can be operatively coupled to a database 115 comprising multiple speech synthesis models. When interacting with users who have specific speaking abilities and / or belong to a specific age group, the automated assistant 120 can employ each speech synthesis model to generate computer speech simulating a specific type of speech, such as a man, woman, cartoon character, speaker with a specific accent, etc.

[0052] In some implementations, the TTS module 116 can employ a specific speech synthesis model on demand. For example, suppose the automated assistant 120 is typically invoked with the phrase "Hey Assistant". In some implementations, the model used to detect this phrase (e.g., the invocation model on the device described earlier) can be modified, for example, to make it more sensitive to "Hey, <entity>The system responds to any utterance such as "(Hey, <entity>)". In some such implementations, the requested <entity> can be used by the TTS module 116 to select the speech synthesis modality to be used. Thus, if the child invokes the auto-assistant by saying something like "Hey, Hypothetical Hippo", the speech synthesis modality associated with the entity "Hypothetical Hippo" can be used.

[0053] The cloud-based STT module 117 can be configured to utilize the potentially greater computing resources of the cloud to convert audio data captured by the speech capture module 110 into text, which can then be provided to the natural language understanding module 135. In various implementations, the cloud-based STT module 117 may employ one or more custom parsers and / or STT models (sometimes referred to herein as "query understanding models") specifically tailored to interpreting the speech of users such as children with limited and / or underdeveloped vocabulary and / or grammar.

[0054] For example, in some implementations, the cloud-based STT module 117 may be operatively coupled to one or more databases 118 storing multiple query understanding models. Each query understanding model may be configured for use with users of a specific age range. In some implementations, each query understanding model may include an artificial intelligence model (e.g., a neural network of various flavors) trained to generate text from speech based on audio input (or data indicating the audio input, such as phonemes and / or other features extracted into the feature vector of the audio input). In some such implementations, the training data for such a model may include audio recordings (or data indicating them) from adults, labeled with actual text, such as spoken speech, and containing injected noise. Additionally or alternatively, in some implementations, the training data may include audio recordings from children labeled with text in children's speech.

[0055] In some implementations, the cloud-based STT module 117 can convert the audio recording of the speech into one or more phonemes, and then convert the one or more phonemes into text. Additionally or alternatively, in some implementations, the STT module 117 may employ a state decoding graph. In some implementations, the STT module 117 can generate multiple candidate text interpretations of the user's utterance. Using a conventional automation assistant 120, the candidate text interpretation with the highest associated confidence score can be accepted as the user's free-form input, provided that the confidence score meets a certain threshold and / or has a confidence score sufficiently better than those associated with other candidate text interpretations. Otherwise, the automation assistant 120 may request clarification and / or disambiguation from the user. Using the STT module 117 configured with selected aspects of this disclosure, such a threshold can be lowered. Therefore, even if the user's utterance (e.g., a statement from a toddler) demonstrates grammatical / lexical deficiencies, the automation assistant 120 may be more likely to attempt to satisfy the user's intent.

[0056] The automated assistant 120 (particularly the cloud-based automated assistant component 119) may include the natural language understanding module 135, the aforementioned TTS module 116, the aforementioned STT module 117, and other components that will be described in more detail below. In some embodiments, one or more modules and / or modules of the automated assistant 120 may be omitted, combined, and / or implemented in components separate from the automated assistant 120. In some embodiments, to protect privacy, one or more components of the automated assistant 120, such as the natural language processor 122, the TTS module 116, the STT module 117, etc., may be implemented at least partially on the client device 106 (e.g., excluded from the cloud).

[0057] In some implementations, the automated assistant 120 responds to a human-computer dialogue session with the automated assistant 120 initiated by the client device 106. 1-N The automated assistant 120 generates response content from various user-generated inputs. The automated assistant 120 may (e.g., via one or more networks when disconnected from the user's client device) provide response content as part of a conversational session. For example, the automated assistant 120 may generate response content in response to free-form natural language input provided via client device 106. As used herein, free-form input is user-defined and not constrained to presenting a set of options for the user to choose from.

[0058] As used herein, a "conversational session" can include a logically self-contained exchange of one or more messages between a user and the automated assistant 120 (and in some cases, other human participants). The automated assistant 120 can distinguish between multiple conversational sessions with the user based on various signals such as the passage of time between sessions, changes in the user's context between sessions (e.g., location, before / during / after a scheduled meeting, etc.), detection of one or more interventional interactions between the user and the client device other than the conversation between the user and the automated assistant (e.g., the user temporarily switches applications, the user walks away and then returns later to a standalone voice-activated product), locking / sleeping of the client device between sessions, changes in the client device used to interact with one or more instances of the automated assistant 120, etc.

[0059] The natural language processor 122 of the natural language understanding module 135 processes natural language input generated by the user via client device 106 and can generate annotated output (e.g., in text form) for use by one or more other components of the automated assistant 120. For example, the natural language processor 122 can process free-form natural language input generated by the user via one or more user interface input devices of client device 1061. The generated annotated output includes one or more annotations to the natural language input and one or more (e.g., all) terms from the natural language input.

[0060] In some embodiments, the natural language processor 122 is configured to recognize and annotate various types of syntactic information in the natural language input. For example, the natural language processor 122 may include a lexical module that can separate individual words into morphemes and / or annotate morphemes, for example, by their categories. The natural language processor 122 may also include a part-of-speech tagger that is configured to annotate terms with their grammatical roles. For example, the part-of-speech tagger can tag a term with the part of speech of each term, such as "noun," "verb," ​​"adjective," "pronoun," etc. Furthermore, for example, in some embodiments, the natural language processor 122 may additionally and / or alternatively include a dependency parser (not described) configured to determine syntactic relations between terms in the natural language input. For example, the dependency parser can determine which terms modify other terms, the subject and verb of a sentence, etc. (e.g., a parse tree)—and can annotate such dependencies.

[0061] In some implementations, the natural language processor 122 may additionally and / or alternatively include an entity annotator (not described) configured to annotate entity references in one or more paragraphs, such as references to people (including, for example, literary figures, celebrities, public figures, etc.), organizations, locations (real and fictional), and so on. In some implementations, data about entities may be stored in one or more databases, such as a knowledge graph (not described). In some implementations, the knowledge graph may include nodes representing known entities (and in some cases, entity attributes), and edges connecting nodes and representing relationships between entities. For example, the "banana" node may (e.g., as a child) be connected to the "fruit" node, which may then (e.g., as a child) be connected to the "produce" and / or "food" nodes. As another example, a restaurant called "Hypothetical Café" may be represented by nodes that also include attributes such as its address, the types of food served, opening hours, contact information, etc. In one implementation, the "Hypothetical Café" node can be connected to one or more other nodes, such as the "restaurant" node, the "business" node, a node representing the city and / or state where the restaurant is located, etc., via edges (e.g., representing the child-to-parent relationship).

[0062] The entity annotator of the natural language processor 122 can annotate entity references at a higher granularity level (e.g., enabling the identification of all references to entity categories such as people) and / or a lower granularity level (e.g., enabling the identification of all references to a specific entity such as a specific person). The entity annotator may rely on the content of the natural language input to resolve specific entities and / or may optionally communicate with a knowledge graph or other entity database to resolve specific entities.

[0063] In some implementations, the natural language processor 122 may additionally and / or alternatively include a coreference parser (not described) configured to group or "cluster" references to the same entity based on one or more contextual cues. For example, a coreference parser may be used to resolve the term "there" in the natural language input "I liked Hypothetical Café lasttime we ate there." to "Hypothetical Café."

[0064] In some implementations, one or more components of the natural language processor 122 may rely on annotations from one or more other components of the natural language processor 122. For example, in some implementations, when annotating all references to a particular entity, the named entity annotator may rely on annotations from the coreference parser and / or dependency parser. Similarly, for example, in some implementations, when clustering references to the same entity, the coreference parser may rely on annotations from the dependency parser. In some implementations, when processing a particular natural language input, one or more components of the natural language processor 122 may use relevant previous inputs and / or other relevant data beyond the particular natural language input to determine one or more annotations.

[0065] The natural language understanding module 135 may also include an intent matcher 136, which is configured to determine the intent of a user participating in a human-computer dialogue session with the automation assistant 120. Although in Figure 1 While depicted separately from the natural language processor 122, in other implementations, the intent matcher 136 may be an integral part of the natural language processor 122 (or more generally, a pipeline including the natural language processor 122). In some implementations, the natural language processor 122 and the intent matcher 136 may together form the aforementioned "natural language understanding" module 135.

[0066] The intent matcher 136 can use various techniques to determine the user's intent, such as based on the output from the natural language processor 122 (which may include annotations and words from the natural language input). In some implementations, the intent matcher 136 may be able to access one or more databases 137, which include, for example, multiple mappings between grammars and response actions (or more generally, intents). In many cases, these grammars may be selected and / or learned over time and may represent the user's most common intents. For example, a grammar like "play" could be used. <artist>The phrase "(Play <Artist>)" maps to the intent of a response action that causes the music of <Artist> to be played on the user-operated client device 106. Another syntax, "[weather|forecast]today", can be matched with user queries such as "what's the weather today" and "what's the forecast for today?". In addition to or in lieu of the syntax, in some implementations, the intent matcher 136 may employ one or more trained machine learning models, alone or in combination with one or more syntaxes. These trained machine learning models may also be stored in one or more databases 137 and may be trained to identify intents, for example, by embedding data indicating the user's utterance into a reduced-dimensional space and then using techniques such as Euclidean distance, cosine similarity, etc., to determine which other embeddings (and therefore, intents) are closest.

[0067] For example in "play" <artist>As seen in the example syntax, some syntaxes have slots that can be filled with slot values ​​(or "parameters") (e.g., <artist>Slot values ​​can be determined in various ways. Often, the user will actively provide the slot value. For example, for the syntax "Order me a..." <topping>"Order me a pizza (with sauce)" is a common phrase in online gaming. Users might say "order me a sausage pizza," in which case... <topping>Automatically populated. Additionally or alternatively, if the user invokes syntax that includes slots to be populated with slot values, without the user actively providing the slot values, the Auto Assistant 120 can request those slot values ​​from the user (e.g., "What type of crust do you want on your pizza?").

[0068] In some implementations, the automated assistant 120 can facilitate (or "arrange") transactions between the user and the agent, which can be independent software processes that receive input and provide response output. Some agents can take the form of third-party applications that may or may not operate on a computing system separate from the computing system operating, for example, the cloud-based automated assistant component 119. One type of user intent that can be identified by the intent matcher 136 is to engage with a third-party application. For example, the automated assistant 120 can provide access to an application programming interface ("API") to a pizza delivery service. The user can invoke the automated assistant 120 and provide a command such as "I'd like to order a pizza." The intent matcher 136 can map this command to syntax that triggers the automated assistant 120 to engage with the third-party pizza delivery service (which may, in some cases, be added to a database 137 by the third party). The third-party pizza delivery service can provide the automated assistant 120 with a minimal list of slots that need to be populated to fulfill the pizza delivery order. The automatic assistant 120 can generate natural language output requesting parameters for the slot and provide it to the user (via the client device 106).

[0069] In various implementations, the intent matcher 136 may have access to a library of multiple sets of grammars and / or trained models, such as in database 137. Each set of grammars and / or models may be designed to facilitate interaction between the automated assistant 120 and users of a specific age range and / or vocabulary level. In some implementations, these grammars and / or models may be part of the aforementioned "query understanding model," supplementing or replacing the model described above with respect to STT module 117 (e.g., stored in database 118). Thus, in some implementations, the query understanding model may include components used during both STT processing and intent matching. In other implementations, the query understanding model may include components used only during STT processing. In still other implementations, the query understanding model may include components used only during intent matching and / or natural language processing. Any combination of these variations is contemplated herein. If a query understanding model is employed during both STT processing and intent matching, then in some implementations, databases 118 and 137 may be combined.

[0070] As an example, a first set of grammars and / or models (or "query comprehension models") stored in database 137 can be configured to facilitate interaction with very young children, such as toddlers under two years old. In some such implementations, this set of grammars / models can be relatively tolerant of errors in grammar, vocabulary, pronunciation, etc., and can, for example, do something like making animal noises when a child says an animal's name. Another set of grammars and / or models can be configured to facilitate interaction with children between two and four years old and / or others with limited vocabulary (e.g., users in the process of learning the language adopted by the automated assistant 120). Such a set of grammars and / or models can be slightly less tolerant of errors, but it can still be relatively lenient. Yet another set of grammars and / or models can be configured to facilitate interaction with the "next" age range and / or vocabulary, such as five to seven years old and / or intermediate speakers, and may be even less tolerant of errors. Another set of grammars and / or models can be configured to facilitate “normal” engagement with adults, older children, and / or other relatively skilled speakers—for such a set of grammars / models, tolerance for errors in grammar, vocabulary, and / or pronunciation can be relatively low.

[0071] The fulfillment module 124 can be configured to receive the predicted / estimated intent and associated slot values ​​(whether actively provided by the user or requested by the user) output by the intent matcher 136, and fulfill (or "parse") the intent. In various implementations, the fulfillment (or "parse") of the user's intent can result in various fulfillment information (also referred to as "response" information or data) being generated / obtained by the fulfillment module 124, for example. As will be described below, in some implementations, the fulfillment information can be provided to a natural language generator ("NLG" in some figures) 126, which can then generate natural language output based on the fulfillment information.

[0072] Fulfillment information can take various forms because intent can be fulfilled in various ways. Suppose a user requests purely informational, such as "Where were the outdoor shots of 'The Shining' filmed?" The user's intent can be determined as a search query, for example, by intent matcher 136. The intent and content of the search query can be provided to fulfillment module 124, which, as... Figure 1 As depicted, one or more search modules 150 can communicate with each other and are configured to search for response information in corpora of documents and / or other data sources (e.g., knowledge graphs, etc.). The fulfillment module 124 can provide the search module 150 with data indicating the search query (e.g., query text, dimensionality reduction embeddings, etc.). The search module 150 can provide response information, such as GPS coordinates or other more explicit information, such as "Timberline Lodge, Mt. Hood, Oregon". This response information can form part of the fulfillment information generated by the fulfillment module 124.

[0073] Additionally or alternatively, the fulfillment module 124 may be configured to, for example, receive the user's intent and any slot values ​​determined by the user or by other means (e.g., the user's GPS coordinates, user preferences, etc.) from the natural language understanding module 135 and trigger a response action. Response actions may include, for example, ordering goods / services, starting a timer, setting a reminder, initiating a phone call, playing media, sending a message, etc. In some such implementations, fulfillment information may include slot values ​​associated with fulfillment, confirmation responses (which in some cases may be selected from predetermined responses), etc.

[0074] In some implementations, and similar to other components such as STT module 117 and intent matcher 136 as described herein, fulfillment module 124 may have access to database 125, which stores a library of rules, heuristics, etc., for various age ranges and / or vocabulary levels. For example, database 125 may store one or more whitelists and / or blacklists of websites, generic resource identifiers (URIs), generic resource locators (URLs), domains, etc., specifying what a user can and cannot access based on their age. Database 125 may also include one or more rules specifying how and / or whether a user can enable the automation assistant 120 to interact with third-party applications, which, as noted above, can be used for, for example, ordering goods or services.

[0075] As noted above, the natural language generator 126 can be configured to generate and / or select natural language output (e.g., words / phrases designed to mimic human speech) based on data obtained from various sources. In some implementations, the natural language generator 126 can be configured to receive performance information associated with the performance of an intent as input and generate natural language output based on the performance information. Additionally or alternatively, the natural language generator 126 can receive information from other sources such as third-party applications (e.g., desired slots), which it can use to compose natural language output for the user.

[0076] If the user's intent is to search for general information, the natural language generator 126 can generate natural language output that conveys the information in response to the user, for example, in sentence form. In some instances, the natural language output can be extracted by the natural language generator 126, without altering the document (e.g., because it is already in complete sentence form) and provided as is. Additionally or alternatively, in some implementations, the response content may not be in complete sentence form (e.g., a request for today's weather may include high temperatures and precipitation chances as separate data pieces), in which case the natural language generator 126 can formulate one or more complete sentences or phrases to present the response content as natural language output.

[0077] In some implementations, the natural language generator 126 may rely on something referred to herein as a "natural language generation template" (or "NLG template") to generate natural language output. In some implementations, the NLG template may be stored in a database 127. The NLG template may include logic (e.g., if / other statements, loops, other programming logic) that specifies how to formulate natural language output in response to various information from various sources, such as data slices containing performance information generated by the performance module 124. Therefore, in some ways, the NLG template can actually constitute a state machine and can be created using any known programming language or other modeling language (e.g., Unified Modeling Language, specification and description languages, Extensible Markup Language, etc.).

[0078] As an example, an NLG template can be configured to respond to a request for weather information in English. The NLG template can specify which of several candidate natural language outputs to provide in various scenarios. For example, suppose the fulfillment message generated by fulfillment module 124 indicates that the temperature will be above, for example, 80 degrees Fahrenheit and there will be no clouds. The logic articulated in the NLG template (e.g., if / otherwise statements) can specify that the natural language output selected by natural language generator 126 is a phrase such as "It's gonna be a scorcher, don't forget your sunglasses." Suppose the fulfillment message generated by fulfillment module 124 indicates that the temperature will be below, for example, 30 degrees Fahrenheit and there will be snow. The logic articulated in the NLG template can specify that the natural language output selected by natural language generator 126 is a phrase such as "It's gonna be chilly, you might want a hat and gloves, and be careful on the road."

[0079] In some implementations, the NLG template may include logic influenced by the detected user's vocabulary level and / or age range. For example, the NLG template might have logic that if the user is predicted to be in a first age range, then a first natural language output is provided, while if the user is predicted to be in another, for example, older age range, then a second natural language output is provided. The first natural language output might use more explanatory words and / or phrases with simpler language, making it more likely that younger users will understand it. Under the assumption that older users do not require as much explanation, the second natural language output might be more concise than the first.

[0080] In the weather example above, different things can be told to adults and children based on the logic in the NLG template. For example, if it's going to be cold, the detected child might be given natural language output reminding them to take precautions that wouldn't require reminding an adult. For instance, the natural language output given to the child could be: "It's going to be cold outside and may even snow! Don't forget your coat, scarf, mittens, and hat. Also, make sure to let a grownup know that you will be outside." In contrast, the natural language output given to an adult in the same situation could be: "It's going to be 30 degrees Fahrenheit, with a 20% chance of snow."

[0081] As another example, NLG templates can be configured to query whether certain entities are true or hypothetical, such as "Is..."<entity_name> The NLG template can respond with "Is the toothfairy real?". Such an NLG template might include logic that selects from multiple options depending on the user's predicted age range. For example, if a user asks "Is the toothfairy real?", and the user is predicted to be a child, the logic within the NLG template could specify an answer like "Yes, the tooth fairy leaves money under pillows of children when they lose teeth." If the user is predicted to be an adult, the logic within the NLG template could specify an answer like "No, the tooth fairy is a make-believe entity employed by parents and guardians to motivate children to pull already-loose teeth in order to get money in exchange."

[0082] Figure 2 Demonstrates how age can be determined based on the user's predicted / estimated age. Figure 1 Examples of various components for processing utterances from user 201 are provided. Components relevant to the techniques described herein are depicted, but this is not intended to be limiting, and may be extended to other areas. Figure 2 Various other components, not described in the text, are used as part of processing the user's speech.

[0083] When a user provides spoken words, the input / output ("I / O") components (such as a microphone) of a client device (e.g., 106) operated by the user 201 can capture the user's words as an audio recording. Figure 2 (The "audio recording" in the text). The audio recording can be provided to a proficiency detector 111, which predicts and / or estimates the age range and / or vocabulary level of user 201 as described above.

[0084] The estimated age range or vocabulary level is then provided to the invocation module 112. Based on the estimated age range and data indicating the user's utterance (e.g., audio recordings, feature vectors, embeddings), the invocation module 112 can classify the user's utterance as either intended to trigger the automatic assistant 120 or not. As noted above, for children with relatively low vocabulary levels or another user, the threshold that must be met for the invocation module 112 to classify the utterance as an appropriate invocation can be reduced.

[0085] In some implementations, the threshold adopted by the invocation module 112 can also be lowered based on other signals such as detected ambient noise, motion (e.g., driving a vehicle), etc. One reason for lowering the threshold based on these other signals is that when these other signals are detected, there may be more noise in the user's speech compared to if the user were speaking in a quiet environment. In such cases, especially if the user is driving or cycling, it may be expected that the autonomous assistant 120 will be invoked more easily.

[0086] Refer again Figure 2 Once the automated assistant 120 is triggered, the STT module 117 can use audio recordings (or data indicating the user's utterances) in conjunction with the estimated age range / vocabulary level to select one or more thresholds and / or models (e.g., from database 113). As noted above, these thresholds / models can be part of a "query understanding model." The STT module 117 can then generate a textual explanation of the user's utterances as output. This textual explanation can be provided to a natural language processor 122, which annotates and / or otherwise processes the textual explanation as described above. The output of the natural language processor 122, along with the estimated age range or vocabulary level of the user 201, is provided to the intent matcher 136.

[0087] Intent matcher 136 may select one or more grammars and / or models from database 137 based on the detected / predicted age of user 201. In some implementations, some intents may be unavailable due to the predicted / estimated age range or vocabulary level of user 201. For example, when the user's estimated age range or vocabulary level is determined to be, for example, below a threshold, intents involving the automatic assistant 120 interacting with third-party applications—particularly those requesting spending money and / or otherwise unsuitable for children—can be disallowed, for example, by one or more rules stored in database 137. Additionally or alternatively, in some implementations, these rules may be stored and / or implemented elsewhere, for example, in database 125.

[0088] The intent determined by the intent matcher 136, along with the estimated age range or vocabulary level and any user-provided slot values ​​(if applicable), can be provided to the fulfillment module 124. The fulfillment module 124 can fulfill the intent according to various rules and / or heuristics stored in the database 125. The fulfillment information generated by the fulfillment module 124 can be passed to the natural language generator 126, for example, along with the user's estimated age range and / or vocabulary level. The natural language generator 126 can be configured to generate natural language output based on the fulfillment information and the user's estimated age range and / or vocabulary level. For example, and as previously described, the natural language generator 126 can implement one or more NLG templates that include logic for generating natural language output based on the user's estimated age range and / or vocabulary level.

[0089] The text output generated by the natural language generator 126, along with the user's estimated age range and / or vocabulary level, can be provided to the TTS module 116. The TTS module 116 can select from the database 115 one or more voice synthesizers to be used by the automated assistant 120 to deliver audio output to the user, based on the user's estimated age range and / or vocabulary level. The audio output generated by the TTS module 116 can be provided to one or more I / O components 250 of a client device operated by the user 201, enabling the audio output to be audibly output via one or more speakers and / or visually output on one or more displays.

[0090] Figure 3A and Figure 3B Describe example scenarios where the techniques described in this article can be employed. Figure 3A In this example, the first user 301A is a relatively young child attempting to interact with an automated assistant 120, which operates at least partially on a client device 306. In this example, the client device 306 takes the form of an assistant device and more specifically, a standalone interactive speaker, but this is not intended to be limiting.

[0091] The first user 301A utters "OK, Assistant, wanna music" which can be captured as an audio recording and / or as a vector of features extracted from the audio recording. The automated assistant 120 can first estimate the age range of user 301A based on the audio recording / feature vector. In some implementations, the proficiency detector 111 can analyze features such as phonemes, pitch, rhythm, etc., to estimate the age of user 301A. Additionally or alternatively, in some implementations, the audio recording can be processed by other components of the automated assistant 120 (such as the STT module 117) to generate a textual interpretation of the user's utterance. This textual interpretation can be analyzed, for example, by the proficiency detector 111 to estimate / predict the grammatical and / or vocabulary proficiency of user 301A (which can, in some cases, be used as a proxy for the user's age).

[0092] Once the automated assistant 120 is invoked, it can analyze the remainder of the utterance "Wanna music" from user 301A. Such a phrase might prompt a conventional automated assistant to seek disambiguation and / or clarification from user 301A. However, the automated assistant 120 configured with selected aspects of this disclosure can have a higher tolerance for grammatical, lexical, and / or pronunciation errors than a conventional automated assistant. Therefore, and using methods such as those previously discussed... Figure 2 The described process involves the various components of the automated assistant 120 using the estimated age range and / or vocabulary level to... Figure 2 At different points in the pipeline, various models, rules, syntaxes, and trial-and-error methods are selected to process and fulfill user requests. In this case, the automated assistant 120 can respond by outputting music for children (e.g., nursery rhymes, songs from children's TV shows and movies, educational music, etc.).

[0093] Figure 3B Another example depicting a child user 301B interacting with the automated assistant 120 again using client device 306. In this example, user 301B again says "Hay assissi, Giggy gat," which cannot be interpreted by a regular automated assistant. However, the automated assistant 120, configured with selected aspects of this disclosure, can estimate / predict that user 301B is in a younger age range, such as 2-3 years old. Therefore, the automated assistant 120 can be more tolerant of grammatical, vocabulary, and / or pronunciation errors. Because it is a child interacting with the automated assistant, the phrase "Giggy gat" can be interpreted more tolerantly as "kitty cat."

[0094] Furthermore, if an adult user invokes the automated assistant 120 and simply says "kitty cat," the automated assistant 120 may fail to respond because the adult user's intent is unclear. However, using the techniques described herein, the automated assistant 120, for example through the intent matcher 136 and / or the fulfillment module 124, can determine the user 301B's intent as wanting to hear a cat's meow based on this interpretation and the user's estimated age range. As a result, the automated assistant 120 can output the sound of a cat meowing.

[0095] Figure 4A Another example depicting a child user 401A interacting with an automated assistant 120 operating at least partially on a client device 406A, which again takes the form of an assistant device, and more specifically, an independent interactive speaker. In this example, user 401A asks, "OK assistant, why grownups putted milk in the refrigerator?" Such a question might not be comprehensible to a conventional automated assistant due to its various grammatical and vocabulary errors, as well as various mispronunciations. However, an automated assistant 120 configured with selected aspects of this disclosure may be more tolerant of such errors and therefore able to handle questions from user 401A. Additionally, in some implementations, the automated assistant 120 may be configured to teach user 401A appropriate grammar and / or vocabulary. Regardless of whether the user is a child (e.g., ... Figure 4A The same applies to users whose native language is not the language used by the Auto Assistant 120 (as shown in the image) or those whose native language is not the language used by the Auto Assistant 120.

[0096] exist Figure 4A In the context of the implementation, before answering the user's question, the automated assistant 120 states, "The proper way to ask that question is, 'Why do grownups put milk in the refrigerator?'" This statement is intended to guide the user 401A regarding appropriate grammar, vocabulary, and / or pronunciation. In some implementations, the automated assistant 120 may monitor a particular user's speech over time to determine if the user's vocabulary is improving. In some such implementations, if the user demonstrates an improved vocabulary level from one conversational session to the next, the automated assistant 120 may congratulate the user or otherwise provide encouragement, for example, to encourage the user to continue improving in language. In some implementations, another user, such as a child's parent, teacher, or guardian, may ask the automated assistant 120 for updates on the child's language progress. The automated assistant 120 may be configured to provide data indicating the child's progress (e.g., as natural language output, as results displayed on the screen, etc.).

[0097] Once the automated assistant 120 has provided language guidance to user 401A, it then provides information in response to the user's query (e.g., obtained from one or more websites by the fulfillment module 124), but in a manner chosen based on an estimated age range and / or vocabulary of user 401A. In this example, the automated assistant 120 says, "because germs grow in milk outside of the refrigerator and make it yucky." Such an answer is clearly tailored for a child. This can be accomplished in various ways. In some implementations, components such as the fulfillment module 124 and / or the natural language generator 126 can replace relatively complex words such as "bacteria," "spoil," "perishable," or "curdle" with simpler words like "germs" and "yucky" based on the estimation that user 401A is a child. If an adult asks the same question, the automated assistant 120 might generate an answer such as: "Milk is aperishable food and therefore is at risk when kept outside the arefrigerator. Milk must be stored in the arefrigerator set to below 40°F for proper safety." In other implementations, one or more NLG templates might be used to influence the output to the user 401A based on their estimated age range and / or vocabulary level.

[0098] In various implementations, different functionalities may be provided (or blocked) depending on the user's predicted age. For example, in some implementations, various games or other activities for children may be provided to child users. In some such implementations, different games may be provided and / or recommended to users based on their predicted age range. Additionally or alternatively, the same games may be available / recommended but may be set to be more difficult for users in a more advanced age range.

[0099] Figure 4B Another example scenario depicting the interaction between a child user 401B and the automated assistant 120. Unlike the previous example, in this example, the automated assistant 120 operates at least partially on a client device 406B in the form of a smart TV or a standard TV equipped with a digital media player dongle (not depicted). In this example, it can be assumed that the automated assistant 120 has predicted / estimated that the user 401B is a young child, for example, in the range of two to four years old.

[0100] In this example, user 401B says "Let's play a game." Based on the user's predicted age range, the automated assistant 120 can, for example, randomly select a game from a set of games tailored to user 401B's age range and / or randomly select a game from all games (where games tailored to user 401B's age range are more weighted). Additionally or alternatively, a list of games suitable for user 401B can be presented to user 401B (e.g., visible on client device 406B).

[0101] The automated assistant 120 then causes the client device 406B to render four zebras (e.g., still images, animations, etc.) visible on its display. The automated assistant 120 asks, "How many zebras do you see?" After counting, the user 401B responds, "I see four." In other implementations, the user can say something like "I see this many" and hold up her fingers. The automated assistant 120 can count the fingers, for example, by analyzing digital images captured by the client device 406B's camera (not depicted). In either case, the automated assistant 120 can determine that "four" matches the number of rendered zebras and can respond with "Great job!"

[0102] exist Figure 4B In this example, user 401B exclusively uses a smart TV (or a standard TV equipped with a smart TV dongle) to interact with the automated assistant 120. However, this is not intended to be limiting. In other implementations, user 401B may use different client devices, such as standalone interactive speakers, to interact with the automated assistant 120, and the visual aspects of this example may still be rendered on the TV screen. For example, the smart TV and the interactive standalone speaker may be part of the same coordinated "ecosystem" of client devices, either controlled by a single user (e.g., a parent or head of household) or associated with all members of the household (or other groups such as colleagues, neighbors, etc.).

[0103] In some implementations, the features detected / predicted / estimated by the user can be used, for example, by the automated assistant 120 as slot values ​​when performing various tasks. These tasks may include, for example, the automated assistant 120 engaging with one or more "agents". As previously noted, and as used herein, an "agent" can refer to, for example, a process that receives input such as slot values, intents, etc., from the automated assistant or elsewhere and provides output in response. A web service is an example of an agent. Third-party applications described above can also be considered agents. Agents can use slot values ​​provided by the automated assistant for various purposes, such as fulfilling user requests.

[0104] Figure 5A and Figure 5B This example illustrates how a user's age range can be used as a slot value, causing different responses to be elicited from the same agent, which in this example is a third-party application in the form of a joke service. In this example, the third-party agent accepts a slot value indicating the user's age and selects an appropriate joke based at least in part on that slot value. Thus, the jokes told by the automated assistant 120 are suitable for the audience.

[0105] exist Figure 5A In this scenario, user 501A is a child. Therefore, when he requests a joke, the automated assistant 120 (again in the form of an independent interactive speaker), which operates at least partially on client device 506, provides the child's predicted / estimated age range to a third-party joke service. The joke service selects and provides an age-appropriate joke: "Why did the birdie go to the hospital? To get a tweet," for output by the automated assistant 120.

[0106] and Figure 5B Conversely, in this case, user 501B is an adult. In this example, user 501B makes the exact same request: "OK, Assistant, tell me a joke." Because the words from user 501B are... Figure 5A The utterance made by user 501A is better formed and / or grammatically correct than that of user 501B, so user 501B is identified as an adult (or at least not a child). Therefore, a joke more age-appropriate for user 501B is given to her: "What's the difference between a taxidermist and a tax collector? The taxidermist only takes the skin."

[0107] The examples described herein involve estimating a user's age (or age range) and vocabulary level and operating the automated assistant 120 in a manner that adapts to these estimates. However, this is not intended to be limiting. Other characteristics of the user can be predicted / estimated based on their utterances. For example, in some implementations, the user's gender can be estimated and used to influence various aspects of how the automated assistant 120 operates. As an example, suppose a user is experiencing a specific medical symptom that represents different things for different genders. In some implementations, the automated assistant 120 can use the predicted user gender to select a piece of information appropriate to the user's gender from multiple potentially conflicting pieces of information obtained from various medical websites. As another example, in Figure 5A and Figure 5B Within the context of the joke, different jokes can be presented to male and female users.

[0108] Figure 6 This is a flowchart illustrating an example method 600 according to the implementation disclosed herein. For convenience, the operations of the flowchart are described with reference to the system performing the operations. This system may include various components of various computer systems, such as one or more components of the computing system implementing the automated assistant 120. Furthermore, although the operations of method 600 are shown in a specific order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0109] At box 602, the system may receive spoken utterances from a user at one or more input components (e.g., microphones) of one or more client devices. At box 604, the system (e.g., via a proficiency detector 111) may apply data indicative of the spoken utterances (e.g., audio recordings, feature vectors) across a trained machine learning model to generate output. At box 606, the system may determine, based on the output generated at box 604, that the user falls into (or should be classified into) a predetermined age group, for example, one of several predefined age groups. In other implementations, the user may be classified into other categories, such as gender.

[0110] At box 608, the system (e.g., invocation module 112) can determine that the user has invoked the automated assistant 120 based on data of the user's predetermined age group and the indicated utterance. For example, an on-device model used by invocation module 112 can be trained to classify the user's utterance as an invocation or not, for example, depending on whether the invocation score generated based on the model meets a certain threshold. In some implementations, this threshold can be adjusted, for example, downwards or upwards, when it is determined that the user is a child or otherwise linguistically deficient. Thus, if the user is detected as a child, it is more likely that the child's utterance will be classified as an invocation than if the user is an adult. In other words, it may be easier for a child (or other user) with a limited vocabulary to invoke the automated assistant 120 than for an adult or other user with a relatively advanced vocabulary level.

[0111] At box 610, the system can select a given query understanding model associated with a predetermined age group from a plurality of candidate query understanding models. As noted above, this query understanding model may include components used by STT module 117 to generate speech recognition output and / or components used by intent matcher 136 to determine the user's intent.

[0112] At box 612, the system, for example via intent matcher 136, can use a given query understanding model to determine the user's intent. As previously described, in some implementations, a given query understanding model may require lower precision when determining the user's intent when the user is a child and their utterance is grammatically incorrect or difficult to understand, compared to when the user is an adult and higher precision is expected.

[0113] At box 614, the system can determine whether the user's intent is resolvable based on a predetermined age group. For example, if the predetermined age group is the age group of young children, the intent to request payment for fulfillment may not be resolvable. At box 616, the system can, for example through fulfillment module 124, fulfill (or parse) the user's intent to generate response data (e.g., the fulfillment data described previously). At box 618, the system can output the response data at one or more output components (e.g., speakers, displays, etc.) of one or more client devices.

[0114] In the example described herein, the Auto Assistant 120 can be switched to Child Mode to be more tolerant of imperfect syntax, provide child-friendly output, and ensure that children cannot access inappropriate content and / or trigger actions such as spending money. However, this is not intended to be restrictive, and other aspects of how the client device operates may be affected by the Auto Assistant 120 switching to Child Mode. For example, in some implementations, volume settings can be limited in Child Mode to protect a child's hearing.

[0115] Additionally, and as previously indirectly mentioned, in some implementations, when in child mode, the automatic assistant 120 can use basic client device features to offer various games, such as guessing animals based on output animal noise, guessing numbers between 1 and 100 (higher / lower), answering riddles or trivia, etc. For example, riddles or trivia questions can be selected based on the difficulty of the riddles or trivia and the user's estimated age range: younger users can receive easier riddles / trivia questions.

[0116] Sometimes, a particular assistant device may be primarily used by children. For example, a standalone interactive speaker could be deployed in a children's play area. In some implementations, such an assistant device can be configured, for example, by a parent, so that the automated assistant using that device to engage is in child mode by default. When in default child mode, the automated assistant may exhibit one or more of the behaviors described previously, such as being more tolerant of grammatical and / or vocabulary errors, being more proactive when engaging with children, etc. If an adult, such as a parent, wishes to use that particular assistant device to engage with the automated assistant 120, the adult can (at least temporarily) switch the assistant device to non-child mode, for example, by speaking a specific invocation phrase baked into the invocation model adopted by invocation module 112.

[0117] In some implementations, instead of predicting or estimating the speaker's age range, the automated assistant 120 can be configured to perform voice recognition to authenticate the speaker. For example, family members can each train the automated assistant 120 to recognize their voices, for example, so that the automated assistant 120 knows who it is speaking to. Users can have profiles, in particular including their age / birthday. In some implementations, the automated assistant 120 can use these profiles to deterministically determine the user's exact age, rather than estimating their age range. Once the user's age is determined, the automated assistant 120 can operate in a manner appropriate to the user's age, as described above.

[0118] Figure 7 This is a block diagram of an example computing device 710, which may optionally be used to perform one or more aspects of the techniques described herein. In some embodiments, one or more of a client computing device, a user-controlled resource module 130, and / or other components may include one or more components of the example computing device 710.

[0119] Computing device 710 typically includes at least one processor 714 that communicates with multiple peripheral devices via a bus subsystem 712. These peripheral devices may include a storage subsystem 724, including, for example, a memory subsystem 725 and a file storage subsystem 726; a user interface output device 720; a user interface input device 722; and a network interface subsystem 716. The input and output devices allow users to interact with computing device 710. The network interface subsystem 716 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.

[0120] User interface input device 722 may include a keyboard; pointing devices such as a mouse, trackball, touchpad, or graphics tablet; a scanner; a touchscreen integrated into a display; an audio input device such as a voice recognition system; a microphone; and / or other types of input devices. Generally, the term "input device" is used to include all possible types of devices and the manner in which information is input into computing device 710 or a communication network.

[0121] User interface output device 720 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or other mechanisms for creating visual images. The display subsystem may also provide non-visual displays, such as via an audio output device. Generally, the term "output device" is used to encompass all possible types of devices and the manner in which information is output from computing device 710 to a user or another machine or computing device.

[0122] Storage subsystem 724 stores the functional programming and data construction of some or all of the modules described herein. For example, storage subsystem 724 may include execution... Figure 6 The selected aspects and implementation of the method Figure 1 and Figure 2 The logic of the various components described in the document.

[0123] These software modules are typically executed by processor 714 alone or in combination with other processors. The memory 725 used in storage subsystem 724 may include multiple memories, including main random access memory (RAM) 730 for storing instructions and data during program execution and read-only memory (ROM) 732 for storing fixed instructions. File storage subsystem 726 can provide persistent storage for program and data files and may include hard disk drives, floppy disk drives, and associated removable media, CD-ROM drives, optical disk drives, or removable media cartridges. Modules implementing the functionality of certain embodiments may be stored in file storage subsystem 726 within storage subsystem 724 or in other machines accessible to processor 714.

[0124] The bus subsystem 712 provides a mechanism for enabling various components and subsystems of the computing device 710 to communicate with each other as intended. Although the bus subsystem 712 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0125] The computing device 710 can be of various types, including workstations, servers, computing clusters, blade servers, server groups, or any other data processing system or computing device. Due to the constantly evolving nature of computers and networks, Figure 7 The description of the computing device 710 depicted herein is intended only as a specific example for illustrating some implementations. Many other configurations of the computing device 710 may have... Figure 7 The computing device depicted in the text has more or fewer components.

[0126] Where the system described herein collects or otherwise monitors personal information about a user, or where personal and / or monitored information may be used, the user may be given the opportunity to control whether a program or feature collects user information (e.g., information about the user's social networks, social behavior or activities, occupation, user preferences, or the user's current geographic location), or to control whether and / or how content more relevant to the user is received from a content server. Similarly, some data may be processed in one or more ways before storage or use to remove personally identifiable information. For example, a user's identity may be treated such that no personally identifiable information can be determined for that user, or geographic location information (such as city, zip code, or state level) may be generalized to make it impossible to determine the user's specific geographic location. Therefore, the user can control how information about the user is collected and / or how information is used. For example, in some implementations, the user may choose not to allow the automated assistant 120 to attempt to estimate their age range and / or vocabulary level.

[0127] While several embodiments have been described and illustrated herein, various other means and / or structures may be utilized for performing functions and / or obtaining results and / or one or more advantages described herein, and each of these variations and / or modifications is considered to be within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and actual parameters, dimensions, materials, and / or configurations will depend on the specific application or use of the teachings. Those skilled in the art will recognize or be able to determine many equivalents of the specific embodiments described herein using no more than conventional experimentation. Therefore, it is to be understood that the foregoing embodiments are presented as examples only, and embodiments may be practiced in ways different from those specifically described and claimed within the scope of the appended claims and their equivalents. Embodiments of this disclosure relate to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of this disclosure if these features, systems, articles, materials, kits, and / or methods do not contradict each other.< / topping> < / topping> < / artist> < / artist> < / artist> < / entity>

Claims

1. A method implemented using one or more processors, comprising: Receive spoken words from a user at one or more input components on one or more client devices; The trained machine learning model applies data instructing the spoken utterance to generate output; Based on the output, it is determined that the user falls into a predetermined age group; Based on the predetermined age group, a tolerance level associated with the predetermined age group is selected, wherein the tolerance level includes at least one of grammatical tolerance, pronunciation tolerance, and vocabulary tolerance; The user's intent is determined using a given query understanding model associated with the tolerance, wherein the tolerance includes a minimum confidence threshold; By determining that the confidence measurement generated for the spoken utterance meets the minimum confidence threshold, and based on the tolerance associated with the predetermined age group, it is determined that the user's intent is parseable, wherein the confidence measurement would fail to meet a higher minimum confidence threshold associated with a different age group. Analyze the user's intent to generate response data; and The response data is output at one or more output components of one or more client devices in the client device.

2. The method of claim 1, further comprising selecting the given query understanding model from a plurality of candidate query understanding models.

3. The method according to claim 2, wherein, The given query understanding model is selected based on the predetermined age group.

4. The method according to claim 3, wherein, The plurality of candidate query understanding models includes at least one candidate query understanding model having a grammar tolerance that is different from the grammar tolerance associated with the predetermined age group.

5. The method according to claim 1, wherein, The data indicating the spoken speech includes an audio recording of the user's speech, and the trained machine learning model is trained to generate an output indicating the user's age based on one or more phonemes contained in the audio recording.

6. The method of claim 1, further comprising selecting a given natural language generation model associated with the predetermined age group from a plurality of candidate natural language generation models, wherein, The selected natural language generation model is used to generate the response data.

7. The method according to claim 6, wherein, The plurality of candidate natural language generation models includes at least one candidate natural language generation model that uses a larger vocabulary than that used by the given natural language generation model.

8. The method according to claim 1, wherein, The spoken utterance includes a request for information, and parsing the intent includes summarizing information from web pages responding to the request for information, wherein the web pages are summarized with a complexity appropriate to the predetermined age group.

9. The method according to claim 1, wherein, The given query understanding model is applied to perform speech-to-text processing of the spoken utterance.

10. The method according to claim 1, wherein, The given query understanding model is applied to perform natural language understanding of the speech recognition output generated from the spoken utterance.

11. A system comprising one or more processors and a memory storing instructions, the instructions being responsive to execution by the one or more processors to cause the one or more processors to: Receive spoken words from a user at one or more input components on one or more client devices; The trained machine learning model applies data instructing the spoken utterance to generate output; Based on the output, it is determined that the user falls into a predetermined age group; Based on the predetermined age group, a tolerance level associated with the predetermined age group is selected, wherein the tolerance level includes at least one of grammatical tolerance, pronunciation tolerance, and vocabulary tolerance; The user's intent is determined using a given query understanding model associated with the tolerance, wherein the tolerance includes a minimum confidence threshold; By determining that the confidence measurement generated for the spoken utterance meets the minimum confidence threshold, and based on the tolerance associated with the predetermined age group, it is determined that the user's intent is parseable, wherein the confidence measurement would fail to meet a higher minimum confidence threshold associated with a different age group. Analyze the user's intent to generate response data; and The response data is output at one or more output components of one or more client devices in the client device.

12. The system of claim 11, further comprising instructions for selecting the given query understanding model from a plurality of candidate query understanding models.

13. The system according to claim 12, wherein, The given query understanding model is selected based on the predetermined age group.

14. The system according to claim 12, wherein, The plurality of candidate query understanding models includes at least one candidate query understanding model having a grammar tolerance that is different from the grammar tolerance associated with the predetermined age group.

15. The system according to claim 11, wherein, The data indicating the spoken speech includes an audio recording of the user's speech, and the trained machine learning model is trained to generate an output indicating the user's age based on one or more phonemes contained in the audio recording.

16. The system of claim 11, further comprising selecting a given natural language generation model associated with the predetermined age group from a plurality of candidate natural language generation models, wherein, The selected natural language generation model is used to generate the response data.

17. The system according to claim 11, wherein, The spoken utterance includes a request for information, and parsing the intent includes summarizing information from web pages responding to the request for information, wherein the web pages are summarized with a complexity appropriate to the predetermined age group.

18. The system according to claim 11, wherein, The given query understanding model is applied to perform natural language understanding of the speech recognition output generated from the spoken utterance.

19. A method implemented using one or more processors, comprising: Receive spoken words from a user at one or more input components on one or more client devices; Perform voice recognition to determine the user's identity; The profile associated with the identity is consulted to determine if the user falls into a predetermined age group; Based on the predetermined age group, a tolerance level associated with the predetermined age group is selected, wherein the tolerance level includes at least one of grammatical tolerance, pronunciation tolerance, and vocabulary tolerance; The user's intent is determined using a given query understanding model associated with the tolerance, wherein the tolerance includes a minimum confidence threshold; By determining that the confidence measurement generated for the spoken utterance meets the minimum confidence threshold, and based on the tolerance associated with the predetermined age group, it is determined that the user's intent is parseable, wherein the confidence measurement would fail to meet a higher minimum confidence threshold associated with a different age group. Analyze the user's intent to generate response data; and The response data is output at one or more output components of one or more client devices in the client device.

20. The method of claim 19, further comprising selecting the given query understanding model from a plurality of candidate query understanding models.

21. The method according to claim 20, wherein, The given query understanding model is selected based on the predetermined age group.

22. The method according to claim 21, wherein, The plurality of candidate query understanding models includes at least one candidate query understanding model having a grammar tolerance that is different from the grammar tolerance associated with the predetermined age group.

23. The method of claim 19, further comprising selecting a given natural language generation model associated with the predetermined age group from a plurality of candidate natural language generation models, wherein, The selected natural language generation model is used to generate the response data.

24. The method according to claim 23, wherein, The plurality of candidate natural language generation models includes at least one candidate natural language generation model that uses a larger vocabulary than that used by the given natural language generation model.

25. The method according to claim 19, wherein, The spoken utterance includes a request for information, and parsing the intent includes summarizing information from web pages responding to the request for information, wherein the web pages are summarized with a complexity appropriate to the predetermined age group.

26. The method according to claim 19, wherein, The given query understanding model is applied to perform speech-to-text processing of the spoken utterance.

27. The method according to claim 19, wherein, The given query understanding model is applied to perform natural language understanding of the speech recognition output generated from the spoken utterance.

28. A method implemented using one or more processors, comprising: Receive spoken words from a user at one or more input components on one or more client devices; Perform voice recognition to determine the user's identity; The profile associated with the identity is consulted to determine if the user falls into a predetermined age group; Select a given query understanding model associated with the predetermined age group from a plurality of candidate query understanding models, wherein the plurality of candidate query understanding models includes at least one candidate query understanding model having a tolerance different from that of the given query understanding model, wherein the tolerance includes at least one of grammatical tolerance, pronunciation tolerance and lexical tolerance, wherein the tolerance includes a minimum confidence threshold; Use the given query understanding model to determine the user's intent; By determining that the confidence measurement generated for the spoken utterance meets the minimum confidence threshold, based on the predetermined age group, it is determined that the user's intent is parseable, wherein the confidence measurement will fail to meet a higher minimum confidence threshold associated with different age groups; Analyze the user's intent to generate response data; and The response data is output at one or more output components of one or more client devices in the client device.

29. The method according to claim 28, wherein, The data indicating the spoken speech includes an audio recording of the user's speech, and a trained machine learning model is trained to generate an output indicating the user's age based on one or more phonemes contained in the audio recording.

30. The method of claim 28, further comprising selecting a given natural language generation model associated with the predetermined age group from a plurality of candidate natural language generation models, wherein, The selected natural language generation model is used to generate the response data.

31. The method according to claim 30, wherein, The plurality of candidate natural language generation models includes at least one candidate natural language generation model that uses a more complex vocabulary than that used by the given natural language generation model.

32. The method of claim 28, further comprising selecting a given speech synthesis model associated with the predetermined age group from a plurality of candidate natural speech synthesis models, wherein, The output of the response data is performed using the given speech synthesis model.

33. The method according to claim 28, wherein, The given query understanding model is applied to perform speech-to-text processing of the spoken utterance.

34. The method according to claim 28, wherein, The given query understanding model is applied to perform natural language understanding of the speech recognition output generated from the spoken utterance.

35. A system comprising one or more processors and a memory storing instructions, the instructions being responsive to execution by the one or more processors to cause the one or more processors to: Receive spoken words from a user at one or more input components on one or more client devices; Perform voice recognition to determine the user's identity; The profile associated with the identity is consulted to determine if the user falls into a predetermined age group; Based on the predetermined age group, a tolerance level associated with the predetermined age group is selected, wherein the tolerance level includes at least one of grammatical tolerance, pronunciation tolerance, and vocabulary tolerance; The user's intent is determined using a given query understanding model associated with the tolerance, wherein the tolerance includes a minimum confidence threshold; By determining that the confidence measurement generated for the spoken utterance meets the minimum confidence threshold, and based on the tolerance associated with the predetermined age group, it is determined that the user's intent is parseable, wherein the confidence measurement would fail to meet a higher minimum confidence threshold associated with a different age group. Analyze the user's intent to generate response data; and The response data is output at one or more output components of one or more client devices in the client device.

36. The system of claim 35, further comprising instructions for selecting the given query understanding model from a plurality of candidate query understanding models.

37. The system according to claim 36, wherein, The given query understanding model is selected based on the predetermined age group.