Tailoring interactive dialogue applications based on author-provided content
By adapting a fixed code with creator-specified variables, the method efficiently generates tailored dialog applications, reducing resource usage and enhancing user interaction quality.
Patent Information
- Application Number
- JP2024113522
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-12-11
- Filing Date
- 2024-07-16
- Publication Date
- 2025-09-10
- Estimated Expiration
- 2038-10-03
AI Technical Summary
Existing interactive dialog applications require significant computational resources and storage space for creating multiple tailored versions, as they often need unique code instances for each version, leading to inefficiencies in resource utilization.
A method for generating tailored versions of dynamic interactive dialog applications using structured content specified by creators, where a fixed code is adapted with creator-specified variables, reducing the need for redundant computational resources and storage by allowing shared code execution across versions.
This approach reduces computational burden and storage requirements by enabling efficient adaptation of dialog applications with persona values, resulting in more understandable and natural user interface outputs, facilitating effective communication and reducing the duration of interactions.
Smart Images

Figure 0007737516000001 
Figure 0007737516000002 
Figure 0007737516000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to tailoring interactive dialog applications based on author-provided content. [Background technology]
[0002] Automated assistants (also known as "personal assistants," "mobile assistants," etc.) can interact with users through a variety of client devices, such as smartphones, tablet computers, wearable devices, automotive systems, standalone personal assistant devices, etc. Automated assistants receive input from users (e.g., typed and / or spoken natural language input) and respond with responsive content (e.g., visual and / or audible natural language output). Automated assistants that interact through client devices can be executed through the client device itself and / or through one or more remote computing devices in network communication with the client device (e.g., computing devices in the "cloud"). Summary of the Invention [Means for solving the problem]
[0003] The present specification generally relates to a method, an apparatus, and a computer-readable medium for executing a tailored version of a dynamic interactive dialog application, where the tailored version is tailored based on structured content specified by a creator of the tailored version. Executing the tailored version of the interactive dialog application may be responsive to receiving, via an assistant interface of an assistant application, a launch phrase assigned to the tailored version and / or other user interface input identifying the tailored version. Executing the tailored version may include generating multiple instances of user interface output for presentation via the assistant interface. Each of the multiple instances of user interface output is for a corresponding dialog turn during execution of the interactive dialog application, and each of the multiple instances is generated through adapting the dynamic interactive dialog application using structured content. For example, various variables of the dynamic interactive dialog application may be populated with values based on the creator-specified structured content, thereby adapting the interactive dialog application to the structured content.
[0004] As described herein, multiple tailored versions of a dynamic interactive dialog application can be executed, where each tailored version is executed based on corresponding structured content specified by a corresponding author. For example, when executing a first tailored version, first values based on first structured content specified by the first author can be utilized for various variables of the dynamic interactive dialog application, when executing a second tailored version, second values based on second structured content specified by the second author can be utilized to populate various variables of the dynamic interactive dialog application, etc.
[0005] In these and other approaches, the same fixed code can be executed for each of multiple tailored versions, adapting only one variable specified by the creator in the structured content for that version, specified by the creator's other user interface input when generating that version, and / or predicted based on the structured content and / or other user interface input for that version, when each tailored version is executed. This can result in a reduction in the computational resources required to create an interactive dialog application. For example, a creator of a tailored version of a dynamic interactive dialog application can utilize computational resources in specifying variables through structured content and / or other input, and these variables are utilized to adapt the interactive dialog application described above (and elsewhere herein). However, the creator does not need to utilize significant computational resources in specifying various codes for the complete execution of the tailored version, because the fixed code of the dynamic interactive dialog application is utilized instead. Moreover, this can result in a reduction in the amount of computer hardware storage space required to store multiple applications. For example, variables for each of multiple tailored versions can be stored without requiring a unique instance of fixed code to be stored for each of the multiple tailored versions.
[0006] In some implementations described herein, a tailored version of a dynamic interactive dialog application runs with one or more persona values specified by a creator of the tailored version and / or predicted based on structured content and / or other inputs provided by the creator when creating the tailored version. The persona values can be utilized for one or more of the interactive dialog application's variables, which also adapts the interactive dialog application based on the persona values. Each of the persona values can affect audible and / or graphical user interface outputs generated when the tailored version runs.
[0007] For example, the one or more persona values may define the tone, intonation, pitch, and / or other voice characteristics of computer-generated utterances that will be provided as natural language user interface outputs upon execution of the adjusted version. Also, for example, the one or more persona values may define terms, phrases, and / or degrees of formality (i.e., not defined in the specified structured content) that will be utilized for various user interface outputs, such as user interface outputs defined in fixed codes. For example, the one or more persona values for a first adjusted version may result in a very formal variety of natural language user interface outputs being provided (e.g., excluding colloquial and / or other casual speech), while the one or more persona values for a second adjusted version may result in a very casual variety of natural language user interface outputs being provided (i.e., less formal). Also, for example, one or more persona values for a first adjusted version may result in a variety of natural language user interface outputs being provided that include terminology specific to a first region (and not including terminology specific to a second region), while one or more persona values for a second adjusted version may result in a variety of natural language user interface outputs being provided that include terminology specific to a second region (and not including terminology specific to the first region). As yet another example, one or more persona values may define music, sound effects, graphical properties, and / or other features to be provided as user interface outputs.
[0008] Executing a tailored version of a dynamic interactive dialog application with persona values for the tailored version can result in the tailored version providing a more understandable and natural user interface output, thereby facilitating more effective communication with the user. For example, techniques described herein can cause the tailored version to convey meaning to a particular user using language and / or phrasing that is more easily understood by the user. For example, as described herein, persona values can be determined based on structured content utilized to execute the tailored version and can be adapted as a result to users who are likely to launch the tailored version. Adapting natural language user interface output based on persona values can shorten the overall duration of the interactive dialog engaged in through execution of the tailored version than would otherwise be necessary, thereby reducing the computational burden on a computing system to execute the tailored version.
[0009] As described above, in various implementations, one or more persona values are predicted based on the structured content and / or other inputs provided by the author when creating the adjusted version. The predicted values can be automatically assigned to the adjusted version and / or presented to the author as suggested persona values, and assigned to the adjusted version if accepted by the author via a user interface input. In many of these implementations, the persona values are predicted based on processing at least some of the structured content and / or other inputs provided when creating the adjusted version using a trained machine learning model. For example, at least some of the structured content can be utilized as at least a portion of the input to the trained machine learning model, the inputs processed using the machine learning model to generate one or more output values, and the persona values selected based on the one or more output values. The trained machine learning model can be trained, for example, based on structured content previously input when generating the corresponding previous adjusted version and based on training instances that are respectively generated based on previously input persona values (e.g., explicitly selected by the corresponding author) for the corresponding previous adjusted version.
[0010] In some implementations described herein, structured content and / or other inputs provided by an author when creating a tailored version of an interactive dialog application are utilized in indexing the tailored version. For example, the tailored version can be indexed based on one or more activation phrases provided by the author for the tailored version. As another example, the tailored version can additionally or alternatively be indexed based on one or more entities determined based on structured content specified by the author. For example, at least some of the entities can be determined based on having a defined relationship (e.g., in a knowledge graph) to multiple entities having aliases included in the structured content. Such entities can be utilized to index the tailored version even if aliases for such entities were not included in the structured content. For example, the structured content may include aliases for numerous points of interest in a given city, but may not include any aliases for the given city. The aliases for the points of interest, and optionally other content, can be utilized to identify entities corresponding to the points of interest in the knowledge graph. Further, it can be determined that all of the entities in the knowledge graph have a defined relationship (e.g., a "located in" relationship) to a given city. Based on the defined relationship and based on a plurality (e.g., at least a threshold) of entities having the defined relationship, a reconciled version can be indexed based on the given city (e.g., indexed by one or more aliases of the given city). A user can then discover the reconciled version through entering a user interface input that references the given city.For example, a user can provide speech input via an automated assistant interface, such as, "I want an application about [alias for a given city]." Based on the tailored version indexed based on the given city, an automated assistant associated with the automated assistant interface can automatically execute the tailored version or can present an output to the user indicating the tailored version as an option for execution, and can execute the tailored version if a positive user interface input is received in response to the presentation. In these and other approaches, tailored versions of dynamic dialog applications that satisfy a user's request can be efficiently identified and executed. This can prevent the user from having to submit multiple requests to identify such tailored versions, thereby saving computational and / or network resources.
[0011] In some implementations, a method performed by one or more processors is provided, including receiving, via one or more network interfaces, instructions for a dynamic interactive dialog application, structured content for executing a tailored version of the dynamic interactive dialog application, and at least one launch phrase for the tailored version of the dynamic interactive dialog application. The instructions, structured content, and at least one launch phrase are transmitted in one or more data packets generated by a user's client device in response to the user's interaction with the client device. The method further includes processing the structured content to automatically select a plurality of persona values for the tailored version of the interactive dialog application, where the structured content does not explicitly indicate a persona value. After receiving the instructions, the structured content, and the at least one launch phrase, and after automatically selecting the plurality of persona values, the method includes receiving a natural language input provided via an assistant interface of the client device or an additional client device, and determining that the natural language input matches a launch phrase for the tailored version of the interactive dialog application. In response to determining that the natural language input matches the activation phrase, the method includes executing an adapted version of the interactive dialog application, including generating a plurality of instances of output for presentation via the assistant interface, each of the plurality of instances of output being for a corresponding dialog turn during execution of the interactive dialog application and generated using the structured content and using a corresponding one or more of the persona values.
[0012] These and other implementations of the techniques disclosed herein can optionally include one or more of the following features.
[0013] In various implementations, processing the structured content to automatically select a plurality of persona values can include utilizing at least some of the structured content as input to a trained machine learning model, processing at least some of the structured content using the trained machine learning model to generate one or more output values, and selecting the persona values based on the one or more output values. In various implementations, the one or more output values can include a first probability of a first persona and a second probability of a second persona, and selecting the persona values based on the one or more output values can include selecting the first persona over the second persona based on the first probability and the second probability, and selecting the persona values based on a persona value assigned to the selected first persona in at least one database. In various other implementations, the method can further include utilizing instructions of a dynamic interactive dialog application as additional input to the trained machine learning model, and processing the instructions and at least some of the structured content using the trained machine learning model to generate one or more output values.
[0014] In various implementations, processing the structured content to automatically select a plurality of persona values may include determining one or more entities based on the structured content, utilizing at least some of the entities as inputs to a trained machine learning model, processing at least some of the entities using the trained machine learning model to generate one or more output values, and selecting the persona values based on the one or more output values.
[0015] In various implementations, before processing at least some of the structured content using the trained machine learning model, the method may further include: identifying a plurality of previous user inputs from one or more databases, each of the previous user inputs including previously input structured content and corresponding previously input persona values, where the previously input persona values are explicitly selected by a corresponding user; generating a plurality of training instances based on the previous user inputs, each of the training instances being generated based on a corresponding one of the previous user inputs and including a training instance input based on the previously input structured content of the corresponding one of the previous user inputs and a training instance output based on the previously input persona value of the corresponding one of the previous user inputs; and training the trained machine learning model based on the plurality of training instances. In some of these implementations, training the machine learning model may include processing training instance inputs for a given training instance from the training instances using the trained machine learning model, generating a predicted output based on the processing, generating an error based on comparing the training instance output and the predicted output for the given training instance, and updating the trained machine learning model based on backpropagation using the error.
[0016] In various implementations, processing the structured content can include parsing the structured content from a document specified by a user.
[0017] In various implementations, the persona value may relate to at least one of the tone of the dialogue, the grammar of the dialogue, and the non-verbal sounds provided by the dialogue.
[0018] In some implementations, a method performed by one or more processors is provided, comprising: receiving, via one or more network interfaces, instructions for a dynamic interactive dialog application and structured content for executing a tailored version of the dynamic interactive dialog application, the instructions and structured content being transmitted in one or more data packets generated by a user's client device in response to the user's interaction with the client device; processing the structured content to determine one or more associated entities; indexing the tailored version of the dynamic interactive dialog application based on the one or more associated entities; after the indexing, receiving natural language input provided via an assistant interface of the client device or an additional client device; determining one or more launch entities from the natural language input; and identifying a mapping of entities, the mapping including at least one of the launch entities and at least one of the associated entities; and identifying the tailored version of the dynamic interactive dialog application based on a relationship between the launch entities and the associated entities in the mapping. Based on identifying the tailored version of the interactive dialog application, the method includes executing a dynamic version of the interactive dialog application, the dynamic version including generating a plurality of instances of output for presentation via an assistant interface, each of the plurality of instances of output being for a corresponding dialog turn during execution of the interactive dialog application and generated using at least some of the structured content of the tailored version of the interactive dialog application.
[0019] These and other implementations of the techniques disclosed herein can optionally include one or more of the following features.
[0020] In various implementations, processing the structured content to determine one or more related entities may include parsing the structured content to identify one or more terms, identifying one or more entities using one or more of the terms as aliases, and determining a given related entity of the one or more related entities based on the given related entity having a defined relationship with a plurality of the identified one or more entities.
[0021] In various implementations, aliases of a given entity may not be included in the structured content. In some of these implementations, determining the given related entity is further based on the given related entity having a defined relationship by at least a threshold amount of the plurality of identified one or more entities.
[0022] In various implementations, the method may further include receiving, via one or more processors, at least one launch phrase for the tailored version of the dynamic interactive dialog application, and further indexing the tailored version of the dynamic interactive dialog application based on the at least one launch phrase.
[0023] In various implementations, the method can further include weighting the related entities based on the structured content and the relationships between the related entities. In some of these implementations, identifying the adjusted version can be further based on the weights of the related entities. In other implementations, the method can further include identifying a second adjusted version of the dynamic interactive dialog application based on the input entities and the related entities, and selecting the adjusted version based on the weights.
[0024] In various implementations, the method can further include identifying a second adjusted version with second structured content and associated entities of the second version based on the input entities and associated entities of the second version, wherein each of a plurality of instances of output is generated using at least some of the structured content and some of the second structured content.
[0025] In some implementations, a method performed by one or more processors is provided, including receiving, via one or more network interfaces, instructions for a dynamic interactive dialog application, structured content for executing a tailored version of the dynamic interactive dialog application, and at least one launch phrase for the tailored version of the dynamic interactive dialog application. The instructions, structured content, and at least one launch phrase are transmitted in one or more data packets generated by a user's client device in response to the user's interaction with the client device. The method further includes, after receiving the instructions, structured content, and at least one launch phrase, receiving a speech input provided via an assistant interface of the client device or an additional client device and determining that the speech input matches a launch phrase for the tailored version of the interactive dialog application. In response to determining that the speech input matches the launch phrase, the method further includes executing the tailored version of the interactive dialog application, including generating multiple instances of output for presentation via the assistant interface, each of the multiple instances of output being for a corresponding dialog turn during execution of the interactive dialog application and generated using the structured content.
[0026] Additionally, some implementations include one or more processors of one or more computing devices, where the one or more processors are operable to execute instructions stored in associated memory, where the instructions are configured to cause performance of any of the aforementioned methods. Some implementations also include one or more non-transitory computer-readable storage media storing computer instructions executable by the one or more processors to perform any of the aforementioned methods.
[0027] It will be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein, for example, all combinations of claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein. [Brief explanation of the drawings]
[0028] [Figure 1] FIG. 1 is a block diagram of an example environment in which implementations disclosed herein can be implemented. [Figure 2] FIG. 1 is a diagram of an example of structured content that can be utilized in the implementations disclosed herein. [Figure 3] FIG. 10 illustrates an example of how persona values can be selected for a request to generate an adapted version of a dynamic interactive dialog application. [Figure 4] 1 is a flow chart illustrating an example method according to implementations disclosed herein. [Figure 5] 1 is a flow chart illustrating an example method for generating a persona selection model according to implementations disclosed herein. [Figure 6] FIG. 1 illustrates a graph with nodes representing entities in a knowledge graph. [Figure 7]FIG. 1 illustrates an example of indexing a tailored version of an application based on entities related to structured content specified for the tailored version of the application. [Figure 8] FIG. 1 illustrates an example of a dialog between a user, a client device, and an automated assistant associated with the client device running an adapted version of an interactive dialog application, according to an implementation disclosed herein. [Figure 9] FIG. 10 illustrates another example of a dialog between a user, a client device, and an automated assistant associated with the client device running another adapted version of the interactive dialog application of FIG. 8, according to an implementation disclosed herein. [Figure 10] FIG. 1 illustrates an example architecture of a computing device. DETAILED DESCRIPTION OF THE INVENTION
[0029] In some cases, a tailored version of the interactive dialog application is generated based on the received structured content and the received instructions for the interactive dialog application. The structured content and instructions can be transmitted to the automated assistant, or a component associated with the automated assistant, in response to one or more user interface inputs provided by the author via interaction with the author's client device. The instructions for the interactive dialog application are utilized to identify the interactive dialog application, and the structured content is utilized in executing the tailored version of the interactive dialog application. Various types of structured content can be provided and utilized in executing the tailored version of the interactive dialog application. For example, the structured content can be a spreadsheet containing prompts and possible responses, such as multiple-choice questions and corresponding answers (e.g., for each question, a correct answer and one or more incorrect answers), jokes and corresponding punchlines (e.g., for each joke, a corresponding punchline), etc. As another example, a structured HTML or XML document, or even an unstructured document that is processed and converted into a structured document, can be provided.
[0030] In some implementations, one or more persona values can be assigned to the adjusted version, and the persona values can be utilized when executing the adjusted version. The persona values can indicate audible characteristics of the speech output, grammatical characteristics to be used in generating natural language for the speech output, and / or specific terms and / or phrases (e.g., in addition to the structured content) to be provided in the speech output. For example, the persona values can collectively define individual personas such as a queen (e.g., a female voice with regular grammar), a robot (e.g., an exaggerated automated voice with a firm speaking tone), and / or a teacher. In some implementations, the persona values can be various characteristics of the presentation voice of the automated assistant that can be changed, such as tone values, grammatical values, and gender values that can be modified to create different personas. In some implementations, one or more of the persona values can be utilized to select a particular voice-to-text model from multiple candidate speech-to-text models that conform to the persona values. In some implementations, one or more of the persona values can be utilized to select corresponding characteristics to utilize during speech-to-text conversion.
[0031] Turning now to the figures, FIG. 1 illustrates an example environment in which the techniques disclosed herein may be implemented. The example environment includes a client device 106, an automated assistant 110, and a tailored application engine 120. In FIG. 1, the tailored application engine 120 is shown as part of the automated assistant 110. However, in many implementations, the tailored application engine 120 may be executed by one or more components separate from the automated assistant 110. For example, the tailored application engine 120 may interface with the automated assistant 110 over one or more networks and, optionally, may interface with the automated assistant 110 using one or more application programming interfaces (APIs). In some implementations in which the tailored application engine 120 is separate from the automated assistant 110, the tailored application engine 120 is controlled by a distinct third party from the party controlling the automated assistant 110.
[0032] The client device 106 may be, for example, a standalone voice-activated speaker device, a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device in a user's vehicle, and / or a user wearable device including a computing device (e.g., a user's watch with a computing device, a user's eyeglasses with a computing device, a virtual or augmented reality computing device). Additional and / or alternative client devices may be provided.
[0033] 1 as separate from the client device 106, in some implementations, all or aspects of the automated assistant 110 can be executed by the client device 106. For example, in some implementations, the input processing engine 112 can be executed by the client device 106. In implementations in which one or more (e.g., all) aspects of the automated assistant 110 are executed by one or more computing devices separate from the client device 106, the client device 106 and these aspects of the automated assistant 110 communicate over one or more networks, such as a wide area network (WAN) (e.g., the Internet). As described herein, the client device 106 can include an automated assistant interface through which a user of the client device 106 interfaces with the automated assistant 110.
[0034] Although only one client device 106 is shown in combination with the automated assistant 110, in many implementations, the automated assistant 110 can be remote and interface with each of a user's client devices and / or with each of multiple users' client devices. For example, the automated assistant 110 can manage communications with each of multiple devices over various sessions and can manage multiple sessions in parallel. For example, the automated assistant 110 in some implementations can be implemented as a cloud-based service using a cloud infrastructure, e.g., using a server farm or cluster of high-performance computers running software suitable for handling large volumes of requests from multiple users. However, for simplicity, many examples herein are described with respect to a single client device 106.
[0035] Automated assistant 110 includes input processing engine 112, output engine 135, and launch engine 160. In some implementations, one or more of the engines of automated assistant 110 can be omitted, combined, and / or implemented in a component separate from automated assistant 110. Moreover, automated assistant 110 can include additional engines not shown herein for brevity. For example, automated assistant 110 can include a dialog state tracking engine, its own dialog engine, etc. (or can share dialog module 126 with tailored application engine 120).
[0036] The automated assistant 110 receives instances of user input from the client device 106. For example, the automated assistant 110 can receive free-form natural language speech input in the form of a streaming audio recording. The streaming audio recording can be generated by the client device 106 in response to a signal received from a microphone of the client device 106 capturing the spoken input of a user of the client device 106. As another example, the automated assistant 110 can receive free-form natural language typed input. In some implementations, the automated assistant 110 can receive non-free-form input from a user, such as a selection of one of multiple options on a graphical user interface element, or structured content provided in generating a tailored version of an interactive dialog application (e.g., in a separate spreadsheet or other document). In various implementations, input is provided at the client device through an automated assistant interface through which a user of the client device 106 interacts with the automated assistant 110. The interface can be an audio-only interface, a graphics-only interface, or an audio-and-graphics interface.
[0037] In some implementations, user input can be generated by the client device 106 and / or provided to the automated assistant 110 in response to explicit activation of the automated assistant 110 by a user of the client device 106. For example, activation can be detection of certain user voice input by the client device 106 (e.g., a hotword / phrase for the automated assistant 110, such as "Hey, assistant"), user interaction with hardware and / or virtual buttons (e.g., tapping a hardware button, selecting a graphical interface element displayed by the client device 106), and / or other specific user interface input. In some implementations, the automated assistant 110 can receive user input that indicates (directly or indirectly) a particular application that can be executed (directly or indirectly) by the automated assistant 110. For example, the input processing engine 112 can receive the input "Assistant, I want to play the Presidents Quiz" from the client device 106. The input processing engine 112 can parse the received audio and provide the parsed content to the activation engine 160. Launch engine 160 can utilize the parsed content to determine (e.g., utilizing index 152) that "Presidents Quiz" is the launch phrase for a tailored version of the dynamic interactive dialog application. In response, launch engine 160 can transmit a launch command to tailored application engine 120 to cause tailored application engine 120 to execute the tailored version and engage in an interactive dialog with a user of client device 106 via the automated assistant interface.
[0038] The automated assistant 110 typically provides instances of output in response to receiving instances of user input from the client device 106. The instances of output can be, for example, audio to be audibly presented by the device 106 (e.g., output through speakers of the client device 106), text and / or graphical content to be graphically presented by the device 106 (e.g., rendered through a display of the client device 106), etc. As described herein, when executing a tailored version of an interactive dialog application, the output provided in a given dialog turn can be generated by the tailored application engine 120 based on the interactive dialog application and on structured content for the tailored version and / or persona values for the tailored version. As used herein, a dialog turn refers to a user utterance (e.g., an instance of speech input or other natural language input) and a response system utterance (e.g., an instance of audible and / or graphical output), or vice versa.
[0039] The input processing engine 112 of the automated assistant 110 processes natural language input received via the client device 106 and generates annotated output for use by one or more other components of the automated assistant 110, such as the launch engine 160, the tailored application engine 120, etc. For example, the input processing engine 112 can process free-form natural language input generated by a user via one or more user interface input devices of the client device 106. The generated annotated output includes one or more annotations of the natural language input and, optionally, one or more (e.g., all) of the terms of the natural language input. As another example, the input processing engine 112 may additionally or alternatively include a voice-to-text module that receives instances of speech input (e.g., in the form of digital audio data) and converts the speech input into text including one or more textual words or phrases. In some implementations, the speech-to-text module is a streaming voice-to-text engine. The speech-text module may rely on one or more stored speech-text models (also called language models), each of which may model the relationships between audio signals and speech units in a language, as well as the sequences of words in the language.
[0040] In some implementations, the input processing engine 112 is configured to identify and annotate various types of grammatical information in the natural language input. For example, the input processing engine 112 may include a portion of a speech tagger configured to annotate terms with these grammatical features. For example, the portion of the utterance tagger may tag each term with its part of the utterance, such as “noun,” “verb,” “adjective,” “pronoun,” etc. Also, for example, in some implementations, the input processing engine 112 may additionally and / or alternatively include a dependency parser configured to determine syntactic relationships between terms in the natural language input. For example, the dependency parser may determine which terms modify other terms, subjects and verbs of sentences, etc. (e.g., a parse tree) and create annotations of such dependencies.
[0041] The output engine 135 provides an instance of output to the client device 106. In some situations, the instance of output can be based on response content generated by the tailored application engine 120 when executing the tailored version of the interactive dialog application. In other situations, the instance of output can be based on response content generated by another application that is not necessarily a tailored version of the interactive dialog application. For example, the automated assistant 110 can itself include one or more internal applications that generate response content and / or can interface with third-party applications that generate response content but are not tailored versions of the interactive dialog application. In some implementations, the output engine 135 can include a text-to-speech engine that converts text components of the response content into audio format, and the output provided by the output engine 135 is in an audio format (e.g., streaming audio). In some implementations, the response content can already be in an audio format. In some implementations, the output engine 135 additionally or alternatively provides textual reply content as output (optionally for conversion to audio by the device 106) and / or provides other graphical content as output for graphical display by the client device 106.
[0042] The tuned application engine 120 includes an indexing module 122, a persona module 124, a dialog module 126, an entity module 128, and a content entry engine 130. In some implementations, modules of the tuned application engine 120 may be omitted, combined, and / or implemented in components separate from the tuned application engine 120. Moreover, the tuned application engine 120 may include additional modules not shown herein for the sake of brevity.
[0043] The content input engine 130 processes content provided by an author to generate a tailored version of the interactive dialog application. In some implementations, the provided content includes structured content. For example, referring to FIG. 2, an example of structured content is provided. The structured content can be transmitted to the content input engine 130 from the client device 106 or from another client device (e.g., another user's client device). The structured content of FIG. 2 is a spreadsheet, and each row 205a-205d of the spreadsheet includes an entry in a question column 210, an entry in a correct answer column 215, and an entry in each of three incorrect answer columns 220. In some implementations, the column headers of the spreadsheet of FIG. 2 can be pre-populated by the tailored application engine 120, and the entries in each of the rows can be populated by an author using a corresponding client device and using the headers for guidance. In some of these implementations, the tailored application engine 120 pre-populates headers based on what the creator indicates they want to create tailored versions of multiple available interactive dialog applications for. For example, the headers in Figure 2 may be pre-populated based on the creator selecting the "Trivia" interactive dialog application. On the other hand, if the user selects the "Jokes" interactive dialog application, the "Jokes" and "Punch Line" headers would be pre-populated instead.
[0044] The content entry engine 130 may receive the structured content of FIG. 2 , optionally process the content, and store the content in the adjusted content database 158 for use in executing an adjusted version of the corresponding interactive dialog application. Processing the content may include annotating and / or storing the entry based on the column and row for the entry provided by the user. For example, the content entry engine 130 may store the entry in column 210, row 205A and annotate it as a “question” entry for the adjusted version. Furthermore, the content entry engine 130 may store the entry in column 215, row 205A and annotate it as a “correct answer” entry for a previously stored “question” entry, and may store the entry in column 220, row 205A and annotate it as an “incorrect answer” entry for a previously stored “question” entry. Processing the content may additionally and / or alternatively include verifying that the structured content values comply with one or more required criteria, and prompting the author to correct them if they do not. The one or more required criteria may include content type (e.g., numbers only, letters only), content length (e.g., X characters and / or Y terms), etc.
[0045] Along with the structured content for the tailored version, the content entry engine 130 can also store instructions for the corresponding interactive dialog application, any provided launch phrases for the tailored version, and any selected and / or predicted persona values for the tailored version. While a spreadsheet is shown in FIG. 2, it is understood that the content entry engine 130 can process other types of structured content. Moreover, in some implementations, the content entry engine 130 can convert unstructured content into a structured format and then process the structured format.
[0046] Dialog module 126 executes the tailored version of the interactive dialog application using the application and structured content for the tailored version, optionally along with additional values for the tailored version (e.g., persona values). For example, dialog module 126 can execute a given tailored version of a given interactive dialog application by retrieving a fixed code for the given interactive dialog application from application database 159 and retrieving structured content and / or other content for the given version from tailored content database 158. Dialog module 126 can then execute the given tailored version utilizing the fixed code for the given interactive dialog application and the tailored content for the tailored version.
[0047] When executing the adapted version of the interactive dialog application, the dialog module 126 engages in multiple dialog turns. In each dialog turn, the dialog module 126 can provide content to the output engine 135, which can provide the content (or a transformation thereof) as user interface output to be presented (audibly or graphically) at the client device 106. The provided output can be based on the interactive dialog application, as well as structured content and / or persona values. Moreover, the provided output in many dialog turns can be based on user utterances (of the dialog turn and / or previous dialog turns) and / or system utterances of previous dialog turns (e.g., a "dialog state" determined based on past user and / or system utterances). The user utterances of a dialog turn can be processed by the input processing engine 112 and output from the input processing engine 122 for use by the dialog module 126 in determining response content to provide. Note that while many instances of output will be based on the interactive dialog application and structured content, some instances of output can be based on the interactive dialog application without referencing structured content. For example, the interactive dialog application can include fixed code that enables responses to various user inputs utilizing only the fixed code and / or referencing other content not provided by the author. That is, many dialog turns during the execution of the adjusted version will be affected by the provided structured content, but some dialog turns will not be affected.
[0048] In some implementations, when executing a tailored version of the interactive dialog application, the dialog module 126 can additionally and / or alternatively customize one or more instances of output for a given user based on the given user's behavior. The given user's behavior can include behavior in one or more previous dialog turns of the current execution of the tailored version and / or behavior in one or more interactive dialogs with the given user in prior executions of the tailored version and / or for prior executions of other tailored versions and / or other interactive dialog applications. For example, if the interactive dialog application is a trivia application and a given user struggles to answer a question correctly, a hint can be proactively provided along with the question in one or more outputs and / or the persona value can be adjusted to be more “encouraging.” For example, the output provided in response to an incorrect answer when executing a tailored version of the trivia application can initially be “Wrong, wrong, wrong,” but can be adapted to “Good answer, but incorrect, try again” in response to the user's incorrect answers to at least a threshold amount of questions. Such adaptation can occur via adaptation of one or more persona values. As another example, the output provided in response to a correct answer when running a tailored version of a trivia application might initially be "That's right, good job," but in response to the user's successful execution and / or a slowdown in the pace of the dialogue, can be adapted to simply "That's right" (e.g., to speed up the dialogue).A given user's behavioral data (e.g., scores, number of errors, time spent, and / or other data) can be persisted throughout execution of the adjusted version and even across multiple instances of execution of the adjusted version (and / or other adjusted versions and / or other applications) and utilized to customize the given user's experience through the adaptation of one or more outputs. In some implementations, such user behavior-specific adaptation when executing the adjusted version can occur through the adaptation of persona values, which can include further adaptation of one or more persona values already adapted based on the structured content and / or other features of the adjusted version.
[0049] The persona module 124 uses the tailored version's author-provided structured content and / or other content to select one or more persona values for the tailored version of the interactive dialog application. The persona module 124 can store the persona values in the tailored content database 158 along with the tailored version. The selected persona values are utilized by the dialog module 126 during execution of the tailored version. For example, the selected persona values can be utilized in selecting the grammar, tone, and / or other aspects of the spoken output to be provided in one or more dialog turns during execution of the tailored version. The persona module 124 can utilize various criteria in selecting persona values for the tailored version. For example, the persona module 124 can utilize the tailored version's structured content, the tailored version's launch phrases, entities associated with the structured content (e.g., as described below for the entity module 128), the corresponding interactive dialog application (e.g., a typing quiz, a joke, a list of facts, or a transportation scheduler), etc. In some implementations, the persona module 124, in selecting one or more persona values, can utilize one or more selection models 156. As described herein, the persona module 124 can automatically select and execute one or more persona values for the adjusted version and / or select one or more persona values and require confirmation by the creator before execution for the adjusted version.
[0050] A persona is an individual personality and is composed of a collection of persona values, some of which may be predefined to reflect this particular type of persona. For example, a persona may include "queen," "king," "teacher," "robot," and / or one or more other distinct types. Each persona may be represented by multiple persona values, each reflecting a particular aspect of the persona, and each individual persona may have a default value for (or be limited to the values that can be assigned to) each. As an example, a persona is a collection of persona values, which may include speaking voice characteristics, grammatical characteristics, and non-verbal sound characteristics (e.g., music played between quiz question rounds). A "queen" persona may have the following persona values: SPEAKING VOICE=(Value 1), GRAMMAR=(VALUE 2), and SOUND=(VALUE 3). A “Teacher” persona may have persona values of SPEAKING VOICE=(VALUE 4), GRAMMAR=(VALUE 5), and SOUND=(VALUE 6). In some implementations, a persona may be selected for a tailored version, and the corresponding persona values may be set based on the persona values that comprise the persona. For example, the techniques described herein for selecting persona values may select persona values by first selecting a persona, identifying persona values that are considered for the selected persona, and setting the tailored version's persona values accordingly. In some implementations, one or more of the persona values may not be part of the persona and may be set independently of the persona when the persona is selected. For example, a “Teacher” persona may not have a “Gender” persona value set, and persona values that have a nature indicating “male” or “female” may be assigned independently to the “Teacher” persona.
[0051] The entity module 128 utilizes the terms and / or other content provided in the received structured content to determine one or more entities referenced in the structured content, and optionally, one or more entities related to such referenced entities. The entity module 128 may utilize the entity database 154 (e.g., a knowledge graph) in determining such entities. For example, the structured content may include the terms “Presidents,” “George Washington,” and “Civil War.” The entity module 128 may identify entities associated with each of the terms from the entity database 154. For example, the entity module 128 may identify an entity associated with the first president of the United States based on the fact that “George Washington” is an alias for this entity. The entity module 128 may optionally identify one or more additional entities related to such referenced entities, such as an entity associated with “Presidents of the United States” based on the fact that “George Washington” has a “belongs to” relationship with the “Presidents of the United States” entity.
[0052] The entity module 128 can provide the determined entities to the indexing module 122 and / or the persona module 124. The indexing module 122 can index the adjusted version corresponding to the structured content based on the one or more entities determined based on the structured content. The persona module 124 can utilize one or more of the entities in selecting one or more persona values for the adjusted version.
[0053] Referring now to FIG. 3 , an example is provided of how persona values may be selected for a request to generate a tailored version of a dynamic interactive dialog application. The content input engine 130 receives an instruction 171 for an interactive dialog application to tailor. This may be, for example, an instruction for a quiz application, an instruction for a joke application, or an instruction for a transportation query application. The instruction may be received in response to an author selecting a graphical element corresponding to the interactive dialog application, speaking a term corresponding to the interactive dialog application, or otherwise indicating a desire to provide structured content to the interactive dialog application. Additionally, the content input engine 130 receives structured content. For example, in response to the instruction 171 for the quiz application, the content input engine 130 may receive a document including the content shown in FIG. 2 . Additionally, the content input engine 130 receives a launch phrase 173. For example, for the structured content of FIG. 2 , the author may provide an instruction phrase 173 for “presidential trivia.”
[0054] The content input engine 130 provides the instructions, at least some of the structured content, and / or activation phrases 174 to the persona module 124 and / or the entity module 128. The entity module 128 utilizes the instructions, at least some of the structured content, and / or activation phrases 174 to identify entities 175 referenced in one or more of these items using the entity database 154 and provides the entities 175 to the persona module 124.
[0055] Persona module 124 uses the selection model and at least one of data 174 and / or 175 to select one or more persona values 176. For example, persona module 124 may utilize one of the selection models that is a machine learning model to process data 174 and / or 175 and, based on the processing, generate an output indicative of persona value 176. For example, the output may indicate a probability for each of a plurality of distinct personas, one of which is selected based on the probability (e.g., a “queen” persona), and persona value 176 can be a collection of persona values that are ascribed to this persona. As another example, the output may include a probability for each of a plurality of persona values and a subset of these persona values selected based on the probability.
[0056] As an example, for structured content intended for an elementary school quiz, persona module 124 may select a persona value that causes the corresponding adjusted version to provide a speech output that is "slower," provide fewer of all possible incorrect answers as options for responding to a question, provide encouraging feedback as a response output even for incorrect answers, and / or limit word use to terms that should be known to children. For example, the selected persona value may be such that "Close but not quite. Please try again" is provided as response content when an incorrect response is received, while for structured content intended for an adult audience, persona module 124 may instead select a persona value that causes the output "Wrong, wrong, wrong" to be provided when an incorrect answer is received.
[0057] In some implementations, the persona module 124 can additionally and / or alternatively select persona values based on attributes of a given user on whom the adjusted version is running. For example, when the adjusted version is running for a given user who possesses the attribute “adult,” the persona module 124 can select persona values based on such attributes (and optionally, additional attributes). In these and other approaches, the persona values of a given adjusted version can be adapted on a “per-user” basis, thereby tailoring each version of the adjusted application to the user on whom it is running. A machine learning model can optionally be trained and utilized in selecting such persona values based on attributes of the user on whom the adjusted version is running. For example, the machine learning model can utilize training examples based on explicit selections by users on whom the adjusted version is running, where the explicit selections indicate one or more persona values that such users desire to be utilized when running the adjusted version. The training examples can optionally be based on the structured content of the adjusted version. For example, a machine learning model can be trained to predict one or more persona values based on the attributes of the user on whom the tailored version is being run, as well as based on the structured content and / or other features of the tailored version.
[0058] The persona values 176 can be stored in the tailored content database 158 along with the structured content 172, instructions 171, and / or launch phrases. Collectively, such values can define a tailored version of the interactive dialog application. A subsequent user can provide a natural language utterance to the automated assistant, which can then identify a term in the natural language input that corresponds to the tailored version's launch phrase 173. For example, the launch engine 160 (FIG. 1) can process the received user interface input to determine which, if any, of multiple previously submitted tailored versions of the interactive dialog application is being launched. In some implementations, when a tailored version of the application is generated, the creator can provide a launch phrase that will be utilized in the future to launch the application. In some implementations, one or more entities can be identified as relevant to the structured content, and the launch engine 160 can select one or more of the tailored versions based on the user input and the associated entities.
[0059] 4 is a flow diagram illustrating an example of how a tailored version of an interactive dialog application is executed for a subsequent user based on the tailored version's structured content and persona values. For convenience, the operations of the flow diagram of FIG. 4 are described with reference to a system that performs the operations. This system may include various components of various computer systems. Moreover, although the operations of the method of FIG. 4 are shown in a particular order, this is not meant to be limiting. One or more operations may be rearranged, omitted, or added.
[0060] Natural language input is received from a user at block 405. The natural language may be received by a component that shares one or more characteristics with the input processing engine 112 of FIG.
[0061] At block 410, one or more terms are identified in the natural language input. For example, the terms may be identified by a component that shares one or more characteristics with the input processing engine 112. Additionally, entities associated with one or more of the terms may optionally be identified. For example, the entities may be determined from an entity database by a component that shares one or more characteristics with the entity module 128.
[0062] At block 415, the previously generated tailored version retains the automatically selected persona value and is identified based on terms and / or entities in the natural language input. In some implementations, the previously generated tailored version can be identified based on the natural language input matching a launch phrase associated with the tailored version. In some implementations, the tailored version can be identified based on an identifying relationship between a segment of the natural language input and one or more entities associated with the tailored version, as described in more detail herein.
[0063] In block 420, a prompt is generated by the dialog module 126. The prompt can then be provided to the user via the output engine 135. For example, the dialog module 126 can generate a text prompt based on the persona values and / or structured content and provide the text to the output engine 135, which can convert the text into an utterance and provide the utterance to the user. When generating the prompt, the dialog module 126 can vary the grammar, word usage, and / or other characteristics of the prompt based on the persona values. Additionally, when providing an utterance version of the text generated by the dialog model 126, the output engine 135 can vary the tone, gender, speaking rate, and / or other characteristics of the output utterance based on one or more of the persona values of the launched, tailored version of the application. In some implementations of block 420, the prompt can be a “starting” prompt that is always provided to the tailored version in the first iteration.
[0064] A user's natural language response is received at block 425. The natural language response can be analyzed by components that share characteristics with the input processing engine 112, which can then determine one or more terms and / or entities from the input.
[0065] Response content is provided at block 430. The response content may be generated based on the received input of block 425, the adjusted version of the structured content, and the adjusted version of the persona values.
[0066] After the response content is provided, additional natural language responses from the user can be received in another iteration of block 425, and additional response content can again be generated and provided in another iteration of block 430. This can continue until the tailored version is complete, until an instance of a natural language response in block 425 indicates a desire to cease interacting with the tailored version, and / or until other conditions are met.
[0067] FIG. 5 is a flow chart illustrating another example of a method according to an implementation disclosed herein. FIG. 5 illustrates an example of training a machine learning model (e.g., a neural network model) for use in selecting a persona value. For convenience, the operations of the flow chart of FIG. 5 are described with reference to a system that performs the operations. The system may include various components of various computer systems. Moreover, although the operations of the method of FIG. 5 are shown in a particular order, this is not meant to be limiting. One or more operations may be rearranged, omitted, or added.
[0068] The system selects structured content and persona values for the tailored version of the interactive dialog application in block 552. As one example, the structured content and persona values may be selected from database 158, and the persona values may have been explicitly dictated by the corresponding creator and / or may have been confirmed by the corresponding creator as desired persona values.
[0069] In block 554, the system generates training instances based on the structured content and persona values. Block 554 includes sub-blocks 5541 and 5542.
[0070] In subblock 5541, the system generates training instance inputs for the training instances based on the structured content and, optionally, on instructions for the interactive dialog application to which the adjusted version corresponds. In some implementations, the system additionally or alternatively generates training instance inputs for the training instances based on entities determined based on the structured content. As one example, the training instance inputs may include instructions for the interactive dialog application and a subset of terms from the structured content, such as the title and the first X terms, or the X most frequently occurring terms. For example, the training instance inputs may include 50 terms from the structured content with the highest TFIDF values, along with a value indicating the interactive dialog application. As another example, the training instance inputs may include instructions for the interactive dialog application and embeddings of some (or all) of the terms from the structured content (and / or other content). For example, the embeddings of the terms from the structured content may be Word2Vec embeddings generated using a separate model.
[0071] For sub-block 5542, the system generates training instance outputs for the training instance based on the persona values. For example, the training instance outputs may include X outputs, each representing a distinct persona. For a given training instance, the training instance outputs may include a "1" (or other "positive" value) for the output corresponding to the distinct persona to which the persona value in block 552 conforms, and a "0" (or other "negative" value) for all other outputs. As another example, the training instance outputs may include Y outputs, each representing a persona characteristic. For a given training instance, the training instance outputs may include, for each of the Y outputs, a value indicating the persona value of the persona value in block 552 for the persona characteristic represented by the output. For example, one of the Y outputs may indicate a degree of "formalism," and the training instance output for this output may be "0" (or other value) if the corresponding persona value in block 552 is "informal," and "1" (or other value) if the corresponding persona value in block 552 is "formal."
[0072] In block 556, the system determines whether there are additional tailored versions of the interactive dialog application to process. If so, the system repeats blocks 552 and 554 using the structured content and persona values from the additional tailored versions.
[0073] Blocks 558-566 may be performed following or in parallel with multiple iterations of blocks 552, 554, and 556.
[0074] In block 558 , the system selects the training instances generated in the iterations of block 554 .
[0075] In block 560, the system applies the training instances as inputs to a machine learning model. For example, the machine learning model may have input dimensions that correspond to the dimensions of the training instance inputs generated in block 5541.
[0076] At block 562, the system generates an output by the machine learning model based on the applied training instance inputs. For example, the machine learning model may have output dimensions that correspond to the dimensions of the training instance outputs generated at block 5541 (e.g., each dimension of the output may correspond to a persona trait).
[0077] In block 564, the system updates the machine learning model based on the generated outputs and the training instance outputs. For example, the system can determine an error based on the generated outputs and the training instance outputs in block 562 and back-propagate the error through the machine learning model.
[0078] At block 566, the system determines whether there are one or more additional unprocessed training instances. If so, the system returns to block 558 to select the additional training instances, and then performs blocks 560, 562, and 564 based on the additional unprocessed training instances. In some implementations, at block 566, the system may determine not to process any additional unprocessed training instances if one or more training criteria have been met (e.g., a threshold number of epochs have occurred and / or a threshold duration of training has occurred). Although method 500 is described with respect to a non-batch learning technique, batch learning can additionally and / or alternatively be utilized.
[0079] A machine learning model trained according to the method of Figure 5 can then be utilized to predict persona values for an adapted version of an interactive dialog application based on structured content and / or other content dictated by the creator of the adapted version. For example, the structured content of Figure 2 can be provided as input to the model, and the personality persona parameters can be values of "Teacher" with probability 0.8 and "Queen" with probability 0.2. This can indicate, based on the structured content, that it is more likely that a user would be interested in a quiz application provided for the Teacher personality than for the Queen personality.
[0080] 5 describes one example of a persona selection model that may be generated and utilized. However, additional and / or alternative persona selection models may be utilized in selecting one or more persona values, such as the alternatives described herein. Such additional and / or alternative persona selection models may optionally be machine learning models trained based on training instances different from those described with respect to FIG. 5.
[0081] As one example, a selection model can be generated based on past explicit selections of persona values by various users, and such selection model can additionally or alternatively be utilized in selecting a particular persona value. For example, in some implementations, an indication of multiple persona values can be presented to a user, and a user selection of one of the multiple persona values can be utilized to select one persona value from the multiple values. Such explicit selections of multiple users can be utilized to generate a selection model. For example, training instances similar to those described above can be generated, but the training instance output for each training instance can be generated based on the persona values selected by the user. For example, for the training instances, a “1” (or other “positive value”) can be utilized for the output dimension corresponding to the selected personality persona value (e.g., the “teacher” personality), and a “0” (or other “negative” value) can be utilized for each of the output dimensions corresponding to all other persona values. Also, for example, for training instances, a "1" (or other "positive value") may be used for output dimensions corresponding to selected persona values, a "0.5" (or other "neutral value") may be used for output dimensions corresponding to other persona values presented to the user but not selected, and a "0" (or other "negative" value) may be used for each of the output dimensions corresponding to all other persona values. In this and other approaches, the user's explicit selection of persona values may be leveraged in generating one or more persona selection models.
[0082] As previously described with respect to FIG. 1 , indexing module 122 receives structured content from a user and indexes the corresponding tailored version in index 152 based on one or more entities related to the structured content. After the tailored version of the application is stored with indications of the associated entities, subsequent user natural language input can be parsed, and terms and entities related to the parsed terms can be identified in entity database 154 by entity module 128. By allowing flexibility in indexing user-created applications, the user and / or subsequent users are not required to know the exact launch phrase. Instead, indexing module 122 enables users to “discover” content, for example, by providing natural language input indicating a desired subject of the served content.
[0083] As an example, a user may provide structured content for a quiz application that includes questions about state capitals. Thus, answers (both correct and incorrect) can be city names, and questions include state names, respectively (or vice versa). The structured content can be received by content entry engine 130, as described above. Additionally, the structured content can be provided with instructions for a dynamic dialog application, and, optionally, launch phrases for launching content in a future, tailored version of the application. After parsing the structured content, content entry engine 130 can provide the parsed content to entity module 128, which can then identify one or more related entities in entity database 154. Returning to the example, and referring to FIG. 6 , a graph of a plurality of nodes is provided. Each of the nodes includes an alias for an entity and represents a portion of entity database 154. The nodes include state capitals, including “Sacramento” 610, “Columbus” 645, “Albany” 640, and “Olympia” 635. Additionally, the graph includes nodes representing related entities. For example, all of the state capital nodes are connected to the "state capital cities" node 625.
[0084] When structured content for a state capitals quiz application is received, the entity module 128 can identify nodes in a graph for the structured content. For example, the structured content can include the question prompt, "What is the capital of California?" with possible answers, "Sacramento" and "Los Angeles." A corresponding node in the graph is then identified. The entity module 128 can then provide an indication of the corresponding node and / or an indication of an entity for the node to the indexing module 122. For example, a state capitals quiz can additionally include the question, "What is the capital of New York?" with the answer choice, "Albany," and the entity module 128 can identify a node for "State Capital Cities" as a general category that links to nodes for "Sacramento" 610 and "Albany" 640.
[0085] In some implementations, the entity module 128 may only identify nodes related to some of the structured content. For example, in a quiz application, the entity module 128 may only identify nodes related to correct answers, but not nodes related to incorrect answers, to avoid associating incorrect entities with the structured content. In some implementations, the entity module 128 may further identify entities related to a launch phrase if provided by a user. For example, a user may provide the launch phrase "capital cities," and the entity module 128 may identify "state capital cities" 625.
[0086] In some implementations, the relationships between the structured content and one or more of the entities can be weighted. For example, the entity module 128 can assign weights to entities identified from correct answers in a quiz application with a score that indicates their relevance over entities related to incorrect answers in the structured content. Additionally, the entity module 128 can weight relationships with categories or other entities based on the number of entities related to both the structured content and the entity. For example, for structured content including "Sacramento" 610, "Olympia" 635, and "Albany" 640, the entity module 128 can weight the relationship with "State Capital City" 625 more than "Western United States Cities" 630 because more of the entities related to the structured content are related to "State Capital City" 625.
[0087] The indexing module 122 then indexes the adjusted version of the application by one or more of the entities. In some embodiments, the indexing module 122 can index the adjusted version of the application by all identified entities. In some implementations, the indexing module 122 can index the adjusted version by only those entities that have a relation score that exceeds a threshold. In some implementations, the indexing module 122 can utilize one or more training models to determine which of the entities to use when indexing the adjusted version of the application.
[0088] The input processing engine 112 can receive natural language input from a user and identify a tailored version of an interactive dialog application to present to the user based on entities indexed in the tailored version. Referring to FIG. 7 , natural language input 181 is received by the input processing engine 112 as described above. The input processing engine 112 parses the input to identify one or more terms within the input. For example, a user can speak the phrase “Show me the state capitals quiz,” and the input processing engine 112 can identify the terms “state,” “state capital,” and “quiz” as parsed input 181. Some of the parsed input 181 can be provided to the entity module 128, which then identifies one or more related entities 183 in an entity database. In some implementations, the entity module 128 can assign weights to the identified entities based on the number of associations between the parsed input 181 and the identified entities.
[0089] The indexing module 122 receives the related entities 183 (and associated weights, if assigned) and identifies one or more adjusted versions 184 of the application indexed by the entities included in the related entities 183. For example, the entity module 128 may identify "state capital cities" as aliases for the related entities, and the indexing module 122 may identify an example of the adjusted version as the version to provide to the user. In some implementations, the indexing module 122 may identify multiple potential versions of the application and select one of the versions based on the weights assigned to the related entities 183 by the entity module 128.
[0090] In some implementations, indexing module 122 can identify multiple potential versions and provide a version to a user that includes content from the multiple potential versions. For example, indexing module 122 can identify a "state capitals" quiz application and can further identify a second "state capitals" quiz in adjusted content database 158 based on entities and related entities associated with the second version. Indexing module 122 can optionally utilize index 152 (FIG. 1) during such identification. Thus, a user can be provided with a composite application that includes structured content from multiple sources seamlessly presented as a single version, even if unrelated users created the two versions.
[0091] 8 illustrates an example of a dialog that may occur between a user 101, a voice-enabled client device 806, and an automated assistant associated with the client device 806 that has access to a tailored version of an interactive dialog application. The client device 806 includes one or more microphones and one or more speakers. One or more aspects of the automated assistant 110 of FIG. 1 may be executed on the client device 806 and / or on one or more computing devices in network communication with the client device 806. Therefore, for ease of explanation, the automated assistant 110 will be referenced in the description of FIG. 8.
[0092] User input 880A is a launch phrase for a tailored version of a dynamic, interactive quiz application. The input is received by input processing engine 112, which identifies “quiz” as relating to the tailored application. Accordingly, input processing engine 112 provides the parsed input to launch engine 160, as described above. In some implementations, launch engine 160 can determine that the input does not include an explicit launch phrase and can provide input to indexing module 122 to determine one or more input entities and identify an associated-entity-indexed version that can be launched by the provided input.
[0093] At output 882A, a prompt is provided. The prompt is provided in the "Teacher" persona and refers to the user as a student in a "Class." At output 882B, a non-verbal sound (i.e., a bell ringing) is included in the prompt and may additionally be part of the "Teacher" persona and / or relate to one or more persona values. The prompt further includes structured content in the form of a question.
[0094] In user input 880B, the user provides an answer. The input processing engine 112 parses the input and provides the parsed input to the dialog module 126. The dialog module 126 verifies that the input is correct (i.e., matches the correct answer in the structured content) and generates a new dialog turn to present to the user. Additionally, output 882C includes the structured content and the dialog generated based on the persona or persona values associated with the version of the application. The user responds incorrectly in user input 880C, and the next output 882D generated by the dialog module 126 reminds the user of the incorrect answer. As an alternative example, if the structured content indicated that the quiz was more likely for young children, output 882D could have provided more encouraging words, allowed a second guess, and / or provided the user with a hint instead of the dialog shown in FIG. 8. In user input 880F, the user indicates a desire to quit the application. This can be a standard activation phrase and / or one of several phrases that instruct the automated assistant to stop sending input to the tailored application.
[0095] Figure 9 illustrates another example of a dialog that may occur between user 101, a voice-enabled client device 906, and an automated assistant associated with client device 906 that has access to an adapted version of an interactive dialog application that has one or more persona values different from the persona values of the dialog in Figure 8, but has the same structured content. At user input 980A, the user launches the adapted version in the same manner as the dialog in Figure 8.
[0096] At output 982A, a prompt is provided. In this dialog, the adjusted version is instead associated with the “Teacher” persona and addresses the user as “Subject,” as opposed to “Class” in the previous example. The dialog module 126 can identify in this output that a title for the user is needed and, based on the persona values associated with the version of the application, determine that the “Queen” persona utilizes “Subject” as the name for the user. At output 982B, a different non-verbal sound (i.e., a trumpet) is included in the prompt and may additionally be part of the “Queen” persona and / or related to one or more persona values. The dialog module 126 can insert different sounds into the prompt depending on one or more of the associated persona values. The prompt further includes the same structured content in the form of a question.
[0097] In user input 980B, the user provides an answer. This is the same answer as provided previously in this user input step in FIG. 8, and dialog module 126 handles the response in the same manner. Additionally, output 982C includes dialog generated based on the structured content and the personas or persona values associated with the version of the application, while tailored to match one or more of the persona values selected for the adjusted version. The user responds incorrectly in user input 980C, and the next output 982D generated by dialog module 126 uses different terminology than in FIG. 8 but acknowledges the incorrect answer. In user input 980F, the user indicates a desire to terminate the application. This can be a standard startup phrase and / or one of several phrases that instruct the automated assistant to discontinue sending input to the adjusted application.
[0098] While several examples of trivia interactive dialog applications are described above, it is understood that various implementations can be utilized with various types of interactive dialog applications. For example, in some implementations, the structured content provided can be a bus timetable, a train timetable, or another transportation timetable. For example, the structured content can be a bus timetable that includes multiple stops (e.g., intersections) and times for each of those stops. The interactive dialog application can include fixed code that enables responses to various queries in a conversational manner. When executing a tailored version based on a bus timetable, the interactive dialog application can utilize the fixed code in determining what types of responses to provide in response to various queries and the structured content in determining at least some of the content to provide in the various responses to the various queries.
[0099] 10 is a block diagram of an example computing device 1010 that may optionally be utilized to perform one or more aspects of the techniques described herein. In some implementations, one or more of the device 106, the automated assistant 110, and / or other components may include one or more components of the example computing device 1010.
[0100] The computing device 1010 typically includes at least one processor 1014 that communicates with several peripheral devices via a bus subsystem 1012. These peripheral devices may include, for example, a storage subsystem 1024 including a memory subsystem 1025 and a file storage subsystem 1026, a user interface output device 1020, a user interface input device 1022, and a network interface subsystem 1016. The input / output devices enable user interaction with the computing device 1010. The network interface subsystem 1016 provides an interface to external networks and is coupled to corresponding interface devices in other computing devices.
[0101] The user interface input devices 1022 may include a keyboard, a pointing device (such as a mouse, trackball, touchpad, or graphics tablet), a scanner, a touchscreen integrated into a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. Overall, use of the term "input device" is intended to encompass all possible types of devices and ways to input information into the computing device 1010 or into a communications network.
[0102] The user interface output devices 1020 may include a display subsystem, a printer, a fax machine, or a non-visual display (such as an audio output device). The display subsystem may include a cathode ray tube (CRT), a flat panel device (such as a liquid crystal display (LCD)), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. Overall, use of the term "output device" is intended to encompass all possible types of devices and manners for outputting information from the computing device 1010 to a user or to another machine or computing device.
[0103] Storage subsystem 1024 stores program and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 1024 may include logic for performing selected aspects of the various methods described herein.
[0104] These software modules are generally executed by the processor 1014 alone or in combination with other processors. The memory 1025 used in the storage subsystem 1024 can include several memories, including a main random access memory (RAM) 1030 for storing instructions and data during program execution, and a read-only memory (ROM) 1032 in which fixed instructions are stored. The file storage subsystem 1026 can provide persistent storage for program and data files and can include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that perform the functions of certain implementations can be stored by the file storage subsystem 1026 in the storage subsystem 1024 or on another machine accessible by the processor 1014.
[0105] The bus subsystem 1012 provides a mechanism for allowing the various components and subsystems of the computing device 1010 to communicate with each other as intended. Although the bus subsystem 1012 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0106] The computing device 1010 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computing device 1010 depicted in Figure 10 is intended to be merely a specific example to illustrate some implementations. Many other configurations of the computing device 1010 can have more or fewer components than the computing device depicted in Figure 10.
[0107] In situations where certain implementations discussed herein may collect or use personal information about a user (e.g., user data extracted from other electronic communications, information about the user's social network, the user's location, the user's time of day, the user's biometric information, and the user's activity and demographic information), the user is provided one or more opportunities to control whether the information is collected, whether the personal information is stored, whether the personal information is used, and how the information about the user is collected, stored, and used. That is, implementations of the systems and methods discussed herein collect, store, and / or use user personal information only upon receiving explicit authorization to do so from the associated user. For example, a user is provided control over whether a program or feature collects user information about this particular user or other users associated with the program or feature. Each user from whom personal information is to be collected is presented with one or more options to enable control over the collection of information associated with that user, to authorize or approve whether information is to be collected and what portions of the information will be collected. For example, a user may be provided with one or more such control options over a communications network. Additionally, certain data may be treated in one or more ways before it is stored or used so that personally identifiable information is removed. As one example, a user's identity may be treated so that personally identifiable information cannot be determined. As another example, a user's geographic location may be blurred to a larger region so that the user's specific location cannot be determined. [Explanation of symbols]
[0108] 101 users 106 Client Device, Device 110 Automated Assistant 112 Input Processing Engine 120 Tuned Application Engine 122 Index Addition Module 124 Persona Module 126 Dialogue Module 128 Entity Module 130 Content Input Engine 135 output engine 152 Index 154 Entity DB, Entity Database 156 Selection Model 158 Adjusted Content, Adjusted Content Database, Database 159 Applications, Application Databases 160 Starting Engine 171 Instructions 172 Structured Content 173 Start phrases, instruction phrases 174 Instructions, Structured Content, and / or Launch Phrase, At least some of the instructions, structured content, and / or launch phrase, data 175 entities, data 176 Persona Value 181 Natural Language Input, Parsed Input 183 Related Entities 184 Adjusted Version, one or more adjusted versions of the application 210 Questions 215 Correct Answers Column 220 Incorrect Answer Column Line 205a Line 205b Line 205c Row 205d 605 California 610 Sacramento 615 Los Angeles 620 California cities 625 state capital cities 630 cities in the Western United States 635 Olympia 640 Albany 645 Columbus 806 voice-enabled client device, client device 906 voice-enabled client device, client device 1010 Computing Devices 1012 Bus Subsystem 1014 processors 1016 Network Interface, Network Interface Subsystem 1020 User Interface Output Device 1022 User Interface Input Device 1024 storage subsystem 1025 Memory Subsystem, Memory 1026 File Storage Subsystem 1030 Main Random Access Memory, RAM 1032 read-only memory, ROM
Claims
1. A method executed by one or more processors, comprising: Through one or more network interfaces, Dynamic interactive dialog application instructions; structured content for executing a tailored version of the dynamic interactive dialog application; receiving a signal from the the structured content comprises a spreadsheet; transmitting the instructions and the structured content in one or more data packets generated by a client device; recording the structured content in one or more databases based on the instructions and based on the columns and rows of the spreadsheet; After receiving the instruction and recording the structured content, receiving natural language input provided via an assistant interface of the client device or an additional client device, the natural language input not explicitly launching the tailored version of the dynamic interactive dialog application; determining that the natural language input corresponds to the adapted version of the dynamic interactive dialog application, said determining step comprising: accessing an index that maps content determined based on the structured content to the adjusted version of the dynamic interactive dialog application; determining that one or more terms of the natural language input correspond to the content; selecting the tailored version of the dynamic interactive dialog application for execution based on a determination that one or more terms of the natural language input correspond to the content; and In response to selecting the adjusted version of the dynamic interactive dialog application for execution, executing the adapted version of the dynamic interactive dialog application; executing the adjusted version of the dynamic interactive dialog application includes generating one or more instances of output for presentation via the assistant interface; each of the one or more instances of output is for a corresponding dialog turn during execution of the dynamic interactive dialog application and is generated using the structured content; A method comprising:
2. The content includes entities inferred from the structured content; 2. The method of claim 1 , wherein determining that one or more terms of the natural language input correspond to the content comprises determining that the one or more terms correspond to the entities inferred from the adjusted version of the structured content.
3. further comprising determining entities inferred from the structured content; determining the entities inferred from the structured content, identifying at least a first entity in the structured content; determining the entity based on the entity being related to the first entity in an entity database; 2. The method of claim 1, comprising:
4. The method of claim 3, wherein the step of determining the entities inferred from the structured content comprises: identifying at least a second entity in the structured content; determining the entity based on the entity being related to the first entity and related to the second entity in the entity database; 4. The method of claim 3, comprising:
5. A method executed by one or more processors, comprising: Through one or more network interfaces, Dynamic interactive dialog application instructions; structured content for executing a tailored version of the dynamic interactive dialog application, the structured content comprising a spreadsheet; additional content for executing the adapted version of the dynamic interactive dialog application, the additional content being additional content to the structured content and not included in the structured content; and receiving the instructions, the structured content, and the additional content in one or more data packets generated by the user's client device in response to the user's interaction with the client device; recording the structured content in one or more databases based on the instructions and based on the columns and rows of the spreadsheet; recording the additional content in one or more of the databases based on the instructions; processing the additional content and the structured content to automatically select a plurality of output response characteristics for the tailored version of the dynamic interactive dialog application, wherein the structured content does not explicitly indicate the output response characteristics; after receiving the instruction, recording the structured content and the additional content, and automatically selecting the plurality of output response characteristics; receiving natural language input provided via an assistant interface on the client device or an additional client device of an additional user; determining that the natural language input corresponds to the adapted version of the dynamic interactive dialog application; In response to determining that the natural language input corresponds to the adjusted version of the dynamic interactive dialog application, executing the adjusted version of the dynamic interactive dialog application, the step including generating one or more instances of output for presentation via the assistant interface, each of the one or more instances of output being for a corresponding dialog turn during execution of the dynamic interactive dialog application and generated based on the structured content and based on one or more of the plurality of output response characteristics; A method comprising:
6. The method of claim 5, wherein the natural language input does not include any launch phrases for the tailored version of the dynamic interactive dialog application.
7. The method described in claim 6, wherein the step of determining that the natural language input corresponds to the adjusted version of the dynamic interactive dialog application includes a step of determining that one or more terms of the natural language input correspond to the structured content of the adjusted version.
8. The method described in claim 7, wherein the step of determining that one or more terms of the natural language input correspond to the structured content of the adjusted version includes the step of determining that the one or more terms correspond to entities inferred from the structured content of the adjusted version.
9. The method of claim 8, wherein determining that the one or more terms correspond to the entities inferred from the structured content comprises: identifying at least a first entity in the structured content based on the one or more terms; determining the entity based on the entity being associated with the first entity in an entity database; 9. The method of claim 8, comprising:
10. The method of claim 1, wherein determining that the one or more terms correspond to the entity inferred from the structured content comprises: identifying at least a second entity in the structured content based on the one or more terms; determining the entity based on the entity being related to the first entity and related to the second entity in the entity database; 10. The method of claim 9, further comprising:
11. The method of claim 5, wherein the plurality of output response characteristics includes one or more grammatical characteristics.
12. A system comprising: one or more databases; a memory for recording instructions; One or more processors and wherein the one or more processors execute instructions stored in the memory, causing the one or more processors to: One or more data packets generated by a client device Dynamic interactive dialog application instructions; structured content for executing a tailored version of the dynamic interactive dialog application, the structured content comprising a spreadsheet; additional content for executing the adapted version of the dynamic interactive dialog application, the additional content being additional content to the structured content and not included in the structured content; and receiving the recording the structured content and the additional content in one or more databases based on the instructions and based on the columns and rows of the spreadsheet; processing the additional content and the structured content to automatically select a plurality of output response characteristics for the tailored version of the dynamic interactive dialog application, wherein the structured content does not explicitly indicate the output response characteristics; and after receiving the instruction, recording the structured content and the additional content, and automatically selecting the plurality of output response characteristics; receiving natural language input provided via an assistant interface on the client device or an additional client device of an additional user; determining that the natural language input corresponds to the adapted version of the dynamic interactive dialog application; In response to determining that the natural language input corresponds to the adjusted version of the dynamic interactive dialog application, executing the tailored version of the dynamic interactive dialog application, wherein in executing the tailored version of the dynamic interactive dialog application, one or more of the processors generate one or more instances of output for presentation via the assistant interface, each of the one or more instances of output being for a corresponding dialog turn during execution of the dynamic interactive dialog application and generated based on the structured content and based on one or more of the plurality of output response characteristics; A system operable to cause 13. The system of claim 12, wherein the natural language input does not include any launch phrases for the tailored version of the dynamic interactive dialog application.
14. In determining that the natural language input corresponds to the adjusted version of the dynamic interactive dialog application, one or more of the processors: The system of claim 13 , further comprising: determining that one or more terms of the natural language input correspond to the structured content in the adjusted version.
15. The system described in claim 14, wherein in determining that one or more terms of the natural language input correspond to the structured content of the adjusted version, one or more of the processors determine that one or more terms correspond to entities inferred from the structured content of the adjusted version.
16. In determining that the one or more terms correspond to the entity inferred from the structured content, one or more of the processors: identifying at least a first entity in the structured content based on the one or more terms; The system of claim 15 , further comprising: determining the entity based on the entity being associated with the first entity in an entity database.
17. In determining that the one or more terms correspond to the entity inferred from the structured content, one or more of the processors: identifying at least a second entity in the structured content based on the one or more terms; The system of claim 16 , further comprising: determining the entity based on the entity being related to the first entity and related to the second entity in the entity database.
18. The system of claim 17, wherein the plurality of output response characteristics include one or more grammatical characteristics.
19. The system of claim 12, wherein the plurality of output response characteristics includes one or more grammatical characteristics.
Citation Information
Patent Citations
Interactive control system, interactive control method, and robot apparatus
JP2003255991A
System and method for providing quiz game capable of presenting quiz created by user
JP2015222561A
A Method for Adaptive Conversational State Management with Filtering Operators Applied Dynamically as Part of a Conversational Interface
JP2016502696A
Providing Virtual Personal Assistance with Multiple VPA Applications
US20140310002A1
Clarifying natural language input using targeted questions
US20140316764A1