Real-time morphing interface for display on a computer screen
The NLP and intent recognition modules of the deformable interface system analyze user input in real time, generate and adjust the user interface, solve the problem of computer systems not responding in time before user input, and improve the user experience and the real-time display of task progress.
Patent Information
- Application Number
- CN202080062354.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-12
- Filing Date
- 2020-09-04
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2040-09-04
AI Technical Summary
Existing computer systems fail to respond or provide applicable tasks in a timely manner before user input, resulting in a poor user experience and lack the ability to display the progress of automated tasks in real time.
Through the deformation interface system, the natural language processing (NLP) pipeline, intent recognition module and entity recognition module are used to analyze user input in real time, generate and dynamically adjust the user interface to match user intent and entity value, and realize real-time deformation of the interface.
It achieves real-time response and dynamic adjustment of the user interface, improves the user experience, ensures that the system quickly adapts to changes in user input, and provides instant feedback and task progress display.
Smart Images

Figure CN114365143B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 895,944, filed September 4, 2019, and U.S. Provisional Application No. 63 / 038,604, filed June 12, 2020, all of which are incorporated herein by reference in their entirety. Background Art
[0003] Computer assistants, such as smart speakers and artificial intelligence programs, are increasingly popular and are used in a variety of user-facing systems. Computerized systems can often be implemented to automate entire processes without requiring the system's human users to have any knowledge of the process. For example, a computer can complete a set of tasks without displaying content on a screen for the user. However, many users prefer to receive feedback on computerized processes, and it may be helpful or necessary for the user to understand the status of a set of tasks if feedback is required at specific steps.
[0004] Additionally, users expect assistive systems to respond as quickly as possible. However, if a system responds to a user before receiving a complete set of instructions from the user, the system may perform a task that is not applicable to the user or may not receive enough information to display content on the screen for the user to view. Therefore, a system that includes, for example, a real-time display of the progress of automated tasks and the ability to adjust the display in response to additional input would be beneficial. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Figure 1 is a high-level block diagram of a system architecture for a morphing interface system according to an example embodiment.
[0006] Figure 2 is a high-level diagram of the interactions between components of the deformation interface system 130 according to an example embodiment.
[0007] Figure 3 is a block diagram illustrating an example NLP signal according to an example embodiment.
[0008] Figure 4A is a flowchart illustrating a process of generating an interface based on user input according to an example embodiment.
[0009] Figure 4B is a flow diagram illustrating a process of deforming an interface upon receiving additional user input according to an example implementation.
[0010] Figure 5A A first layout for an interface display associated with a flight booking intent is shown according to an example embodiment.
[0011] Figure 5B A second layout for display of an interface associated with a flight reservation intent is shown in accordance with an example embodiment.
[0012] Figure 5C A third layout for display of an interface associated with a flight reservation intent is shown in accordance with an example embodiment.
[0013] Figure 5D A fourth layout for display of an interface associated with a flight reservation intent is shown in accordance with an example embodiment.
[0014] Figure 5E A fifth layout for display of an interface associated with a flight reservation intent is shown in accordance with an example embodiment.
[0015] Figure 5F A sixth layout for display of an interface associated with a flight reservation intent is shown in accordance with an example embodiment.
[0016] Figure 5G An email confirmation of execution of a flight reservation intent is shown in accordance with an example embodiment.
[0017] Figure 6A Receipt of a first portion of user input associated with a pizza ordering intent is shown in accordance with an example embodiment.
[0018] Figure 6B A layout for display of an interface associated with ordering a pizza is shown in accordance with an example embodiment.
[0019] Figure 6C A layout for display of an interface associated with purchasing a pizza T-shirt is shown in accordance with an example embodiment.
[0020] Figure 6D Additional example interfaces that can be associated with a T-shirt purchase intent are shown in accordance with an example embodiment.
[0021] Figure 7 is a block diagram showing components of an example machine, capable of reading instructions from a machine-readable medium and executing the instructions in one or more processors, in accordance with an example embodiment.
[0022] A letter following a reference number (e.g., "105A") indicates that the text refers specifically to an element having that particular reference number, while a reference number without a letter following it (e.g., "105") indicates that the text means either any or all of the elements in the figures having that reference number.
[0023] The drawings depict various embodiments for the purpose of illustration only. One skilled in the art will readily recognize from the following discussion that alternative embodiments of the structures and methods illustrated herein can be employed without departing from the principles described herein. DETAILED DESCRIPTION
[0024] The drawings and the following description are directed to preferred embodiments. It should be noted that alternative embodiments disclosed herein will be readily apparent to those skilled in the art, and the use of such alternatives will be considered within the scope of the principles described herein, in accordance with the following discussion, which is merely illustrative.
[0025] Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying drawings. Note that wherever possible, the same or similar reference numbers are used in the drawings and the following description to refer to the same or similar functionality unit components. The drawings depict embodiments of the disclosed system (or method) for purposes of illustration only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein can be employed without departing from the principles described herein.
[0026] Configuration Overview
[0027] Systems (and methods and computer program code stored on non-transitory computer readable media) configured to generate and enable display of a user interface on a computer screen are disclosed. In one example embodiment, the system can include one or more computer processors for executing computer program instructions and a non-transitory computer readable storage medium including stored instructions that can be executed by the at least one processor. In an example embodiment, the instructions can include instructions that, when executed, cause the processor to receive a first input including an input string from a first user device and generate a set of natural language processing signals based on the first user input. The instructions can also include instructions for selecting a user intent that matches the first user input, the selection based on the natural language processing signals; identifying an interface associated with the intent; and extracting a set of values associated with entities of the interface from the user input. The entities can be variables of the interface that can be assigned values. In one example embodiment, the instructions can also include instructions for enabling display of the interface on the user device as a display, the interface for display including values from the set of values, the display occurring approximately at the instant the first user input was received. In one example embodiment, the instructions that can be executed by the at least one processor of the system can also include instructions that, when executed by the processor, cause the processor to receive a second user input including a text string from the user device and generate an updated set of natural language processing signals based on a combination of the first user input and the second user input. The instructions can also include instructions for selecting an intent that matches the combination of the first user input and the second user input based on the updated set of natural language processing signals; identifying a second interface associated with the newly selected intent; extracting a second set of values associated with entities of the second interface from the combination of the first user input and the second user input; and enabling the second interface for display on the user device, the second interface including values from the second set of values.
[0028] In various example embodiments, the first user input and / or the second user input can be a voice input. Further, the first interface and the second interface can include associated sets of entities and can be the same or different interfaces in various example embodiments. The input string can be a text string, an audio input, and / or another form of user input.
[0029] Example System Configuration
[0030] Figure 1 is a high level block diagram of a system architecture for a malleable interface system according to an example embodiment. Figure 1 includes a malleable interface system 130, a network 120, and a client device 110. For clarity, Figure 1Only one client device and one morphing interface system 130 are shown. Alternative embodiments of the system environment can have any number of client devices 110 and multiple morphing interface systems 130. In different embodiments, the functions performed by Figure 1 the various entities can vary. The client device 110 and the morphing interface system 130 can include some or all of the components of a computing device (e.g., the components depicted) and a suitable operating system. Figure 7
[0031] In example embodiments, the morphing interface system 130 generates (or renders or enables for rendering) a user interface for display to the user in response to user input (e.g., a typed or spoken string of text). In some embodiments, the system can also receive visual input, e.g., from a camera or camera roll of the client device 110, to implement a search process on an online marketplace. The morphing interface system 130 determines a user intent. The user intent corresponds to a machine (e.g., a computer or computing system) prediction of what the user can want based on the received user input. Thus, the user intent can be a computer-executable function or request that corresponds to the received user input and / or is described by the received user input. The executable function can be instantiated by generating and / or populating (e.g., at the time of rendering) one or more user interfaces for functions that can be executed and that correspond to the function that can be the predicted intent.
[0032] As the morphing interface system 130 receives additional user input (e.g., more words added to the typed or spoken string of text), the morphing interface system 130 reevaluates whether the determined user intent is still the best match associated with the user input. If another user intent is better suited to the updated user input, the morphing interface system 130 generates and populates a new user interface that is appropriate for the new intent. That is, the user interface “morphs” from one interface to another as the morphing interface system 130 receives more input information about what user intent is best suited to the input (i.e., which function or request best handles the user input). In cases where the morphing interface system 130 determines multiple equally likely intents, the morphing interface system 130 can prompt the user with an interface preview (e.g., by providing information for rendering the interface preview at the client device 110) so that the user can choose between the equally likely intents, or the morphing interface system 130 can automatically select an intent based on stored user preferences (e.g., user preferences learned based on past user interactions with the system).
[0033] A user can input user input, such as typed text or spoken voice input, via a client device 110. The client device 110 can be any personal or mobile computing device, such as a smartphone, tablet, notebook computer, laptop computer, desktop computer, and smartwatch, and any home entertainment device, such as a television, video game console, television box, and receiver. The client device 110 can present information received from the metamorphic interface system 130 to the user, for example, in the form of a user interface. In some implementations, the metamorphic interface system 130 can be stored and executed from the same machine as the client device 110.
[0034] The client device 110 can communicate with the metamorphic interface system 130 via a network 120. The network 120 can include any combination of local and wide area networks employing wired or wireless communication links. In some implementations, all or a portion of the communication over the network 120 can be encrypted.
[0035] The metamorphic interface system 130 includes various modules and data stores to determine intent and / or generate interfaces. The metamorphic interface system 130 includes a natural language processing (NLP) pipeline 135, an intent recognition module 140, an entity recognition module 145, an intent model store 150, an interface store 155, and an entity recognition model store 160. Computer components such as web servers, network interfaces, security functions, load balancers, failover servers, management and network operations consoles, and the like are not shown so as to not obscure the details of the system architecture. Furthermore, the metamorphic interface system 130 can contain more, less, or different components than shown, and the functions of the components described herein can be distributed among the components in a different manner than described herein. Note that the pipeline and modules can be implemented as program code (e.g., software or firmware), hardware (e.g., an application specific integrated circuit (ASIC), field programmable gate array (FPGA), controller, processor), or a combination thereof. Figure 1
[0036] The NLP pipeline 135 receives user input in the form of, for example, text or audio, and generates NLP signals that the morphing interface system 130 can use for intent recognition and for extracting entities. In some implementations, the NLP pipeline 135 performs tokenization, part-of-speech tagging, stemming, lemmatization, stopword identification, dependency parsing, entity extraction, chunking, semantic role labeling, and coreference resolution. In one implementation, the input to the NLP pipeline is a set of one or more words in the form of, for example, a complete or partially complete sentence or phrase. In one implementation, the NLP pipeline 135 produces an annotated version of the input set of words. In another implementation, the NLP pipeline 135 constructs or looks up a numerical representation or feature embedding for immediate use by a downstream module such as the intent recognition module 140 or the entity recognition module 145, which can use a neural network. For example, the input to the NLP pipeline 135 can be a partial sentence, and the output can be the partial sentence with accompanying metadata about the partial sentence.
[0037] The intent recognition module 140 identifies what a user's intent can be based on input received from the user (via the client device 110). In particular, the intent recognition module 140 predicts available intents (i.e., functions) that the morphing interface system 130 can perform. The available intents correspond to a set of words that make up the user input. The user input can match one or more predefined intents. For ease of discussion, the system is described in the context of words. However, it is noted that the principles described herein can also be applied to any set of signals, which can include sound actions (e.g., voice commands or audio tones), video streams (e.g., in ambient computing scenarios), and other potential forms of information input. In different implementations, the intent recognition module 140 can use various machine learning models to determine intents that can be associated with user input. For ease of description, the system will be described in the context of supervised machine learning. However, it is noted that the principles described herein are also applicable to semi-supervised and unsupervised systems.
[0038] In one example implementation, the intent recognition module 140 can use text classification to predict the intent to which a user input most likely corresponds. In this example implementation, a text classification model can be trained using labeled examples of input strings. For example, the morphing interface system 130 can store labeled example input strings. The labels associate each example input string with one of the available intents. The training data can include example input strings in the form of words, partial sentences, partial phrases, complete sentences, and complete phrases. The classification model can also be trained to use various natural language processing signals produced by the NLP pipeline 135, and the training data can additionally include natural language processing signals. The classification model can also utilize signals from the entity recognition module 145, such as using an "airline" entity identified by the entity recognition module 145 to determine that the intent or function is "book a flight." Thus, the classification model is trained to predict which of a set of available intents is most likely to correspond to a given user input string, e.g., using semantic similarity, i.e., determining the closest matching query from a set of example queries for each action.
[0039] In another example implementation, the intent recognition module 140 can use a model that computes semantic similarity scores between a user input and a set of example inputs across a set of available intents. That is, rather than training a model to directly predict the applicable intent based only on labeled training data, the intent recognition module 140 can also compare a user input to some or all previously received user input strings when determining the intent that best matches a given user input. For example, the morphing interface system 130 can store a record of past matched intents and input strings, and if the intent recognition module 140 determines that a current user input string is the same as, or has the same sentence structure as, or has related NLP signals to, a stored previous user input string, the intent recognition module 140 can predict the same intent for the current user input string. In addition to comparing to correctly matched past user input strings, the user input can also be compared to computer-generated strings created by both rule-based methods based on semantics and generative deep learning algorithms.
[0040] In another example implementation, the intent recognition module 140 can utilize a method based on simpler rules to infer the most likely intent of a user input string. This can include regular expression matching, i.e., recognizing certain pre-defined syntactic and grammatical patterns in the input string to determine the user's intent. This can also include utilizing signals from the NLP pipeline 130 (e.g., dependency parsing, constituent parsing, chunking, and / or semantic role labeling) to find the verb, subject, predicate, etc. of the query and match them with data from a stored knowledge base. For example, if the user's input is "get me bananas," the intent recognition module 140 can determine that the word "bananas" is the direct object of the query and retrieve a match for its entry "banana" from the knowledge base to learn that "banana" is a food or ingredient - which, for example, can indicate a match with the intent to purchase food groceries.
[0041] In some example implementations, the morphing interface system 130 includes an intent model store 150. The intent model store 150 can store program code for computer models trained and applied by the intent recognition module 140 to predict the most likely intent related to a given user input string. In some implementations, labeled training data as well as records of previously matched intents and user inputs can be stored in the intent model store 150. The intent model store 150 can also store a list of available intents, which are tasks that the morphing interface system's tasks 130 can perform for the user in response to user input. The intent model store 150 can also store a list of unavailable intents, which are tasks that the morphing interface system 130 is currently unable to perform but have been identified as tasks independent of the available intents. In addition, the intent model store 150 can store custom intents built by users, which are only available to those users. For example, the user input string "open device" can not be in the list of globally available intents, but a user can have created this intent for their own use and the intent logic will be stored in the intent model store 150.
[0042] In one implementation, the interface store 155 stores program code for a user interface for each available intent that can be performed by the morphing interface system 130. The interfaces stored by the interface store 155 can include a layout for displaying the interface on the client device 110, instructions for performing the intent, and a list of entities associated with populating the layout and performing the intent. In various implementations, the user interface can be a customized interface that has been tailored for each potential intent. In other implementations, the interface store 155 can contain customized interfaces for custom intents designed by users and only available to those users.
[0043] The entity recognition module 145 predicts a set of entity values associated with a given user input. The entity values can be used to execute an intent that matches the user input. In various implementations, the entity recognition module 145 takes as input a user input string, an associated NLP signal from the NLP pipeline 135, and a matched intent from the intent matching module 140. The entity recognition module 145 can also access the interface store 155 to use an interface associated with the matched intent as input to obtain a list of entity values needed by the morphing interface system 130 to execute the intent. The entity recognition module 145 can apply trained computer models to extract a set of values from the user input string and associate the extracted values with entities of the matched intent. In one implementation, the entity recognition module 145 first extracts high-level entity values from the input string and then detailed entity values. For example, the entity recognition module 145 can apply a model that determines that the user input string includes a title and can apply different models to predict that the title is a movie title, a book title, etc. In one implementation, one or more computer models applied by the entity recognition module 145 are classifiers or sequence taggers trained on example training data that includes example user input strings with tagged entity values. These classifiers or sequence taggers can be further trained on large amounts of unstructured, untagged text information from the internet using multiple objectives (language modeling, autoencoding, etc.) to incorporate real-world knowledge as well as an understanding of syntax, grammar, and semantics.
[0044] In other example implementations, the entity recognition module 145 can apply rule-based methods to extract a set of values from the user input string, such as matching against regular expression patterns. This can again help the entity recognition module 145 to quickly customize the extraction of values for new, custom intents designed by users.
[0045] The models and training data applied by the entity recognition module 145 can be stored in the entity recognition model store 160. The entity recognition model store 160 can also include tagged training data used to train the computer models used by the entity recognition module 145.
[0046] In some implementations, the entity recognition module 145 and the intent recognition module 140 can be the same system. That is, the entity recognition module 145 and the intent recognition module 140 can be configured as a joint intent and entity recognition system such that the two systems can make more accurate decisions in collaboration.
[0047] Example Morphing Interface System
[0048] Figure 2is a high-level diagram of the interaction between components of the morphing interface system 130 according to example implementations. The morphing interface system 130 receives a user query 210. The user query can be, for example, a complete sentence or concept expressed by a user in the form of typed text or speech audio, or a partial sentence or phrase. In implementations where the input is received as an audio file or audio stream, automatic speech recognition or other types of speech models can be used to produce an input string representing the input, e.g., represented as text. While the user is still providing input, the morphing interface system 130 can begin responding to the user through a display interface. Thus, in some cases, the user query 210 received by the morphing interface system 130 can be only the first portion of the user's input, e.g., one word or set of words.
[0049] The user query 210 is provided as input to the NLP pipeline 135, which analyzes the user query 210 and outputs a corresponding NLP signal 220. The NLP signal 220 and the user query 210 are provided to the intent recognition module 140. The intent recognition module 140 outputs a predicted user intent 230. That is, the intent recognition module 140 predicts what intent the user is requesting or intends to perform. The predicted intent 230 or function, the NLP signal 220, and the user query 210 are provided to the entity recognition module 145, which generates extracted entity values 240 associated with the predicted intent. The morphing interface system 130 can use predicted intent information 250 about the predicted user intent, the extracted entities 240, and additional generated metadata to enable display of a user interface (on a screen of a computing device, e.g., a client device) related to the system prediction corresponding to the user intent, and to populate fields of the user interface with the extracted entity values. Thus, the interface to be generated and enabled (or provided) for display on the client device can advantageously begin changing substantially in real-time.
[0050] In some implementations, the components of the morphing interface system 130 can be configured to enable display of the user interface in a different manner than Figure 2in the manner shown in the example of FIG. 1. In one implementation, the metamorphic interface system can be configured to include a feedback loop between the intent recognition module 140 and the entity recognition module 145. For example, the intent recognition module 140 can provide information about a predicted intent to the entity recognition module 145, the entity recognition module 145 can use the information about the predicted intent as input to identify entities and potential entity types in the user query 210, and information about the identified entities can be provided back to the intent recognition module 140 for regenerating a prediction of the intent that should be associated with the user’s query 210. In some implementations, the entity recognition module 145 can analyze the NLP signals or user input and predict entities associated with the user input, then provide the predicted entities and entity types to the intent recognition module 140 in addition to the input and NLP signals. In this case, the intent recognition module 140 can then use the predicted entity information to predict the intent type associated with the user input. A feedback loop between the intent recognition module 140 and the entity recognition module 145 can also exist in the present implementation (i.e., the intent recognition module 140 can send predicted intent information back to the entity recognition module 145 to improve or add to existing predictions about entities). In some implementations where the entity recognition module 145 receives input data before the intent recognition module 140, the intent recognition module 140 can filter the extracted entities provided by the entity recognition module 145 into entities that correspond to a predicted intent.
[0051] In other example implementations, one module can be configured to perform the functions of both the intent recognition module 140 and the entity recognition module 145. For example, a model can be trained to perform both intent recognition and entity recognition. In another example implementation, the metamorphic interface system 130 can include a sub-model associated with entity recognition for each intent type (i.e., each domain). That is, the metamorphic interface system 130 can store different entity recognition models for determining entity values associated with each potential intent type, and can use transfer learning to automatically create a model for a new potential intent type based on a collection of past entity recognition models. For example, if the intent recognition module 140 predicts an intent to order a pizza, the entity recognition module 145 can then access and use an entity recognition model trained to identify entities associated with an intent to order food. In another example implementation, the intent recognition module 140 can be configured in the form of a hierarchical model, where a first model infers a higher-level domain (e.g., “food”), and a sub-model of that domain then infers a specific intent of the user within that predicted domain (e.g., whether the user wants to reserve a table, order takeout, search for a recipe, etc.).
[0052] In another example implementation, the morphing interface system can not include the NLP pipeline 135. In such an implementation, the intent identification module 140 and the entity identification module 145 are trained to predict an intent and determine an entity directly based on the user query 210.
[0053] Figure 3 is a detailed block diagram illustrating example NLP signals according to an example implementation. The NLP pipeline 135 can include tokenization, part-of-speech (POS) tagging, text chunking, semantic role labeling (SRL), and coreference resolution functionality for generating various NLP signals from a user input string. Some other example NLP signals include lemmatization, stemming, dependency parsing, entity extraction, and stopword identification. In various implementations, different combinations of NLP signals can be used in the NLP pipeline 135, and the signals can be determined in various orders. For example, in another example implementation, the NLP pipeline 135 can determine NLP signals in the order of tokenization, stemming, lemmatization, stopword identification, dependency parsing, entity extraction, chunking, SRL, and then coreference resolution.
[0054] Figure 4A is a flowchart illustrating an example process of generating an interface from a user input according to an example implementation. The morphing interface system 130 receives 405 a first portion of a user input. The first portion of the user input can be, for example, one or more words at the beginning of a sentence, and can be received by the morphing interface system 130 in a variety of input forms including text or speech input. The morphing interface system 130 generates 410 natural language processing signals based on the received first portion of the user input. The NLP pipeline 135 can generate the natural language processing signals using a variety of analysis techniques including tokenization, part-of-speech (POS) tagging, dependency parsing, entity extraction, stemming, lemmatization, stopword identification, text chunking, semantic role labeling, and coreference resolution.
[0055] The morphing interface system 130 selects 420 an intent that matches the first portion of the user input. In some implementations, the intent identification module 140 applies a trained computer model to predict which intent is most appropriate for responding to the first portion of the user input. That is, the intent identification module 140 selects an intent that the received user input implies.
[0056] The morphing interface system 130 extracts 425 entity values associated with the predicted intent from the first portion of the user input. In one implementation, the entity recognition module 145 applies a trained computer model to extract relevant values from the received user input. In some implementations, the morphing interface system 130 is configured to include a feedback loop such that information about the extracted entities is sent to the intent recognition module 140 as additional input. In some implementations, this can also include an automatic retraining and self-improvement loop. The morphing interface system 130 obtains 430 an interface associated with the selected intent. In some implementations, the selected interface is used in the process of extracting entities associated with the intent. For example, the interface associated with the intent can include input fields for values relevant to that particular intent, and the entity recognition module 145 can use information about the input fields to identify values from the user input. The extracted entity values associated with the selected intent are used to populate 435 the interface. In some implementations, the interface can be only partially populated, for example if the portions of the user input received so far only include some of the information needed to complete the input fields included in the interface layout. For example, in the case where the user pauses after providing input including “book me a flight,” the morphing interface system 130 can highlight a calendar and prompt the user with a query such as “when do you want to board the plane?” to receive more user input to further populate the user interface. The morphing interface system 130 displays 440 the populated interface to the user, for example via the client device 110.
[0057] In some implementations, when multiple intents have similar or equivalent predicted likelihoods that apply to the first portion of the user input, the morphing interface system 130 can select 420 more than one applicable intent. For example, the first portion of the user input can be "give me some coffee," and the morphing interface system 130 can determine that an intent to purchase a bag of coffee beans is equally likely as an intent to order a cup of coffee from a coffee shop to be a likely applicable response. That is, the morphing interface system 130 can determine that the user's intent is equally likely to be beverage delivery and grocery ordering. In this case, the morphing interface system 130 can extract 425 entity values associated with the two or more intents that have equivalent likelihoods, and can obtain 430 preview interfaces associated with the multiple likely intents, populate 435 the preview interfaces with the extracted entity values associated with the intents, and provide the populated preview interfaces for display 440 to the user at the client device 110. In such implementations, the morphing interface system can wait for input from the user, selecting between one of the multiple likely intents before continuing to analyze additional user input. In some implementations, the morphing interface system 130 can store information about the user's selection of a preview interface provided to the user input, so that in the future, the system can increase the likelihood that the user selects a particular intent based on patterns of the user input. Thus, as user preferences and input patterns are stored, after some history of interacting with the morphing interface system 130, fewer intent previews can be presented to the user.
[0058] Figure 4B is a flowchart illustrating a process of morphing an interface as additional user input is received, according to example implementations. Upon receiving a portion of user input for analysis, the morphing interface system 130 can enable (or provide) an interface for display at the client device 110. Enabling (or providing) an interface for display at the client device 110 can include providing code, commands, and / or data to an operating system of the client device 110 to display a user interface corresponding to a morphing structure and / or intent defined by the morphing interface system 130. As additional user input is received, the morphing interface system 130 repeatedly reevaluates whether the selected intent is still the best match for the growing string of user input. If the morphing interface system determines 130 that the same intent is still applicable, additional entity values can be extracted from the augmented user input. However, if the morphing interface system determines 130 that a different intent is more applicable given the additional user input, the system will enable (or provide) display of a change (e.g., a "morph") to the user interface for the different, more applicable intent. The change in the user interface can be, for example, a gradual visual transition to a full refresh of the user interface on the screen (e.g., of the client device 110). Moreover, the morphing interface system can enable the user interface to change any number of times as user input is added over time.
[0059] In particular, after a first portion of user input is received and analyzed (as described in Figure 4A , the morphing interface system 130 continues to analyze the growing string of user input as additional portions of user input are received (as shown in Figure 4B ). The morphing interface system 130 receives 445 additional portions of user input. For example, the first portion of user input can be the first word of a sentence, and Figure 4B the process of FIG. 4 can begin in response to receiving the second word of the user input sentence.
[0060] The morphing interface system 130 generates 450 a natural language processing signal based on the combination of previously received portions of user input and the additional portions of user input. For example, an intent matching the combination of user input is selected 455 by the intent matching module 140.
[0061] The morphing interface system 130 determines 460 whether the matching intent is the same as the most recently matched intent or whether the matching intent is a new intent. A new intent can be an intent that is different from the intent associated with the interface currently displayed at the client device 110.
[0062] If the intent is the same intent, the morphing interface system 130 extracts 465 additional entity values associated with the intent from the combination of previous user input and the current (i.e., most recent) user input. The interface is populated 470 with the additional entity values associated with the intent.
[0063] If the intent is not the same intent as the previous intent matched with the previous set of user input, the morphing interface system 130 extracts 475 entity values associated with the new intent from the combination of user input. For example, an interface associated with the new intent is obtained 480 from the interface store 155. The obtained interface is populated 485 with the extracted entity values associated with the new intent.
[0064] Regardless of whether the interface is the same as before or a newly obtained interface, the morphing interface system 130 causes (or provides for) display 490 of the populated interface to the user, e.g., via the client device 110. The process shown in FIG. 4 can be repeated each time additional portions of user input are received 445 by the morphing interface system 130. Figure 4B
[0065] The following Figures 5A-5G illustrates an example of an interface for an intent that is executed upon receiving additional user input, according to an implementation. In one implementation, the interface is a user interface that is presented for display on a screen of a computing device (e.g., a client device 110 such as a smartphone, tablet, laptop, or desktop computer). Figures 5A-5G An example is shown in which user input received, e.g., via client device 110, has been matched to a general intent of booking a flight (i.e., a function or user request). In response to receipt of additional user input, it is determined that the displayed interface morphs and changes layout as additional entity values associated with the selected interface. However, in Figures 5A-5G the example of FIG. 6, the morphing interface system 130 does not determine that a new intent better matches the user input, and the displayed layout accordingly remains associated with the flight booking interface.
[0066] Figure 5A A first layout displayed for an interface associated with a flight booking intent is shown in accordance with example embodiments. The morphing interface system 130 selects the interface 500 associated with flight booking based on received user input 530A. In Figures 5A-5E the example of FIG. 6, the morphing interface system 130 selects the interface 500 associated with flight booking based on received user input 530A. In Figure 5A the example of FIG. 6, the morphing interface system 130 selects the interface 500 associated with flight booking based on received user input 530A. In Figure 5B the example of FIG. 6, the morphing interface system 130 selects the interface 500 associated with flight booking based on received user input 530A. In Figures 5A-5G the example of FIG. 6, the morphing interface system 130 selects the interface 500 associated with flight booking based on received user input 530A. In Figures 6A-6D the example of FIG. 6, the morphing interface system 130 selects the interface 500 associated with flight booking based on received user input 530A. In Figure 5A Example types of user input 530 shown in FIG. 6 include voice input 510 and typing input 520.
[0067] In Figure 5AIn the example of FIG. 5, the morphing interface system 130 receives an initial user input 530A comprising the text string "book me a flight." The morphing interface system 130 determines that the user input 530A is most likely associated with the intent of booking a flight, and the flight booking interface 500 is accordingly proximately and instantaneously displayed. The interface 500 can include a widget 540. A widget can be a part of the layout of an interface, the widget displaying or collecting information related to an entity associated with the interface. In some cases, the widget 540 can be an input field (e.g., a text field, a checkbox, or other data input format). The interface 500 can display the widget 540, which can be populated with an entity value determined by the entity recognition module 145. In various implementations, the interface 500 can even display some or all of the widgets 540 associated with the interface 500 before the entity recognition module 145 has determined a value for populating the input field. For example, Figure 5A includes a widget 540A comprising a space for inputting a destination value for the flight booking. Figure 5A Similarly includes a departure and return date widget 540B comprising a space for inputting a date value for the flight booking and a widget 540C indicating a number of passengers, for which the entity recognition module 145 has predicted that the value of "1 passenger, economy" will be the most likely input value, and these values are accordingly represented as populated in the widget 540C.
[0068] Figure 5B A second layout for display of an interface associated with a flight booking intent is shown in accordance with example implementations. In Figure 5B the user input 530B includes additional information. In particular, the user has added more input such that the user input 530B now includes "book me a flight from San Francisco." The morphing interface system 130 determines that the selected intent should still be flight booking, and identifies additional entity value information to further populate the widgets 540 in the interface 500. Accordingly, the user interface 500 morphs from Figure 5A the layout of Figure 5B the layout shown in FIG. 5B. For example, Figure 5B the interface 500 shown in FIG. 5C includes the departure city of "San Francisco" inputted for the flight departure city information in the widget 540A. In some implementations, morphing an interface from one layout to another can include the display of an animation, such as a moving expansion of a section of the layout, which can be further populated with newly received information from the user input as new entity information is extracted from the user input.
[0069] Figure 5C A third layout for display of an interface associated with a flight booking intent is shown in accordance with example implementations. In Figure 5C, user input 530C also includes more additional information. In particular, the user has added input such that user input 530C includes "Book me a flight from San Francisco to Los Angeles." Transformed interface system 130 determines that the selected intent should still be flight booking. Transformed interface system 130 identifies additional entity value information for further populating widgets 540 in interface 500, including, for example, the destination "Los Angeles" in the destination field of widget 540A. Thus, user interface 500 changes from Figure 5B The layout shown is morphed into Figure 5C The layout shown.
[0070] Figure 5D A fourth layout for an interface display associated with a flight booking intent is shown in accordance with an example embodiment. Figure 5D , user input 530D includes additional information. In particular, the user adds more input, such that user input 530D now includes "book me a flight from San Francisco to Los Angeles on April 22". The transformed interface system 130 determines that the selected intent should still be flight booking, and extracts additional entity value information to further populate the interface 500. Accordingly, the user interface 500 changes from Figure 5C The layout shown is morphed into Figure 5D For example, Figure 5D An expanded widget 540D is shown showing departure date information regarding the flight requested by the user, with April 22 selected.
[0071] Figure 5E A fifth layout of an interface display associated with a flight booking intent is shown according to an example embodiment. When the morphing interface system 130 determines that the entity values associated with the selected intent have all been extracted and applied to the intent, the intent can be executed to generate a response to the user. The layout of the interface 500 can then include the display of the response. For example, when the entity identification module 145 identifies the values associated with all entities required to book a flight, the intent is executed and the layout of the interface 500 is presented, which identifies possible flights that meet the criteria specified in the user input 530. For example, the user interface 500 changes from Figure 5D The layout shown is morphed into Figure 5E The user can then select a flight from the presented options (e.g., from widgets 540E, 540F, 540G, 540H, and 540I that match possible flights under the specified criteria).
[0072] Figure 5F A sixth layout for an interface display associated with a flight booking intent is shown in accordance with an example embodiment. Figure 5FIn the example of FIG. 5B, the interface 500 displays selected flight information so that the user can confirm the data before submitting an order to purchase the flight. In some implementations, the malleable interface system 130 can use various personalization techniques to determine values for entities that are relevant to the user. In one implementation, the malleable interface system 130 can store user profile information for use in performing intents. For example, the malleable interface system 130 can store a user name, credit card number, home and work locations, etc. For example, in Figure 5F In the example of FIG. 5B, the interface 500 displays selected flight information so that the user can confirm the data before submitting an order to purchase the flight. In some implementations, the malleable interface system 130 can use various personalization techniques to determine values for entities that are relevant to the user. In one implementation, the malleable interface system 130 can store user profile information for use in performing intents. For example, the malleable interface system 130 can store a user name, credit card number, home and work locations, etc. For example, in
[0073] Figure 5G A portion of a user interface displayed on a screen is shown. The user interface portion here is an email confirmation of performing a flight reservation intent, according to example implementations. Figure 5G The example of FIG. 5B depicts an email 550 that a user receives confirming that the user purchased a flight. For certain intents, such an email confirmation can not be included as part of the intent execution.
[0074] The following Figures 6A-6D An example of an interface performing an intent as additional user input is received, according to implementations, is shown. Figures 6A-6D An example is shown in which the user input 530E has been matched to an intent to order a pizza. The layout of the displayed interface is malleated and changes as additional entity values associated with the selected interface are determined in response to receipt of additional user input. In Figures 6A-6D In the example of FIG. 5B, the interface 500 displays selected flight information so that the user can confirm the data before submitting an order to purchase the flight. In some implementations, the malleable interface system 130 can use various personalization techniques to determine values for entities that are relevant to the user. In one implementation, the malleable interface system 130 can store user profile information for use in performing intents. For example, the malleable interface system 130 can store a user name, credit card number, home and work locations, etc. For example, in
[0075] Figure 6A Receipt of a first portion of user input associated with a pizza ordering intent, according to example implementations, is shown. The malleable interface system 130 selects the interface 500 associated with pizza ordering based on the user input 530E that has been received. In Figure 6AIn the particular example, the morphing interface system 130 receives an initial user input 530E that includes the text string (on 520) "I want to order pizza" or a voice input (on 510) that corresponds to "I want to order pizza." The morphing interface system 130 determines that the user input 530E is most likely associated with the intent to order pizza and the interface 500 begins to transition (e.g., morph) to display a pizza ordering layout accordingly (as shown by the graphical example 501 on the interface 500).
[0076] Figure 6B A layout is shown for an interface associated with ordering pizza, according to an example implementation, as the interface is transitioning from Figure 6A to Figure 6B As shown, the interface 500 can include a widget 540 that can be populated with entity values determined by the entity recognition module 145. For example, Figure 5A The widget 540L is included that includes a space for inputting a pizza restaurant. Additional information can also begin to appear on the morphed screen, for example, the price of the pizza and the delivery time for the pizza from the specified restaurant. In Figure 6B the example, the entity recognition module 145 has predicted a pizza restaurant that the user can want to order from and inputs information about the restaurant in the widget 540 of the user interface 500. Figure 6B
[0077] Figure 6C A layout is shown for an interface associated with purchasing a pizza t-shirt, according to an example implementation, as the interface continues to morph from Figure 6A to Figure 6C In this example, the user input 530F includes additional information. In particular, the user has added input so that the user input 530F now includes "I want to order a pizza t-shirt." The intent matching module 140 analyzes the additional user input and determines that the previously selected pizza ordering intent is no longer the most applicable intent and that the new most associated intent is the t-shirt ordering intent. The morphing interface system 130 selects an interface 500 that is appropriate for the t-shirt purchase intent and morphs the display at the client device 110 to display a layout associated with the selected intent. For example, Figure 6C Instead of displaying pizza restaurant suggestions, the interface 500 of Figure 6D is morphed from the interface displaying pizza ordering shown in Figure 6C to the interface displaying pizza t-shirt purchase options shown in The widget 540M associated with the example t-shirt purchase intent can include a picture of a pizza t-shirt that the user can then select to purchase.
[0078] Figure 6D 1 shows additional example interfaces that may be associated with a T-shirt purchase intent according to an example embodiment. For example, once the morphing interface system 130 determines that all entity values required to execute the intent are available, the process may be taken over by the user. Figure 6D In the example of , a user can select one of a series of pizza-themed T-shirts (as shown in interface 500A), the user can view additional information about the selected item (as shown in interface 500B), and the user can confirm the order details and place an order for the pizza T-shirt.
[0079] Figures 5A-5G as well as Figures 6A-6D The example advantageously reflects a user interface that changes rapidly (e.g., deforms) as the received user input is gradually expanded with additional information, and these user interfaces change through substantially (or almost) simultaneous refreshes. Unlike conventional systems, here the user does not need to parse through potential recommendations, which are presented to the user by conventional systems when the user provides user input. Furthermore, unlike conventional systems, the user does not need to wait for the complete selection of the user input to begin viewing what user interface is displayed for presentation at the moment corresponding to the enhanced input string (e.g., currently provided). Furthermore, unlike conventional systems, the user interface enabled for display on the screen of the computing device begins to reflect the partial input almost instantaneously (or immediately), and as additional wording is contextually added to the user input, it quickly evolves to reflect the current input, and ends with the appropriate final user interface corresponding to the complete user input. That is, the user interface enabled for display at text input TX0+TX1 is substantially immediately updated from the user interface enabled for display of the original user input TX0.
[0080] Example Computing System
[0081] Figure 7 is a block diagram illustrating components of an example machine capable of reading instructions from a machine-readable medium and executing the instructions in one or more processors (or controllers) according to an example embodiment. Figure 7 A diagrammatic representation of the deformation interface system 130 is shown in the example form of a computer system 700. The computer system 700 may be configured to execute instructions 724 (e.g., program code or software) for causing the machine to perform any one or more of the methods (or processes) described herein. In alternative embodiments, the machine operates as a standalone device or as a connected (e.g., networked) device connected to other machines. In a networked deployment, the machine may operate in the capacity of a server or a client user machine in server-client user network environment, or as a peer machine in a peer-to-peer (or distributed) network environment.
[0082] The machine can be a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a smartphone, an internet appliance (la), a network router, switch or bridge, or any machine capable of executing instructions 724 (sequential or otherwise) that specify actions to be taken by that machine. Further, while only a single machine is illustrated, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute instructions 724 to perform any one or more of the methodologies discussed herein.
[0083] Example computer system 700 includes one or more processing units (typically processors) 702. Processors 702 are, for example: central processing units (CPUs), graphical processing units (GPUs), digital signal processors (DSPs), controllers, state machines, one or more application specific integrated circuits (ASICs), one or more radio-frequency integrated circuits (RFICs) or combinations of such. Processors execute an operating system of the computing system 700. The computer system 700 also includes a main memory 704. The computer system can include a storage unit 716. The processor 702, the memory 704, and the storage unit 716 communicate via a bus 708.
[0084] Additionally, the computer system 706 can include a static memory 706, a graphical display 710 (e.g., to drive a plasma display panel (PDP), a liquid crystal display (LCD), or a projector), an alphanumeric input device 712 (e.g., a keyboard), a cursor control device 714 (e.g., a mouse, a trackball, a joystick, a motion sensor, or other pointing instrument), a signal generation device 718 (e.g., a speaker), and a network interface device 720, which are also configured to communicate via the bus 708.
[0085] The storage unit 716 includes a machine-readable medium 722 on which is stored instructions 724 (e.g., software) embodying any one or more of the methodologies or functions described herein. The instructions 724 can include, for example, instructions to implement the NLP pipeline 135, the function matching module 140, and / or the entity recognition module 145. The instructions 724 can also reside completely, or at least partially, within the main memory 704 and / or within the processor 702 during execution thereof by the computer system 700, the main memory 704 and the processor 702 also constituting machine-readable media. The instructions 724 can be transmitted or received via the network interface device 720 over the network 726 such as the network 120. Further, for a client device (or user device), the received instructions can be instructions from a server system that enable functionality on the client device. For example, how a user interface will be displayed can include receiving code as to how the user interface should be enabled (e.g., presented) for display based on how the code properly interacts with the operating system of the client device.
[0086] While the machine-readable medium 722 is illustrated in an example embodiment to be single medium, the term“machine-readable medium” should be taken to include a single medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) able to store the instructions 724. The term“machine-readable medium” shall also be taken to include any medium that is capable of storing the instructions 724 for execution by a machine, and that cause the machine to perform any one or more of the methodologies
[0087] Additional Considerations
[0088] The foregoing description of implementations has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the patent rights to the precise form disclosed. Many modifications and variations are possible in light of the above disclosure.
[0089] Some portions of this specification are presented in terms of algorithms and symbolic representations of operations on data stored as bits or binary digital signals within a machine memory. These algorithms and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. These and similar operations, while commonly based on logical steps, are not meant to be limiting, unless specifically so limited. Furthermore, in the context of this specification, unless otherwise specifically stated, the term“some” refers to one or more.
[0090] Any of the steps, operations, or processes described herein can be performed or implemented with one or more hardware modules or software modules, alone or in combination with other devices. In one embodiment, a software module is implemented with a computer program product comprising a computer-readable medium containing computer program code, which can be executed by one or more computer processors for performing any or all of the steps, operations, or processes described.
[0091] Embodiments can also relate to an apparatus for performing the operations herein. This apparatus can be specially constructed for the required purposes, and / or it can comprise a computer device selectively activated or reconfigured by a computer program stored in the computer. Such a computer program can be stored in a non-transitory, tangible computer readable storage medium, or any type of media suitable for storing electronic instructions, which can be coupled to a computer system bus. For example, a computer device coupled to a data storage device that stores the computer program can correspond to a special-purpose computing device. Additionally, any of the computing systems mentioned in this specification can include a single processor or can be architectures employing multiple processor designs for increased computing capability.
[0092] Embodiments can also relate to a product produced by the computing process described herein. Such a product can include information resulting from the computing process, where the information is stored on a non-transitory, tangible computer readable storage medium and can include any embodiment of a computer program product or other data combination described herein.
[0093] Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it can not have been selected to delineate or circumscribe the patent rights of the application. It is intended that the scope of the patent rights be defined by any claims issued by the patenting authority and that meanings assigned to features described herein are consistent with the principles of patent law.
Claims
1. A computer-implemented method comprising: receiving, from a user device, a first user input comprising an input string; generating a set of natural language processing signals based on the first user input; selecting an intent matching the first user input, the selection based on the natural language processing signals, the intent corresponding to a computer-executable function; identifying a first interface associated with the intent; extracting, from the first user input, a set of values associated with entities of the first interface; enabling the first interface for display on the user device, the display of the first interface including values from the set of values; predicting one or more intents associated with the selected intent; updating the first interface to include one or more items associated with the predicted one or more intents; and when the first user input is received, enabling a gradual visual transition from the first interface to an updated first interface for display on the user device, the gradual visual transition beginning immediately reflecting the first interface and rapidly evolving to reflect current input as additional phrasing is contextually added to the first user input, and ending with the updated first user interface corresponding to the full first user input.
2. The computer-implemented method of claim 1, further comprising: receiving, from the user device, a second user input comprising a text string; generating an updated set of natural language processing signals based on a combination of the first user input and the second user input; selecting an intent matching the combination of the first user input and the second user input, the selection based on the updated set of natural language processing signals; identifying a second interface associated with the newly selected intent; extracting, from the combination of the first user input and the second user input, a second set of values associated with entities of the second interface; and enabling the second interface for display on the user device, the second interface for display including values from the second set of values.
3. The computer-implemented method of claim 1, wherein, The first user input is a voice input.
4. The computer-implemented method of claim 1, wherein, Selecting an intent matching the user input comprises comparing the first user input to one or more previously received user input strings, and in response to the first user input matching the one or more previously received user input strings, selecting an intent selected in response to at least one of the previously received user input strings.
5. The computer-implemented method of claim 1, wherein, Selecting an intent matching the user input comprises applying a trained computer model to predict a most applicable intent.
6. The computer-implemented method of claim 1, wherein, The interface includes an associated set of entities.
7. The computer-implemented method of claim 2, wherein, The first interface and the second interface are the same interface.
8. The computer-implemented method of claim 2, wherein, The first interface and the second interface are different interfaces.
9. The computer-implemented method of claim 1, wherein, The input string is a text string.
10. The computer-implemented method of claim 1, wherein, The input string is an audio input.
11. A computer system comprising: one or more computer processors to execute computer program instructions; and a non-transitory computer-readable storage medium comprising stored instructions executable by at least one processor, the instructions, when executed, causing the processor to: receiving, from a first user device, a first user input comprising an input string; generating a set of natural language processing signals based on the first user input; selecting an intent matching the first user input, the selection based on the natural language processing signals, the intent corresponding to a computer-executable function; identifying a first interface associated with the intent; extracting, from the first user input, a set of values associated with entities of the first interface; enabling the first interface for display on the user device, the display of the first interface including values from the set of values; predicting one or more intents associated with the selected intent; updating the first interface to include one or more items associated with the predicted one or more intents; and when the first user input is received, enabling a gradual visual transition from the first interface to an updated first interface for display on the user device, the gradual visual transition immediately starting to reflect the first interface and rapidly evolving to reflect current input as additional phrasing is contextually added to the first user input, and ending with the updated first user interface corresponding to the full first user input.
12. The computer system of claim 11, further comprising stored instructions that, when executed, further cause the processor to: receive, from the user device, a second user input comprising a text string; generate an updated set of natural language processing signals based on a combination of the first user input and the second user input; select an intent matching the combination of the first user input and the second user input, the selection based on the updated set of natural language processing signals; identify a second interface associated with the newly selected intent; extract, from the combination of the first user input and the second user input, a second set of values associated with entities of the second interface; and enable the second interface for display on the user device, the second interface for display including values from the second set of values.
13. The computer system of claim 11, wherein, The first user input is a voice input.
14. The computer system of claim 11, wherein, The instructions that cause the processor to select an intent matching the user input include instructions to compare the first user input to one or more previously received user input strings; and in response to the first user input matching the one or more previously received user input strings, select an intent that was selected in response to at least one of the previously received user input strings.
15. The computer system of claim 11, wherein, The instructions that cause the processor to select an intent matching the user input include instructions to cause the processor to apply a trained computer model to predict a most applicable intent.
16. The computer system of claim 11, wherein, The interface includes an associated set of entities.
17. The computer system of claim 12, wherein, The first interface and the second interface are the same interface.
18. The computer system of claim 12, wherein, The first interface and the second interface are different interfaces.
19. The computer system of claim 11, wherein, The input string is a text string.
20. The computer system of claim 11, wherein, The input string is an audio input.
Citation Information
Patent Citations
Intelligent automated assistant
WO2014197635A2
Information processing device and information processing method
WO2019107145A1