System and method for intelligent conversation based on knowledge tracing
By receiving language understanding results and multimodal information processing, dynamically updating dialogue strategies is solved, and the problem of insufficient adaptability of traditional dialogue systems is achieved, and more attractive and effective user interaction is achieved.
Patent Information
- Application Number
- CN202080053996.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-17
- Filing Date
- 2020-06-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-06-09
AI Technical Summary
Traditional computer-assisted dialogue systems are difficult to adapt to the personalized needs of different human users, and may cause users to lose interest or generate stimulation during the conversation, and lack the ability to combine adaptability and dynamic information.
By receiving language understanding results and their evaluation, based on the parameter set of multiple probability update dialogue strategies, the knowledge tracking unit, multiple probability estimators and information status updaters are used to realize adaptive dialogue management, and combine multimodal information processing and dynamic information tracking to dynamically update dialogue strategies to improve interaction effect.
It realizes more attractive and effective dialogue management, and can adjust conversation content and strategies according to users' personalized needs and improve user interaction experience.
Smart Images

Figure CN114270435B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to the following patent applications: U.S. Provisional Patent Application 62 / 862,268, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503582); U.S. Provisional Patent Application 62 / 862,253, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503572); U.S. Provisional Patent Application 62 / 862,257, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503574); U.S. Provisional Patent Application 62 / 862,261, filed on June 19, 2019 (Attorney Docket No.: 047437 - 0503575); U.S. Provisional Patent Application 62 / 862,264, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503578); U.S. Provisional Patent Application 62 / 862,265, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503581); U.S. Provisional Patent Application 62 / 862,273, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503579); U.S. Provisional Patent Application 62 / 862,275, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503580); U.S. Provisional Patent Application 62 / 862,279, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503584); U.S. Provisional Patent Application 62 / 862,282, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503585); U.S. Provisional Patent Application 62 / 862,286, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503586); U.S. Provisional Patent Application 62 / 862,290, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503587), and U.S. Provisional Patent Application 62 / 862,296, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503589), the contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] This teaching generally relates to computers. More specifically, this teaching relates to human - machine dialogue management. Background Art
[0004] With the advancement of artificial intelligence technology and the explosive growth of Internet-based communications due to ubiquitous Internet connectivity, computer-aided dialogue systems have become increasingly popular. For example, more and more call centers are deploying automated dialogue robots to handle customer calls. Hotels are starting to install various kiosks that can answer questions from tourists or guests. In recent years, automated human-machine communication in other fields has also become increasingly popular.
[0005] Traditional computer-aided dialogue systems are usually pre-programmed with certain dialogue content, such as questions and answers based on well-known conversation patterns in related fields. Unfortunately, some conversation patterns may be suitable for some human users, but may not be suitable for other human users. In addition, human users may go off-topic during the conversation, and continuing with a fixed conversation pattern without considering what the user says may cause irritation or loss of interest, which is not desirable.
[0006] When planning a conversation, human designers usually need to manually create the content of the conversation based on known knowledge, which is time-consuming and tedious. Considering the need to create different conversation patterns, even more labor is required. When creating dialogue content, any deviation from the designed conversation pattern may need to be noted and used to determine how to continue the conversation. Previous dialogue systems have not effectively solved such problems.
[0007] With the latest developments in the field of AI, observed dynamic information can be adaptively incorporated into learning and used to guide the progress of human-machine interaction sessions. How to develop a knowledge representation that can incorporate dynamic information in different dimensions and sometimes in different modalities is a challenging problem. Since this knowledge representation is the basis for the dynamic conversation process between humans and machines, it needs to be fully configured to support adaptive conversation in a relevant manner.
[0008] In order to converse with humans, an automated dialogue system may need to achieve different levels of understanding of what humans say linguistically, what the semantics of what is said are, sometimes the emotional state of the person, and the mutual causal relationship between what is said and the conversation environment. Traditional computer-aided dialogue systems are not sufficient to solve such problems.
[0009] Therefore, methods and systems are needed to address such limitations. Summary of the Invention
[0010] The teachings disclosed herein relate to methods, systems, and programming for advertising. More specifically, the teachings relate to methods, systems, and programming related to exploring the sources of advertising and their utilization.
[0011] In one example, a method for adaptive dialogue management implemented on a machine including at least one processor, a memory, and a communication platform capable of connecting to a network. Receive a language understanding result and its evaluation. The language understanding result is derived from the utterance of a user participating in a dialogue, which is targeted at a topic and governed by a dialogue strategy. The evaluation is obtained for the expected result represented in the dialogue strategy. Derive a plurality of probabilities based on the language understanding result and the associated evaluation. Update a set of parameters associated with the dialogue strategy based on the plurality of probabilities, wherein a first set of parameters parameterizes the dialogue strategy with respect to the user and characterizes the effectiveness of the dialogue with the user under the dialogue strategy.
[0012] In different examples, a system for adaptive dialogue management is disclosed, which includes a knowledge tracking unit, a plurality of probability estimators, and an information state updater. The knowledge tracking unit is configured to receive a language understanding result together with its evaluation, wherein the language understanding result is derived from the utterance of a user participating in a dialogue targeted at a topic, which is governed by a dialogue strategy, and the evaluation is obtained for the expected result in the dialogue strategy. The plurality of probability estimators are configured to determine a plurality of probabilities based on the language understanding result and the associated evaluation. The information state updater is configured to update a first set of parameters associated with the dialogue strategy based on the plurality of probabilities, wherein the first set of parameters parameterizes the dialogue strategy with respect to the user and characterizes the effectiveness of the dialogue with the user under the dialogue strategy.
[0013] Other concepts relate to software for implementing this teaching. A software product according to this concept includes at least one machine-readable non-transitory medium and information carried by the medium. The information carried by the medium can be executable program code data, parameters associated with the executable program code, and / or information related to the user, request, content, or other additional information.
[0014] In one example, a machine-readable non-transitory and tangible medium on which data for adaptive dialogue management is recorded, wherein when the medium is read by a machine, the machine is caused to perform a series of steps required by the disclosed method for dialogue management. Receive a language understanding result together with its evaluation. The language understanding result is derived from the utterance of a user participating in a dialogue, which is targeted at a topic and governed by a dialogue strategy. The evaluation is obtained for the expected result represented in the dialogue strategy. Derive a plurality of probabilities based on the language understanding result and the associated evaluation. Update a set of parameters associated with the dialogue strategy based on the plurality of probabilities, wherein a first set of parameters parameterizes the dialogue strategy with respect to the user and characterizes the effectiveness of the dialogue with the user under the dialogue strategy.
[0015] Additional advantages and novel features will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the following and the accompanying drawings, or may be learned by the production or operation of examples. The advantages of the present teachings may be realized and obtained by practicing or using various aspects of the methods, tools, and combinations set forth in the detailed examples discussed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The methods, systems, and / or programs described herein are further described according to exemplary embodiments. These exemplary embodiments are described in detail with reference to the accompanying drawings. These embodiments are non-limiting exemplary embodiments, where like reference numerals represent similar structures in several views of the drawings, and where:
[0017] Figure 1A An exemplary configuration of a dialogue system centered on an information state that captures dynamic information observed during a dialogue according to an embodiment of the present teachings is depicted;
[0018] Figure 1B is a flowchart of an exemplary process of a dialogue system that uses an information state capturing dynamic information observed during a dialogue according to an embodiment of the present teachings;
[0019] Figure 2A An exemplary construction of an information state according to an embodiment of the present teachings is depicted;
[0020] Figure 2B Illustrates a representation of how different estimated mindsets are connected in a dialogue with a robot tutor teaching users addition of fractions according to an embodiment of the present teachings;
[0021] Figure 2C Shows an exemplary relationship between an estimated representation of the mindset of an agent, a shared mindset, and the mindset of a user in an information state according to an embodiment of the present teachings;
[0022] Figure 3A Shows an exemplary relationship between different types of And-Or-Graphs (AOGs) for representing the estimated mindsets of the parties involved in a dialogue according to an embodiment of the present teachings;
[0023] Figure 3B Depicts an exemplary association between a Spatial AOG (S-AOG) and a Temporal AOG (T-AOG) in an information state according to an embodiment of the present teachings;
[0024] Figure 3C Illustrates an exemplary S-AOG and its associated T-AOG according to an embodiment of the present teachings;
[0025] Figure 3D Illustrates an exemplary relationship between S-AOG, T-AOG, and C-AOG according to an embodiment of the present teachings;
[0026] Figure 4A Illustrates an exemplary S-AOG according to an embodiment of the present teachings, which partially represents the mindset of an agent for teaching different mathematical concepts;
[0027] Figure 4B Illustrates an exemplary T-AOG according to an embodiment of the present teachings, which represents a dialogue strategy partially associated with the mindset of an agent for teaching the concept of fractions;
[0028] Figure 4C Shows an exemplary dialogue content for teaching concepts associated with fractions according to an embodiment of the present teachings;
[0029] Figure 5A Illustrates an exemplary temporal parsed graph (T-PG) within a T-AOG according to an embodiment of the present teachings, which represents the shared mindset between a user and a machine;
[0030] Figure 5B Illustrates a part of a dialogue between a machine and a person along a dialogue path according to an embodiment of the present teachings, where the dialogue path represents the current representation of the shared mindset;
[0031] Figure 5C Depicts an exemplary S-AOG according to an embodiment of the present teachings, which has nodes parameterized with measurements related to the mastery levels of different underlying concepts to represent the mindset of a user;
[0032] Figure 5D Shows exemplary types of user personality traits according to an embodiment of the present teachings, which can be estimated based on observations from a dialogue;
[0033] Figure 6A Depicts a general S-AOG for a tutoring dialogue according to an embodiment of the present teachings;
[0034] Figure 6B Depicts a specific T-AOG for a dialogue regarding greetings according to an embodiment of the present teachings;
[0035] Figure 6C Shows different types of parameterization alternatives for different types of AOGs according to an embodiment of the present teachings;
[0036] Figure 6DIllustrated is an S-AOG according to an embodiment of the present teachings, which has different nodes parameterized by rewards that are updated based on dynamic observations from a conversation;
[0037] Figure 6E Illustrated is an exemplary T-AOG generated according to an embodiment of the present teachings by merging different graphs via graph matching using parameterized content;
[0038] Figure 6F Illustrated is an exemplary T-AOG according to an embodiment of the present teachings, which has parameterized content associated with nodes;
[0039] Figure 6G Shown is a T-AOG according to an embodiment of the present teachings, which has each node parameterized by one or more content sets;
[0040] Figure 6H Illustrated is exemplary data in different content sets associated with different nodes of a T-AOG according to an embodiment of the present teachings;
[0041] Figure 6I Illustrated is an exemplary T-AOG according to an embodiment of the present teachings, in which different paths traverse different nodes, and these different paths are parameterized by rewards that are updated based on dynamic observations from a conversation;
[0042] Figure 7A Depicted is a high-level system diagram of a knowledge tracking unit according to an embodiment of the present teachings;
[0043] Figure 7B Illustrated is how knowledge tracking enables adaptive dialogue management according to an embodiment of the present teachings;
[0044] Figure 7C Is a flowchart of an exemplary process of a knowledge tracking unit according to an embodiment of the present teachings;
[0045] Figure 8A Shown is an example of a utility-driven tutoring (node) plan for an S-AOG according to an embodiment of the present teachings;
[0046] Figure 8B Illustrated is an example of a utility-driven path plan for a T-AOG according to an embodiment of the present teachings;
[0047] Figure 8C Illustrated is a dynamic state in utility-driven adaptive dialogue management based on a parameterized AOG according to an embodiment of the present teachings;
[0048] Figure 9ADepicts an exemplary mode of creating an AOG with authored content according to an embodiment of the present teachings;
[0049] Figure 9B Depicts an exemplary high-level system diagram of a content creation system for automatically creating an AOG via machine learning according to an embodiment of the present teachings;
[0050] Figure 9C Shows different types of topic-based AOGs derived from machine learning according to an embodiment of the present teachings;
[0051] Figure 9D Is a flowchart of an exemplary process of a content creation system for creating an AOG via machine learning according to an embodiment of the present teachings;
[0052] Figure 10A Illustrates an exemplary visual programming interface configured for content creation associated with an AOG according to an embodiment of the present teachings;
[0053] Figure 10B Illustrates an exemplary visual programming interface configured for authoring content for a parameterized AOG according to an embodiment of the present teachings;
[0054] Figure 10C Is a flowchart of an exemplary process for creating an AOG and content creation via visual programming according to an embodiment of the present teachings;
[0055] Figure 11A Illustrates exemplary code for generating scenes associated with an S-AOG obtained via automatic / semi-automatic content creation according to an embodiment of the present teachings;
[0056] Figure 11B Illustrates exemplary code for generating a T-AOG obtained via automatic / semi-automatic content creation according to an embodiment of the present teachings;
[0057] Figure 12A Depicts an exemplary high-level system diagram of a system for authoring content based on multimodal input from a user according to an embodiment of the present teachings;
[0058] Figure 12B Illustrates different types of metadata that can be automatically generated based on multimodal input from a user according to an embodiment of the present teachings;
[0059] Figure 12C Is a flowchart of an exemplary process of a system configured for authoring content based on multimodal input from a user according to an embodiment of the present teachings;
[0060] Figure 13is an illustrative diagram of an exemplary mobile device architecture that can be used to implement a dedicated system for practicing the present teachings; and
[0061] Figure 14 is an illustrative diagram of an exemplary computing device architecture that can be used to implement a dedicated system for practicing the present teachings. DETAILED DESCRIPTION
[0062] In the following detailed description, numerous specific details are set forth by way of example in order to facilitate a thorough understanding of the relevant teachings. However, it will be apparent to those skilled in the art that the present teachings may be practiced without these details. In other instances, well-known methods, procedures, components, and / or circuitry have been described at a relatively high level without detailed elaboration in order to avoid unnecessarily obscuring aspects of the present teachings.
[0063] The present teachings are directed to addressing the deficiencies of traditional human-machine dialogue systems and providing methods and systems that enable a rich representation of multimodal information from a conversation environment to allow a machine to have an improved sense of the context of the conversation in order to better adapt the dialogue with more effective conversation and enhanced interaction with the user. Based on such a representation, the present teachings further disclose different modes of creating such a representation and authoring the content of the dialogue in processing the representation. In addition, to allow for adaptation of the representation based on the dynamics occurring during the conversation, the present teachings also disclose mechanisms for tracking the dynamics of the conversation and updating the representation accordingly, and then the machine uses these representations to conduct the dialogue in a utility-driven adaptive manner to achieve maximized results.
[0064] Figure 1A Depicts an exemplary configuration of a dialogue system 100 according to an embodiment of the present teachings, the dialogue system 100 centered around an information state 110 that captures the dynamic information observed during the conversation. The dialogue system 100 includes a multimodal information processor 120, an automatic speech recognition (ASR) engine 130, a natural language understanding (NLU) engine 140, a dialogue manager (DM) 150, a natural language generation (NLG) engine 160, and a text-to-speech (TTS) engine 170. The dialogue system 100 interacts with a user 180 to conduct a conversation.
[0065] During a conversation, multimodal information is collected from the environment (including from user 180), which captures surrounding information of the conversation environment, the user's voice, and (facial or body) expressions, etc. The multimodal information thus collected is analyzed by the multimodal information processor 120 to extract relevant features of different modalities in order to estimate different characteristics of the user and the environment. For example, the voice signal can be analyzed to determine voice-related features such as speaking speed, pitch, or even accent. Visual signals related to the user can also be analyzed to extract, for example, facial features or body postures, etc., in order to determine the user's expression. By combining acoustic features and visual features, the multimodal information processor 120 can also be able to infer the user's emotional state. For example, a high pitch, rapid speech plus an angry facial expression can indicate that the user is unhappy. In some embodiments, the observed user activities can also be analyzed to better understand the user. For example, if the user points to or walks towards a specific object, then it can reveal what the user is referring to in his / her speech. Such multimodal information can provide useful context to understand the user's intention. The multimodal information processor 120 can continuously analyze the multimodal information and store such analyzed information in the information state 110, and then different components in the dialogue system 100 use the analyzed information to facilitate decisions related to dialogue management.
[0066] In operation, the voice information from user 180 is sent to the ASR engine 130 to perform speech recognition. The speech recognition can include discerning the language and the words spoken by user 180. To understand the semantics of what the user said, the result from the ASR engine 130 is further processed by the NLU engine 140. This understanding can depend not only on the words spoken, but also on other information (such as the user 180's expression and posture) and / or other context information (such as what was said previously). Based on the understanding of the user's utterance, the dialogue manager 150 determines how to respond to the user. Then the determined response can be generated in text form via the NLG engine 160 and further transformed from text form into a voice signal via the TTS engine 170. Then the output of the TTS engine 170 can be delivered to the user 180 as a response to the user's utterance. Through such back-and-forth interaction, the process for the machine dialogue system continues to conduct a conversation with the user 180.
[0067] As Figure 1AAs seen, the components in the dialogue system 100 are connected to the information state 110, which, as discussed herein, captures the dynamics surrounding the dialogue and provides relevant and rich context information that can be used to facilitate speech recognition (ASR), natural language understanding (NLU), and various dialogue-related determinations, including what is an appropriate response (DM), what language features to apply to the text response (NLG), and how to convert the text response into a speech form (TTS) (e.g., what accent). As discussed herein, the information state 110 can represent dialogue-related dynamics obtained based on multimodal information, which is related to the user 180 or to the surroundings of the dialogue.
[0068] After receiving multimodal information from the dialogue scenario (regarding the user or regarding the dialogue environment), the multimodal information processor 120 analyzes the information and characterizes the dialogue surrounding environment in different dimensions, e.g., acoustic characteristics (e.g., the user's pitch, speed, accent), visual characteristics (e.g., the user's facial expression, objects in the environment), physical characteristics (e.g., the user waving or pointing at an object in the environment), the estimated user's mood and / or mindset, and / or the user's preferences or intentions. This information can then be stored in the information state 110.
[0069] The rich media context information stored in the information state 110 can facilitate different components to play their respective roles, enabling the conversation to proceed in an adaptive, more engaging, and more effective manner regarding the intended goal. For example, the rich context information can improve the understanding of the user 180's utterance based on what is observed in the conversation scenario, evaluate the performance of the user 180 and / or estimate the user's utility (or preference) according to the intended goal of the conversation, determine how to respond to the user 180's utterance based on the estimated emotional state of the user, and deliver the response in the most appropriate manner considered based on the understanding of the user, etc. For example, if accent information represented in the form of sound (e.g., a special way of uttering certain phonemes) and visual form (e.g., the user's special visual elements) is captured in the information state regarding the user, then the ASR engine 130 can utilize this information to determine the words spoken by the user. Similarly, the NLU engine 140 can also utilize the rich context information to determine the semantics referred to by the user. For example, if the user points to a computer placed on the table (visual information) and says "I like this", then the NLU engine 140 can combine the output of the ASR engine 130 (i.e., "I like this") and the visual information that the user points to the computer in the room to understand that the user's "this" refers to the computer. As another example, if the user 180 repeatedly makes mistakes during a tutoring session, and at the same time, it is estimated based on the tone of voice and / or the user's facial expression (e.g., they are determined based on multimodal information) that the user is annoyed, then instead of continuing to push forward the tutoring content, the DM 140 can decide to temporarily change the topic based on the user's known interests (e.g., talking about Lego games) to continue to engage the user. The decision to distract the user can be determined based on, for example, the utility of what has worked (e.g., intermittently distracting the user has worked in the past) and what has not worked (e.g., continuing to pressure the user to do better) previously observed regarding the user.
[0070] Figure 1B is a flowchart of an exemplary process of the dialogue system 100 according to an embodiment of the present teachings, where the information state 110 captures dynamic information observed during the conversation. As Figure 1BAs can be seen, the process is an iterative process. At 105, multimodal information is received, and then this multimodal information is analyzed by the multi-information processor 170 at 125. As discussed herein, multimodal information includes information related to the user 180 and / or information related to the context of the conversation. Multimodal information related to the user may include the user's utterances and / or visual observations of the user, such as body postures and / or facial expressions. Information related to the context of the conversation may include information related to the environment, such as the objects present, the spatial / temporal relationships between the user and such observed objects (e.g., the user stands in front of a table), and / or the dynamic relationships between the user's activities and the observed objects (e.g., the user walks towards the table and points to the computer on the table). Then, the understanding of the multimodal information captured from the conversation scenario can be used to facilitate other tasks in the dialogue system 100.
[0071] Based on the information stored in the information state 110 (representing the past state) and the analysis results from the multimodal information processor 120 (regarding the current state), the ASR engine 130 and the NLU engine 140 perform speech recognition at 125 respectively to determine the words spoken by the user and the language understanding based on the recognized words. ASR and NLU can be performed based on the current information state 110 and the analysis results from the multimodal information processor 120.
[0072] Based on the results of multimodal information analysis and language understanding (i.e., what the user said or what it means), the change in the dialogue state is tracked at 135, and such change is used to update the information state 110 accordingly at 145 to facilitate subsequent processing. To execute the dialogue, the DM 140 determines a response at 155 based on the dialogue tree designed for the underlying dialogue, the output of the NLU engine 140 (the understanding of the utterance), and the information stored in the information state 110. Once the response is determined, the NLG engine 160 generates a response based on the information state 110, for example, in its text form. When the response is determined, there may be different ways to say it. The NLG engine 160 can generate a response at 165 in a style based on the user's preference or a style known to be more suitable for the specific user in the current dialogue. For example, if the user answers a question incorrectly, there may be different ways to point out that the answer is incorrect. For a specific user in the current dialogue, if it is known that the user is sensitive and easily frustrated, a milder way can be used to tell the user that his / her answer is incorrect to generate a response. For example, instead of saying "This is wrong", the NLG engine 160 can generate a text response of "It's not entirely correct".
[0073] The text response generated by the NLG engine 160 can then be rendered into an audible form, e.g., an audio signal form, by the TTS engine 170 at 175. Although standard or common TTS techniques can be used to perform TTS, the present teachings disclose that the response generated by the NLG engine 160 can be further personalized based on the information stored in the information state 110. For example, if a slower talking speed or a softer talking style is known to be more effective for the user, then the generated response can be rendered by the TTS engine 170 at 175 into an audible form, e.g., having a lower speed and pitch. Another example is to render the response with an accent consistent with the known accent of the student based on the personalized information of the user in the information state 110. The rendered response can then be delivered to the user as a response to the user's utterance at 185. After responding to the user, the dialogue system 100 then tracks additional changes in the dialogue and updates the information state 110 accordingly at 195.
[0074] Figure 2A Depicts an exemplary construction represented by the information state 110 according to an embodiment of the present teachings. Non - limitingly, the information state 110 includes a representation for an estimated mindset. As shown, different representations can be estimated to represent, for example, the mindset 200 of the agent, the mindset 220 of the user, and the shared mindset 210, along with other information recorded therein. The mindset 200 of the agent can refer to the (one or more) intended goals to be achieved by the dialogue agent (machine) in a particular dialogue. The shared mindset 210 can refer to a representation of the current dialogue situation, which is a combination of the agent's execution of the intended agenda based on the agent's mindset 200 and the actual performance of the user. The mindset 220 of the user can refer to a representation of the agent's estimate of the state of the student with respect to the intended purpose of the dialogue based on the shared mindset or the user's performance. For example, if the current task of the agent is to teach a student user the concept of fractions in mathematics (which may include sub - concepts for building an understanding of fractions), then the mindset of the user can include an estimated level of mastery of the user over various related concepts. Such an estimate can be derived based on an assessment of the student's performance at different stages of tutoring such related concepts.
[0075] Figure 2BIllustrated is how such a representation of different mindsets is connected in an example where a robotic tutor 205 teaches a concept 215 related to fraction addition to a student user 180 according to an embodiment of the present teachings. As can be seen, the robotic tutor 205 interacts with the student user 180 via multimodal interaction. The robotic tutor 205 can start tutoring based on an initial representation of the agent's mindset 200 (e.g., a lesson on fraction addition that can be represented as an AOG). During the tutoring, the student user 180 can answer questions from the robotic tutor 205 and such answers to the questions form a specific dialogue path, enabling an estimation of the representation of the shared mindset 210. Based on the user's answers, the user's performance is evaluated, and a representation of the user's mindset 220 is estimated with respect to different aspects, e.g., whether the student has mastered the taught concept and what kind of dialogue style works for this particular student.
[0076] As Figure 2A seen, the representation of the estimated mindset is based on some graph-related forms (including but not limited to the spatio-temporal-causal AND-OR graph STC-AOG 230, the STC parse graph (STC-PG) 240), and can be used in combination with other types of information stored in the information state, such as the dialogue history 250, the dialogue context 260, event-centered knowledge 270, the common sense model 280, … and the user profile 290. These different types of information can belong to multiple modalities and constitute different aspects of the dynamics of each dialogue for each user. Thus, the information state 110 captures the general information of various dialogues as well as the personalized information about each user and each dialogue, and interconnects them together to facilitate different components in the dialogue system 100 to perform corresponding tasks in a more adaptive, personalized, and engaging manner.
[0077] Figure 2C Illustrated is an exemplary relationship between the agent's mindset 200, the shared mindset 210, and the user's mindset 220 represented in the information state 110 according to an embodiment of the present teachings. As discussed herein, the shared mindset 210 represents the state of a dialogue achieved via the interaction between the agent and the user, and is a combination of the agent's intended intention (according to the agent's mindset) and the user's performance in following the agent's intended agenda. Based on the shared mindset 210, the dynamics of the dialogue can be traced with respect to what the agent can achieve and what the user can achieve up to that point.
[0078] Tracking this dynamic knowledge enables the system to estimate what the user has achieved up to that point and in what way the student user has grasped which concepts or sub - concepts (i.e., which dialogue paths have worked and which may not work). Based on what the student user has achieved so far, the user's mindset 220 can be inferred or estimated, which will be used to determine how the agent can further facilitate adjusting or updating the dialogue strategy to achieve the desired goal or adjust the agent's mindset to suit the user. The process of adjusting the agent's mindset enables the derivation of the updated agent's mindset 200. Based on the dialogue history, the dialogue system 100 learns the user's preferences or what is more effective (utility) for the user. Once this information is incorporated into the information state, it will be used to adjust the dialogue strategy via a utility - driven (or preference - driven) dialogue plan. The updated dialogue strategy drives the next step in the dialogue, which in turn leads to a response from the user and subsequent updates to the shared mindset, the user's mindset, and the agent's mindset. This process iterates so that the agent can continue to adjust the dialogue strategy based on the dynamic information state.
[0079] According to this teaching, different mindsets are represented based on, for example, STC - AOG and STC - PG. Figure 3A Exemplary relationships between different types of AND - OR graphs (AOGs) for representing the estimated mindsets of the parties involved in a dialogue according to an embodiment of this teaching are shown. An AOG is a graph with AND (conjunction) branches and OR (disjunction) branches. Branches associated with a node in an AOG and connected by an AND relationship represent tasks that all need to be traversed. Branches from a node in an AOG and connected by an OR relationship represent tasks that can be selectively traversed. As discussed herein, STC - AOG includes an S - AOG corresponding to a spatial AOG, a T - AOG corresponding to a temporal AOG, and a C - AOG corresponding to a causal AOG. According to this teaching, an S - AOG is a graph that includes nodes, each of which can correspond to a topic to be covered in a dialogue. A T - AOG is a graph that includes nodes, each of which can correspond to a temporal action to be taken. Each T - AOG can be associated with a topic or node in the S - AOG, i.e., represents steps to be performed during a dialogue about the topic corresponding to the S - AOG node. A C - AOG is a graph that includes nodes, each of which can be linked to a node in the T - AOG and the corresponding node in the S - AOG, thus representing the action that occurs in the T - AOG and the causal impact of that action on the corresponding node in the S - AOG.
[0080] Figure 3BDepicts an exemplary relationship between nodes in the S-AOG according to an embodiment of the present teachings and nodes of the associated T-AOG represented in information state 110. In this illustration, each K-node corresponds to a node in the S-AOG and represents a skill or topic to be taught in the conversation. The assessment regarding each K-node can include "mastered" or "not yet mastered", for example, the corresponding probabilities P(T) and 1 - P(T), i.e., P(T) represents the transition probability from the not yet mastered to the mastered state. P(L0) represents the probability of prior learning skills or prior knowledge regarding the topic, i.e., the likelihood that the student has mastered the concept before the tutoring session begins. To teach the skill / concept associated with each K-node, the robot tutor can pose multiple questions according to the T-AOG associated with the K-node, and then the student will answer each question. Each question is shown as a Q-node and the student's answer is represented as an A-node in Figure 3B as seen in
[0081] During the conversation between the agent and the user, the student's answer can be a correct answer A(c) or an incorrect answer A(w), as seen in Figure 3B Based on each answer received from the user, additional probabilities are determined based on various knowledge or observations collected, for example, during the conversation. For example, if the user provides a correct answer (A(c)), then the probability P(G) that the answer is a guess can be determined, which represents the likelihood that the student did not know the correct answer but guessed it correctly. Conversely, 1 - P(G) is the probability that the user knew the correct answer and answered it correctly. For an incorrect or wrong answer A(w), the probability P(S) can be determined, which represents the likelihood that the student gave an incorrect answer but actually knew the concept. Based on P(S), the probability 1 - P(S) can be estimated, which represents the likelihood that the student gave an incorrect answer because they did not know the concept. Such probabilities can be calculated for each node along the path experienced based on the actual conversation and can be used to estimate when the student has mastered the concept and to estimate what might work and what might not work in teaching this particular student regarding each specific topic.
[0082] Figure 3CIllustrated is an exemplary S-AOG and its associated T-AOG according to an embodiment of the present teachings. In the exemplary S-AOG 310 for guiding the concept of fractions, each node corresponds to a topic or concept to be taught to a student user during a conversation. For example, S-AOG 310 includes a node P0 or 310-1 representing the concept of fractions, a node P1 or 310-2 representing the concept of division, a node P2 or 310-3 representing the concept of multiplication, a node P3 or 310-4 representing the concept of addition, and a node P4 or 310-5 representing the concept of subtraction. In this example, the different nodes in S-AOG 310 are related. For example, to master the concept of fractions, at least some of the other concepts of addition, subtraction, multiplication, and division need to be mastered first. To teach the concepts (e.g., fractions) represented by the S-AOG nodes, the agent may need to perform a series of steps or processes during a conversation session with the student user. Such a process or series of steps corresponds to the T-AOG. In some embodiments, for each node in the S-AOG, there may be multiple T-AOGs, and each T-AOG may represent a different way of teaching the student and may be invoked in a personalized manner. As shown, the S-AOG node 310-1 has multiple T-AOGs 320, one of which is shown as 320-1, which corresponds to a series of time steps such as question / answer 330, 340, 350, 360, 370, 380... etc. In each tutoring session for teaching the concept of fractions, the choice of which T-AOG to use can vary and can be determined based on various considerations (e.g., the user in the session (personalized), the degree of mastery of the current concept (e.g., P(L0)), etc.).
[0083] The representation capture of a dialogue based on STC-AOG captures entities / objects / concepts related to the dialogue (S-AOG), possible actions observed during the dialogue (T-AOG), and the impact of each of these actions on these entities / objects / concepts (C-AOG). The actual dialogue activities that occur during the dialogue (voice) cause traversal of the corresponding graphical representation or STC-AOG, resulting in a parse graph (PG) corresponding to the traversed portion of the STC-AOG. In some embodiments, the S-AOG may model the spatial decomposition of the objects and scenarios of the dialogue. In some embodiments, the S-AOG may model the decomposition of concepts and sub-concepts as discussed herein. In some embodiments, the T-AOG may model the temporal decomposition of events / sub-events / actions that can be performed or have occurred in a dialogue related to certain entities / objects / concepts represented in the corresponding S-AOG. The C-AOG may model the decomposition of the events represented in the T-AOG and their causal relationships with the corresponding entities / objects / concepts represented in the S-AOG. That is, the C-AOG describes the changes to the nodes in the S-AOG caused by the events / actions taken in the dialogue and represented in the T-AOG. This information is about different aspects of the dialogue and is captured in the information state 110. That is, the information state 110 represents the dynamics of the dialogue between the user and the dialogue agent. This is shown in Figure 3D illustrated.
[0084] As discussed herein, based on the actual dialogue session, the specific path traversed based on the conversation results in different types of corresponding parse graphs (PGs). For example, it can be applied to the S-AOG to produce an S-PG, to the T-AOG to produce a T-PG, and to the C-AOG to produce a C-PG. That is, based on the STG-AOG, the actual dialogue results in a dynamic STC-PG that at least partially represents the different mental states of the parties participating in the dialogue session. To illustrate this, Figure 4A - 4C an exemplary S-AOG / T-AOG is shown, which is associated with the agent's mind for teaching fraction-related concepts; Figure 5A - 5B an exemplary representation of the shared mental state is provided via the T-PG, which is generated based on the dialogue in a specific tutoring session; Figure 6A - 6B an exemplary representation of the user's mental state in terms of the estimated mastery of different concepts taught in the dialogue with the dialogue agent is shown.
[0085] Figure 4AShows an exemplary representation of an agent's mindset regarding fraction tutoring according to an embodiment of the present teachings. As discussed herein, the representation of the agent's mindset can reflect what the agent expects or is designed to cover in a conversation. The agent's mindset can be adjusted based on the user's performance / behavior during a conversation session such that the representation of the agent's mindset captures such dynamics or adjustments. As Figure 4A shown, the exemplary representation of the agent's mindset includes various nodes, each representing a sub-concept related to the concept of fractions. For example, there are sub-concepts related to: "Understanding Fractions" 400, "Comparing Fractions" 405, "Understanding Equivalent Fractions" 410, "Expanding and Simplifying Equivalent Fractions" 415, "Finding Factor Pairs" 420, "Applying the Properties of Multiplication / Division" 425, "Adding Fractions" 430, "Finding the LCM" 435, "Solving for Unknowns in Multiplication / Division" 440, "Multiplication and Division within 100" 445, "Simplifying Improper Fractions" 450, "Understanding Improper Fractions" 455, and "Addition and Subtraction" 460. These sub-concepts can form the landscape of fractions, and some sub-concepts may need to be taught before other sub-concepts. For example, "Understanding Improper Fractions" 455 may need to be covered before "Simplifying Improper Fractions" 450, and "Addition and Subtraction" 460 may need to be mastered before "Multiplication and Division within 100" 445, and so on.
[0086] Figure 4B Illustrates an exemplary T-AOG according to an embodiment of the present teachings, which represents the agent's mindset when teaching concepts related to fractions. As discussed herein, the T-AOG includes various steps associated with a conversation, some of which relate to what the agent says, some of which relate to what the user responds, and some of which correspond to certain evaluations of the conversation performed by the agent. There are branches in the T-AOG that represent decisions. For example, at 470, the corresponding action is for the agent to highlight the numerator and denominator boxes, which could be, for example, after teaching a student what a numerator and denominator are. After link 480, the agent advances to 490 to request user input, for example, asking the user to tell the agent which of the highlighted ones is the denominator. Based on the answer received from the student, the agent follows two links combined by an OR (plus sign), where each link represents a path taken by the user. For example, if the user correctly answers which one is the denominator, then the agent advances to 490-3, for example, further asking the user to evaluate the denominator. If the user answers incorrectly, then the agent advances to 490-4 to provide a hint about the denominator to the user and then returns to 490 along link 490-2, again asking for user input about which one is the denominator.
[0087] If the agent asks the user to evaluate the denominator at 490-3, then there are two associated outcomes, one being the wrong answer and the other being the correct answer. The former leads to 490-5, where the agent indicates to the user that the answer is incorrect and then follows the link 490-1 back to 490, asking the user for the input again. If the answer is correct, then the agent follows another path forward to 490-6, letting the user know that he / she is correct and continuing along that path to further set the denominator and clear the highlighting at 490-7 and 490-8 respectively. As can be seen, the steps at 490 represent the time actions planned by the agent, related to the concept of teaching the denominator, which are related to the Figure 4A S-AOG in which the concept representing the agent's plan to teach the student the concept of fractions is shown in. Thus, they together form part of the representation of the agent's mindset. Figure 4C Illustrates exemplary dialogue content created for teaching concepts associated with fractions according to an embodiment of the present teachings. By using a similar dialogue strategy, the conversation is intended to be executed in a question-and-answer flow.
[0088] Figure 5A Illustrates an exemplary representation of the shared mindset in the form of a T-PG according to an embodiment of the present teachings (corresponding to the path in the T-AOG in Figure 4B . The highlighted steps form the specific path taken by the actions performed by the dialogue agent in the conversation based on the answers from the user. Compared with the T-AOG shown in Figure 4B , Figure 5A shows the T-PG of the individual highlighted steps (e.g., 470, 510, 520, 530, 540, 550...) along the highlighted path. Figure 5A The T-PG shown in represents the instantiated path traversed based on the actions of both the agent and the user, and thus represents the shared mindset. Figure 5B Illustrates a part of the created dialogue content between the agent and the user according to an embodiment of the present teachings, from which the representation of the shared mindset can be obtained. As discussed herein, the representation of the shared mindset can be derived based on the flow of the conversation, which forms a specific path or T-PG traversed along the existing T-AOG.
[0089] As discussed herein, during the conversation, the dialogue agent estimates the mindset of the user participating in the conversation based on the observation of the conversation with the user, such that both the conversation and the representation of the estimated user's mindset are adjusted based on the dynamics of the conversation. For example, in order to determine how to proceed with the conversation, the agent may need to evaluate or estimate the user's mastery of a particular topic based on the observation of the user. As regarding Figure 3BAs discussed, the estimate can be probabilistic. Based on such probabilities, the agent can infer the current level of mastery of the concept and determine how to proceed further in the conversation. For example, if the estimated level of mastery is insufficient, then continue to instruct on the current topic, or if the estimated user's mastery of the current concept is sufficient, then move on to other concepts. The agent can evaluate periodically during the conversation and annotate the PG (parameterized) during this process to facilitate the decision-making of the next move when traversing the graph. Such an annotated or parameterized S-AOG can produce an S-PG, that is, for example, indicating which nodes in the S-AOG have been sufficiently covered and which have not been sufficiently covered. Figure 5C depicts an exemplary S-PG of a corresponding S-AOG according to an embodiment of the present teachings, which represents the estimated state of mind of the user. The underlying S-AOG is shown in Figure 4A In this illustrated example, during the conversation, each node in this S-AOG is evaluated based on the conversation and parameterized or annotated based on such evaluation. As shown in Figure 5C nodes representing different sub-concepts related to the fraction are annotated with corresponding different parameters (these parameters indicate, for example, the level of mastery of the corresponding node).
[0090] As shown in Figure 5C nodes in the initial S-AOG ( Figure 4A ) are now annotated in Figure 5C with different weights, each weight indicating the evaluated level of mastery of the sub-concept corresponding to that node. As can be seen, the nodes in Figure 5C are presented in different shades, which are determined according to the weights representing different levels of mastery of the underlying sub-concepts. For example, the now-dotted nodes can correspond to those sub-concepts that have been mastered and thus do not require further traversal. Nodes 560 and 565 (corresponding to "understanding fractions" and "understanding improper fractions") can correspond to sub-concepts that have not reached the required level of mastery. All nodes connecting to these two nodes between these two nodes (for example, mastered and unmastered) can be considered the reasons why the user has not mastered the concepts of fractions and improper fractions.
[0091] Such an estimated level of mastery of the corresponding nodes in the original S-AOG results in an annotated S-PG that represents the state of mind of the estimated user, which will indicate the degree of understanding of the concepts associated with such nodes. This provides a basis for the dialogue agent to understand the relevant schema of the user, e.g., what the user has understood of what has been taught and what the user still has questions about. As can be seen, the representation of the user's state of mind is dynamically estimated based on, for example, the user's performance and activities during the ongoing conversation. In addition to estimating the level of mastery of the concepts associated with different nodes to understand the user's state of mind, context observations and information about the user that can be collected during the ongoing conversation can also be used to estimate other characteristics or behavioral metrics of the user as part of understanding the user's state of mind. Figure 5D Exemplary types of user personality traits that can be estimated based on information observed during a conversation in accordance with an embodiment of the present teachings are shown. As depicted, during a conversation with a user, based on observations of the user's behavior or expressions (whether verbal, visual, or physical), the agent can estimate various characteristics of the user in different dimensions, such as whether the user is extroverted, how mature the user is, whether the user is naughty, whether the user is easily excitable, whether the user is generally cheerful, how confident or secure the user feels about him / herself, whether the user is reliable, meticulous, etc., via multimodal information processing (e.g., by multimodal information processor 120). Once such information is estimated, it forms a profile of the user, which can influence the dialogue system 100 to determine how to adjust the dialogue strategy of the dialogue system 100 when needed and in what manner the agent of the dialogue system 100 should converse with the user.
[0092] Both the S-AOG and the T-AOG can have certain structures that are organized based on, for example, the topic, concept, or flow of the conversation. Figure 6A An exemplary general structure of an S-AOG related to a tutoring conversation in accordance with an embodiment of the present teachings is depicted. As Figure 6AThis general structure shown is not specific to a topic, but can be used to teach any topic. Exemplary structures include different stages involved in a tutoring conversation, which are represented as different nodes in the S-AOG. As shown, a greeting node 600 is used for conversations related to greetings, a chat node 605 is used for chats about, for example, weather or health, a review node 610 is used for conversations to review previously learned knowledge (e.g., as a basis for teaching the intended topic), a teaching node 615 is used to teach the intended topic, a test node 620 is used to test the student user on the taught topic, and an evaluation node 625 is used for conversations to evaluate the student user's mastery of the taught topic based on the test. Different nodes can be connected in ways that cover different flows between underlying sub-conversations, but the specific flow within each conversation can be determined dynamically based on the situation. Some branches coming out of a node can be related via an AND relationship, and some branches coming out of a node can be related via an OR relationship.
[0093] As Figure 6A seen, a conversation related to tutoring can start with a greeting conversation 600, such as "Good morning", "Good afternoon", or "Good evening". There are three branches coming out of the greeting node 600, including going to the chat node 605 for a short chat, going to the review node 610 for reviewing previously learned knowledge, and going to the teaching node 615 to start teaching directly. These three branches are ORed together, i.e., the conversation agent can proceed to follow any one of these three branches. After the chat session 605, there are also three branches, one going to the teaching node 615, one going to the test node 620, and one going to the review node 610. The review node 610 also has two branches, one going to the teaching node 615 and the other going to the test node 620 (the prior knowledge or prior mastery level of the topic of the student can be tested first before teaching). In this illustrated embodiment, the teaching and test nodes are required conversations, such that the branches from the chat node 605 and the review node 610 to the teaching node 615 and the test node 620 are related via AND.
[0094] Teaching and testing can be iterative, as indicated by the two-way arrows between teaching node 615 and testing node 620. As needed, either teaching node 615 or testing node 620 can proceed to evaluation node 625. That is, the evaluation can be performed based on the teaching results from teaching node 615 or the test results from testing node 620. Based on the evaluation results, the conversation can proceed to one of three alternative options (related by OR), including teaching 615 (reviewing the concepts again), testing 620 (retesting), or review 610 (strengthening the user's understanding of some concepts), or even chatting 605 (e.g., if it is found that the user is frustrated, then the dialogue system 100 can switch topics to continue engaging the user rather than losing the user). This general S-AOG for tutoring-related conversations is provided by way of illustration and not limitation. The S-AOG for tutoring can be derived according to any logical flow required by the application.
[0095] As Figure 6A seen, each node is itself a conversation and, as discussed herein, can be associated with one or more T-AOGs, each T-AOG representing a conversation flow for an intended topic. Figure 6B Depicted is an exemplary T-AOG having conversation content regarding greetings authored for S-AOG greeting node 600 according to an embodiment of the present teachings. The T-AOG can be defined as the conversation strategy of the conversation. Following the steps defined in the T-AOG is to execute a strategy for achieving certain intended purposes. In Figure 6B it, the content in each rectangular box represents what the agent is to say, and the content in the ellipse represents what the user responds with. As seen, in Figure 6B the T-AOG for greetings shown, the agent first says one of three alternative greetings, namely, good morning 630-1, good afternoon 630-2, and good evening 630-3. The user's response to such a greeting can vary. For example, the user can repeat what the agent said (i.e., good morning, good afternoon, or good evening). Some people will repeat and then add "you too" 635-1. Some people will say "thank you, and you?" at 635-2. Some people will say both 635-1 and 635-2. Some people can just remain silent 635-3. There can be other alternative ways to respond to the agent's greeting. After receiving a response from the user, the dialogue agent can then answer the user's response. For each alternative response from the user, each answer can correspond to the user's response. This is illustrated by the content at 640-1, 640-2, and 640-3 in Figure 6B it.
[0096] Figure 6B The T-AOG shown in it can cover multiple T-AOGs. For example, Figure 6B630-1, 635-2, and 640-2 in Figure 6B can form a T-AOG for greeting. Similarly, 630-1, 635-1, 640-1 can correspond to another T-AOG for greeting; 630-2, 635-1, 640-1 can form another T-AOG; 630-1, 635-3, 640-3 form a different T-AOG; 630-2, 635-3, and 640-3 can form another different T-AOG, and so on. Although different, these alternative T-AOGs all have a substantially similar structure and common content. This commonality can be used to generate a simplified T-AOG that has flexible content associated with each node. This can be achieved via, for example, graph matching. For example, the different T-AOGs related to greeting mentioned above, although having different authored content regarding greeting, all have a similar structure, that is, an initial greeting plus a response from the user and plus a response to the user's response to the greeting. In this sense, the T-AOG in
[0097] may not correspond to the most simplified general T-AOG for greeting.
[0097] To facilitate flexible conversation content and enable the dialogue system 100 to adjust the dialogue in a personalized manner, the AOG can be parameterized. According to different embodiments of the present teachings, such parameterization can be applied to both the S-AOG and the T-AOG with respect to both the parameters associated with the nodes in the AOG and the parameters associated with the links between different nodes. Figure 6C illustrates different exemplary types of parameterization according to an embodiment of the present teachings. As shown, the parameterized AOG includes a parameterized S-AOG and T-AOG. For the parameterized S-AOG, each of its nodes can be parameterized with a reward, which, for example, represents the reward obtained by covering the topic or topic / concept associated with the node. In the context of tutoring, the higher the reward associated with a node in the S-AOG, the greater the value for the agent to teach the concept associated with that node to the student user. Conversely, if the student user is already familiar with the concept associated with a node in the S-AOG (e.g., has already mastered the concept), then the lower the reward assigned to that node, because there is no further benefit in teaching the associated concept to the student. This reward associated with the node can be dynamically updated during the course of the tutoring dialogue. This is shown in Figure 6D where the S-AOG 310 has nodes associated with relevant mathematical concepts related to fractions. As can be seen, each node representing a relevant concept is parameterized with a reward that is estimated to indicate whether it is rewarding to teach the concept to the student.
[0098] Each node in the S-AOG can have different branches, and each branch leads to another node associated with a different topic. Such branches can also be associated with parameters such as the probability of taking the corresponding branch, as shown in Figure 6C . Figure 6D Also illustrated in Figure 6D is the parameterization of paths in the AOG. Teaching fractions may require building knowledge starting from addition and subtraction, and then moving on to multiplication and division. Along each connection between different concepts, there is a probability of moving from one to the other. For example, as shown in the figure, from the "addition" node 310-4 to the "subtraction" node 310-5, the parameterized probability P a,s can indicate the likelihood of successfully teaching a student to understand the concept of "subtraction" if the concept of "addition" is taught first. Conversely, the probability P s,a can indicate the likelihood of successfully teaching a student to understand addition if subtraction is taught first. As another example, the transitions from "addition" to "multiplication" / "division" are parameterized with probabilities P a,m and P a,d respectively. Similarly, the transitions from "subtraction" to "multiplication" / "division" are also parameterized with probabilities P s,m and P s,d respectively. With such probabilities, the dialogue agent can maximize the probability of successfully teaching the intended concept by choosing an optimized path in an order that may work better. Such probabilities can also be dynamically updated based on, for example, observations from the dialogue. In this way, the optimal process of teaching a student can be adjusted in real time based on individual circumstances.
[0099] Parameterization can also be applied to the T-AOG, as indicated in Figure 6C . As discussed herein, the T-AOG represents a dialogue strategy for a specific topic. Each node in the T-AOG represents a specific step in the dialogue, which is often related to what the dialogue agent is going to say to the user, or what the user is going to respond to the dialogue agent, or the evaluation of a transition. As discussed herein, it often happens that the same thing can be said in different ways, and any way of saying it should be considered to convey the same thing. Based on this observation, the content associated with the nodes in the T-AOG can be parameterized. According to an embodiment of the present teachings, this is shown in Figure 6E . As shown in Figure 6B , there are different ways to conduct a greeting dialogue. Even for such a simple topic, there can be many different ways to express almost the same thing. The content of the greeting dialogue can be parameterized in a more simplified T-AOG. Figure 6E shows the Figure 6BThe exemplary parameterized T-AOG corresponding to the T-AOG shown in [___]. The initial greeting is now parameterized as "[___] Good!" 650, where the content in the brackets is parameterized with possible instances "Morning", "Afternoon", and "Evening". The user's response to the initial greeting is now divided into two cases, one is a verbal response 655-1, and the other is no verbal response or silence 655-2. The verbal response 655-1 can be parameterized with different content selections in response to the initial greeting, as shown in the braces associated with 655-1. That is, any content included in the parameterized set 655-1 can be recognized as a possible answer from the user's response to the initial greeting from the agent. Similarly, in response to the user's answer, the content of this response from the agent at 660-1 can also be parameterized as a set of all possible responses. In the case of the user's silence, the response content 660-2 of the agent can also be parameterized similarly.
[0100] Figure 6F Another example of parameterizing the content associated with the nodes in the T-AOG is illustrated in [___]. This example relates to a T-AOG for testing students on the concept of "addition". As shown, the T-AOG for such a test may include the following steps: posing a question (665), asking the student for an answer (667), the student providing an answer (670-1 or 675-1), responding to the user's answer (670-2 or 675-2), and then evaluating the reward associated with the S-AOG node for "addition" (677). For Figure 6F For each node of the T-AOG in [___], the content associated with it is parameterized. For example, for step 665, the parameters involved include X, Y, Oi, where X and Y are numbers and Oi refers to an object of type i. By instantiating specific values of these parameters, many questions can be formed. In Figure 6F In this example in [___], the first step at 665 of the test is to present X objects of type 1 ("o1") and Y objects of type 2 ("o2"), where X and Y are instantiated with numbers (3 and 4) and "o1" and "o2" can be instantiated with object types (such as apples and oranges). Based on this instantiation of the parameters, specific test questions can be generated. In Figure 6FAt 667 of the T-AOG in the Chinese context, to test students, the dialogue agent will ask the user what the sum of X objects of type "o1" and Y objects of type "o2" is. When X, Y, o1, and o2 are instantiated with specific values, such as X = 3, Y = 4, o1 = apples, and o2 = oranges, the text question will be presented as "3 apples, 4 oranges" (or even its picture), and the test question can be asked by instantiating the parameterized question to ask for the sum of X + Y. For example, "How many fruits are there?" or "Can you tell me the total number of fruits?" In this way, flexible test questions can be generated under the general and parameterized T-AOG. Similarly, the parameterized test questions can also facilitate the generation of the expected correct answers. In this example, since X and Y are instantiated as 3 and 4 respectively, the expected correct answer for the summation test question can be dynamically generated as X + Y = 7. Then this dynamically generated expected correct answer can be used to evaluate the answer from the student user in response to the question. In this way, the T-AOG can be parameterized with a simpler graph structure while enabling the dialogue agent to flexibly configure different dialogue contents in the parameterized framework to perform the expected tasks.
[0101] As discussed in this article, when the parameterized content is instantiated, the dialogue agent can also dynamically derive the basis for the evaluation of the test. In this example, the expected correct answer 7 is formed based on the instantiation of X = 3 and Y = 4. When an answer is received, the answer can be classified as an unanswered answer at 670-1 (for example, when the user does not respond at all or the answer does not contain a number) or an answer with a number at 675-1 (correct or incorrect number). The response to the answer can also be parameterized. Both the unanswered answer and the incorrect answer can be considered as non-correct answers, and the parameterized response can be used to respond to this non-correct answer at 670-2.
[0102] The response to the non-correct answer can also be classified into different cases, and in each case, the response content can be parameterized using the appropriate content suitable for that classification. For example, when the non-correct answer is an incorrect total, the response to it can be parameterized to address the incorrect answer, which may be due to an error or a guess. If the non-correct answer is because the user simply does not know (for example, does not answer at all), then the response can be parameterized to directly target that situation using appropriate response alternatives. Similarly, the response to the correct answer can also be parameterized to address the fact that it is indeed the correct answer or an estimated lucky guess.
[0103] As Figure 6FAs shown, after the dialogue agent responds to the answer (or no answer) from the user in different situations, T-AOG includes steps at 677 for evaluating the current reward for mastering the "addition" concept. After this evaluation, the process can return to 665 to test the student on more questions. If the evaluation reveals that the student does not fully understand the concept, then the process can also continue with "teaching", or if the student is considered to have mastered the concept, then the process exits. In some cases, the process can also encounter an exception. For example, if it is detected that the student is simply unable to complete the assignment, then the system can consider temporarily switching topics, such as regarding Figure 6A as discussed.
[0104] In addition, since T-AOG corresponds to the dialogue strategy (which indicates alternative possible flows of the conversation), the actual conversation can traverse a part of T-AOG by following a specific path in T-AOG. Different users in different dialogue sessions or the same user can produce different paths embedded in the same T-AOG. This information can be useful for allowing the dialogue system 100 to personalize the dialogue by parameterizing the links along different paths regarding different users, and such parameterized paths can indicate what works and what does not work for each user. For example, for each link between two nodes in T-AOG, the reward for that link can be estimated based on the performance of each student user's understanding of the underlying concept being taught. This path-centered reward can be calculated based on the probabilities associated with different branches of each node along the path. Figure 6G Illustrates a T-AOG associated with a user according to an embodiment of the present teachings, having different paths between different nodes, the different paths being parameterized with rewards updated based on dynamic information observed during the conversation with the user. In this exemplary parameterized T-AOG (similar to Figure 6F the T-AOG presented in), after presenting X object 1s and Y object 2s to the student user at 680, the agent asks the user at 685 for the total number of X+Y. Based on previous teaching or testing of the same user, there can be an estimated likelihood of how the student will proceed in this round, i.e., for each possible outcome (690-1, 680-2, and 690-3), there are associated rewards R11, R12, and R11, respectively.
[0105] If the answer from the student is incorrect (690-2), then there can be different ways to respond to it, such as 695-1, 695-2, and 695-3. Based on the user's past experience or known personality (also estimated in a personalized manner),
[0106] There can be different reward scores R22, R23, and R24 associated with each possible response. For example, if it is known that the user is sensitive and performs better in an encouraging or positive manner, then the reward associated with response 695-2 can be the highest. In this case, the dialogue system 100 can choose to utilize response 695-2 to respond to an incorrect answer, which is more positive in terms of encouragement. For example, the dialogue agent can say "Almost there. Think again." Different users may prefer not to be told of their mistakes, in which case, for that user, the reward R22 linked to response 695-1 can be the highest compared to R23 and R24. Such reward scores associated with alternative paths of the T-AOG are personalized based on knowledge of the specific user and / or past interactions with the user. By configuring the AOG by leveraging parameters for both nodes and paths, the dialogue system 100 can dynamically configure and update the parameters during each dialogue to personalize the AOG, and thus these dialogues can be conducted in a flexible (the content is parameterized), personalized (the parameters are calculated based on the information for personalization), and therefore more efficient manner.
[0107] As discussed herein and in Figure 2A shown, the information state 110 is represented based not only on the AOG but also on various types of information (such as the dialogue context 260, the dialogue history 250, the user profile 290, event-centric knowledge 270, and some commonsense models 280). The representation of different mental states 200-220 is determined based on the dynamically updated AOG and other information from 250-290. For example, while the AOGs are used to represent different mental states, their corresponding PGs (the results of traversing the AOGs based on the dialogue) are generated based on the actual traversal (nodes and paths) in the AOGs and the dynamic information collected during the dialogue. For example, the values of the parameters associated with the nodes / links in the AOG can be dynamically estimated based on the ongoing dialogue, the dialogue history, the dialogue context, the user profile, the events that occur during the dialogue, etc. Given this, to update the information state 110, different types of information (such as knowledge about events, the surrounding environment, the characteristics of the user, activities, etc.) can be tracked, and then this tracked knowledge can be used to update different parameters and ultimately update the information state 110.
[0108] As discussed herein, AOG / PG is used to represent different mental states, including the mental state of a robotic agent (designed according to what is expected to be accomplished), the representation of a shared mental state between the robotic agent and the user (derived based on the actual conversation that takes place), and the mental state of the user (estimated based on the conversation that takes place and the user's performance in the conversation). When AOG and PG are parameterized, the values of the parameters associated with the nodes and links can be evaluated based on, for example, information related to the conversation, the user's performance, the user's characteristics, and optionally one or more events that occur during the conversation, etc. Based on such dynamic information, the representation of such mental states can be updated over time during the conversation based on changing circumstances.
[0109] Figure 7A FIG. depicts a high-level system diagram of a knowledge tracking unit 700 according to an embodiment of the present teachings, the knowledge tracking unit 700 being used to track information and update rewards associated with nodes / paths in AOG / PG. As discussed herein, the nodes in AOG can be parameterized with state-related rewards respectively, and the paths in PG can also be parameterized with path-related rewards or utilities respectively. The reward / utilities associated with the states or nodes in AOG can include rewards / utilities representing the degree of mastery of the concepts associated with those nodes. The higher the degree of mastery of the concept of an AOG node, the lower the state reward / utility associated with that node, i.e., the reward / utility for teaching a concept that has already been mastered is quite low. The reward / utilities are personalized and obtained based on an assessment of the user's performance, and such an assessment can be continuously carried out instantaneously or periodically during the conversation.
[0110] As discussed herein, each S-AOG node associated with a concept (e.g., to be taught in tutoring) can have one or more T-AOGs, and each T-AOG can correspond to a specific way of teaching that concept. The parsing path or PG is formed based on the nodes and links in the T-AOGs traversed during the conversation. The reward / utility associated with the path in the T-AOG or T-PG can represent the likelihood that this path will lead to a successful tutoring session or successful mastery of the concept. Given this, the better the user's performance evaluated while traversing the path, the higher the path reward / utility associated with that path. Such path-related reward / utilities can also be determined based on the performance of multiple users, which statistically indicates which teaching style works better for a group of users. When determining which branch path to take to continue the conversation, such estimated reward / utilities along different branch paths can be particularly helpful in the conversation session and can guide the conversation agent to select the path that statistically has a better chance of leading to better performance (i.e., reaching a higher degree of mastery of the concept faster).
[0111] Figure 7AThe illustrated embodiment shown in [Figure 0] tracks status and path rewards during a conversation. In this illustrated embodiment, the rewards associated with nodes and paths are determined based on different probabilities estimated based on conversation-based dynamic situations. For example, for a node in the S-AOG associated with a concept such as "addition", its reward is a state-based reward that represents whether there is a return or reward in teaching the "addition" concept to a specific user. For each student user registered to learn mathematics from the robot agent, the reward value for each node in the S-AOG on the mathematical concept is calculated adaptively. The reward for each node in such an S-AOG (e.g., for the mathematical concept "addition") can be assigned an initial reward value, and the reward value can continue to change as the user engages in a conversation indicated by the associated T-AOG (conversation flow regarding the "addition" concept). During the conversation prescribed by the T-AOG, the robot agent can ask the user questions, and then the user answers the questions. The answers from the user can be continuously evaluated, and the probability that the user is learning or making progress is estimated. Such probabilities estimated when traversing the T-AOG can be used to estimate the rewards associated with the nodes in the S-AOG (i.e., the nodes representing the concept "addition"), which indicates whether the user has mastered the concept. That is, the reward value associated with the node representing the concept is updated during the conversation. If the teaching is successful, the reward can be reduced to a low value, which indicates that there is no further value or reward in teaching the student this concept because the student has mastered the concept. As can be seen, such state-based rewards are personalized because they are calculated based on each user's performance in the conversation.
[0112] There are also rewards associated with different paths in the T-AOG, which includes different nodes, each of which can have multiple branches that represent alternative pathways. The selection of different branches results in different traversals of the underlying T-AOG, and each traversal yields a T-PG. In a tutoring application, to track the effectiveness of tutoring, at each node of the T-AOG (traversed during a conversation), different branches can be associated with corresponding measurements that can indicate the likelihood of achieving the desired goal when the corresponding branch is selected. The higher the measurement associated with a branch, the more likely it is to lead to a path that meets the desired purpose. However, optimizing the selection of branches emerging from each individual node may not result in an overall optimal path. In some embodiments, instead of optimizing the individual selection of branches at each node, optimization can be performed on a path basis, i.e., the optimization is performed with respect to a path (of a specific length). In operation, this path-based optimization can be implemented as a look-ahead operation, i.e., what is the best choice at the current branch of the current node when considering the next K selections along the path. This look-ahead operation selects a branch based on a composite measurement along each possible path, which is determined based on the measurements accumulated on the links starting from the current node along each possible path. The length of the look-ahead can vary and can be determined based on application needs. The composite measurement associated with all alternative paths (emanating from the current node) can be referred to as the path-based reward. Then, the branch starting from the current node can be selected by maximizing the path-based rewards for all possible traversals starting from the current node.
[0113] The rewards along the paths of the T-AOG can be determined based on multiple probabilities determined based on the observed performance of the user during the conversation. For example, at the current node in the T-AOG, the dialogue agent can pose a question to the student user and then receive an answer in response to the question from the user, where the answer corresponds to a branch emanating from the node in the T-AOG for that question. Then, the measurement associated with the reward for that branch can be estimated based on probabilities. Such measurements as well as the path-based rewards are personalized as they are calculated based on the personal information observed from a conversation involving a specific user. The measurements associated with different branches along the T-AOG paths (associated with S-AOG nodes) can be used to estimate the rewards of the S-AOG nodes regarding the student's mastery level. These rewards (including node-based rewards and path-based rewards) can constitute the "utility" or preference of the user and can be used by the robotic agent to adaptively determine how to continue the conversation in a utility-driven dialogue plan. This is in Figure 7BAs shown, this figure shows how knowledge can be tracked so that the dialogue system 100 can instantaneously adapt relevant knowledge based on a "shared mind" (which represents an actual conversation) and use this tracked knowledge to dynamically update models (parameters in a parameterized AOG, e.g., rewards for the S-AOG regarding the student / user's mastery of basic concepts in the mind, and / or rewards for different paths in the T-AOG in the agent's mind), and these models can then be used (by the agent) to execute a utility-driven dialogue plan according to the dynamics of the conversation with a specific user.
[0114] Return reference Figure 7A , to perform knowledge tracking and update of the information state 110 based on the tracked knowledge, the knowledge tracking unit 700 includes an initial know probability estimator 710, a know affirmative probability estimator 720, a know negative probability estimator 730, a guess probability estimator 740, a state reward estimator 760, a path reward estimator 750, and an information state updater 770. Figure 7C is a flowchart of an exemplary process of the knowledge tracking unit 700 according to an embodiment of the present teachings. In operation, the initial know probability of nodes in the relevant AOG representation can be estimated first at 705. This can include the initial know probability of each relevant node in the S-AOG as well as each branch of each T-AOG associated with the S-AOG node.
[0115] Using the estimated initial probabilities, the dialogue agent can have a conversation with the user about a specific topic, which is represented by the relevant S-AOG node and a specific T-AOG for that S-AOG node, and the associated probabilities have been initialized. To initiate the conversation, the robotic agent starts the conversation by following the T-AOG. When the user responds to the robotic agent, the NLU engine 140 can analyze the response and generate a language understanding output. In some embodiments, to understand the user's utterance, the NLU engine 140 can also perform language understanding based on information outside the utterance (e.g., information from the multimodal information analyzer 702). For example, the user may say "This is a robotic toy" while pointing at a toy on the table. To understand the semantics of this utterance (i.e., what "this" means), the multimodal information analyzer 702 can analyze audio and visual information to combine clues in different modalities to facilitate the NLU engine 140 in understanding the user's meaning and outputting the user's response, which has an assessment of the correctness of the response based on, for example, the T-AOG.
[0116] When the knowledge tracking unit 700 receives a response from the user with the assessment at 715, to track knowledge based on what has happened in the conversation, different modules can be invoked to estimate corresponding probabilities based on the received input. For example, if the user's response corresponds to the correct answer, then the know-for-sure probability estimator 720 can be invoked to determine the probability associated with knowing the correct answer for sure; the know-for-sure probability estimator 730 can be invoked to estimate the probability associated with not knowing the answer; the guess probability estimator 740 can be invoked to determine the probability of assessing that the user is merely making a guess. If the user's response corresponds to an incorrect answer, then the know-for-sure probability estimator 720 can also determine the probability associated with knowing for sure but the user making a mistake; the know-for-sure probability estimator 730 can estimate the probability associated with not knowing and the user still answering incorrectly; the guess probability estimator 740 can determine the probability that the answer is merely a guess. These steps are performed at 725, 735, and 745 respectively.
[0117] As discussed herein, for T-AOG, when a user interacts with a dialogue agent, such interactions form a parse graph that continues to grow as the conversation progresses. An example is shown in Figure 5A Given the parse graph or history of the interaction between the robot agent and the user, the probability that the user knows the underlying concept can be adaptively updated based on the estimated probabilities. In some embodiments, the probability of initially knowing (or knowing) a concept at time t+1 can be updated based on the observations. In some embodiments, it can be calculated based on the following formula:
[0118]
[0119] where P(L t+1 |obs = correct) represents the probability of initially knowing at time t+1 given the observed correct answer, P(L t+1 |obs = wrong) represents the probability of initially knowing at time t+1 given the observed wrong answer, P(L t ) is the probability of initially knowing at time t, P(S) is the probability of a slip, and P(G) represents the probability of a guess. Thus, using the probabilities estimated based on the observations of the conversation, the probability of prior knowledge can be dynamically updated, as shown in the examples herein. Then, this prior knowledge probability associated with the nodes in S-AOG can be used by the state reward estimator 760 at 755 in Figure 7C to calculate the state-based reward or node-based reward associated with the nodes in S-AOG, which represent the user's mastery of the relevant skills associated with the concept nodes.
[0120] Based on the probabilities calculated for each node along the PG path in the T-AOG for different branches (e.g., some corresponding to the correct answer and some corresponding to the wrong answer), a path-based reward can be calculated by the path reward estimator 750 for each path at 765 in Figure 7C . Based on this estimated state-based reward and path-based reward, the information state updater 770 can then continue to update the parameterized AOG in the information state 110 at 775. When the parameters associated with the AOG in the information state 110 are updated, the updated parameterized AOG can then be used to control the conversation based on the user's utility (preference).
[0121] In some embodiments, the different parameters used to parameterize the AOG can be learned based on observations and / or calculated probabilities. In some implementations, unsupervised learning methods can be employed to learn such model parameters. This includes, for example, knowledge tracing parameters and / or utility / reward parameters. Such learning can be performed online or offline. Below, an exemplary learning scheme is provided:
[0122]
[0123] α1(j) = π j b j (o1), j ∈ [1, N]
[0124]
[0125] β T (i) = 1, i ∈ [1, N]
[0126]
[0127] Figure 8A - 8B Depicts a utility-driven dialogue plan based on dynamically calculated AOG parameters according to an embodiment of the present teachings. The utility-driven dialogue plan can include a dialogue node plan and a dialogue path plan. The former can refer to selecting a node in the S-AOG for continuing the dialogue session. The latter can refer to selecting a path in the T-AOG for conducting the dialogue. Figure 8A Shows an example of a utility-driven tutoring plan for a parameterized S-AOG according to an embodiment of the present teachings. Figure 8B Shows an example of a utility-driven path plan in a parameterized T-AOG according to an embodiment of the present teachings.
[0128] For the node plan, as Figure 6D shown, the exemplary S-AOG 310 is used to teach various mathematical concepts and each node corresponds to a concept. In Figure 8ADifferent nodes are shown, where the nodes have reward-related parameters associated with them, and some of the nodes can be parameterized with conditions formulated based on the rewards of the connected nodes. As Figure 8A seen, node 310-4 is used to teach the concept "addition", node 310-5 is used to teach the concept "subtraction",... and so on. Each node is parameterized, for example, by indicating the reward for teaching the concept, which is related to, for example, the current mastery level of the concept. The rewards associated with some of the nodes in the S-AOG 310 are expressed as a function of the reward parameters from their connected nodes.
[0129] Some concepts may need to be taught subject to requirements or conditions (e.g., prerequisites) that the user has already mastered some other concepts. For example, in order to teach a student the concept of "division", it may be required that the user has already mastered the concepts of "addition" and "subtraction". This can be indicated by requirement 820, which is expressed as R d = F d (R a , R s ), where the reward R d associated with node 310-3 is a function F a of R s and R d , R a and R s represent the rewards associated with node 310-4 for "addition" and node 310-5 for "subtraction", respectively. For example, an exemplary condition for teaching the "division" concept 310-3 can be that its reward level must be high enough (i.e., the user has not yet mastered the concept of "division") and the rewards R a for "addition" (310-4) and R s for "subtraction" (310-5) must be low enough (i.e., the user has already mastered the prerequisite concepts of "addition" and "subtraction"). The mathematical formula of the function Fd can be designed according to the application requirements to meet these conditions.
[0130] A plan based on the nodes can be set such that the dialogue (T-AOG) associated with a node conditional on a certain reward criterion in the S-AOG may not be scheduled until the reward condition associated with the node is satisfied. In this way, initially, when the user knows no concepts, the only unconditional nodes that can be scheduled are 310-4 and 310-5. During the dialogue for "addition" or "subtraction", the associated rewards (R a or R s ) can be continuously updated and propagated to nodes 310-2 and 310-3, such that R m or R d is also updated according to Fm or F d is updated. At a point when the user has grasped the concepts of "addition" and "subtraction", the rewards R a and R s become low enough that there is no need to schedule conversations associated with nodes 310-4 and 310-5. At the same time, the low R a and R s can be substituted into F m or F d , such that the conditions associated with nodes 310-2 and 310-3 can now be satisfied to activate 310-2 and 310-3, because R m or R d can now become high enough that they are ready to be selected to conduct conversations on the topics of multiplication and division. When this occurs, the associated T-AOG can be used to initiate conversations for teaching the corresponding concepts.
[0131] This also applies to the nodes for "fractions". It can be required that the user has grasped the concepts of "multiplication" and "division" (the rewards for 310-2 and 310-3 are low enough), and the rewards for the node "fractions" become reasonably high accordingly. In this way, the state-based rewards associated with the nodes in the S-AOG can be used to dynamically control how to traverse between the nodes in the S-AOG in a personalized manner, for example, in a manner adaptive to the situation related to each individual. That is, in an actual conversation with different users, the traversal can be adaptively controlled in a personalized manner based on the observation of the actual conversation situation. For example, in Figure 8A , depending on the stage of teaching, different nodes can have different rewards at different times. As shown, the node 310-4 for "addition" is the darkest, indicating, for example, the lowest reward value, which can indicate that the user has grasped the concept of "addition". The node 310-5 for "subtraction" has a reward value in between, for example, indicating that the user has not yet grasped the concept but is close. The nodes 310-1, 310-2, and 310-3 are light-colored, thus indicating, for example, high reward values, which represent that the user has not yet grasped the corresponding concepts.
[0132] The path-related or path-based rewards associated with the paths in the T-AOG can also be dynamically calculated based on the observation of the actual conversation and can also be used to adjust how to traverse the T-AOG (how to select branches) during the conversation. Figure 8BIllustrated is an example of a utility-driven path plan for T-AOG according to an embodiment of the present teachings. As shown, when traversing the T-AOG, at each moment, e.g., at time t, after receiving an answer from the user, the robotic agent needs to determine how to respond. During moments 1,..., t, the conversation traverses the parse graph pg 1...t , the states s1, s2,..., s t . To respond, from state s t there can be multiple branches leading to the next state s t+1 .
[0133] To determine which branch to take, look-ahead operations can be performed along alternative paths based on path-based rewards. For example, to look-ahead one step, the rewards associated with the alternative branches originating from s t (taking one step forward) can be considered, and the branch representing the best path-based reward can be selected. To look-ahead two steps, consider the rewards associated with each branch in the first set of alternative branches originating from s t and the rewards associated with each secondary alternative branch (originating from each branch in the first set of alternative branches), and the branch resulting in the best path-based reward is selected as the next step. Deeper look-ahead can also be implemented based on the same principle. Figure 8B The example shown in is a scheme implementing two-step look-ahead, i.e., at time t, the look-ahead scope includes multiple paths at t + 1 and each path among the multiple paths originating from each path at t + 1 at t + 2. Then, branches are selected via look-ahead to optimize the path-based reward.
[0134] The path-based rewards associated with the branches can be initialized first and then updated during the conversation. In some embodiments, the initial path-based rewards can be calculated based on the previous conversations indicated by the user. In some embodiments, such initial path-based rewards can also be calculated based on the previous conversations of multiple users with similar situations. Then, based on how each branch selection leads to the satisfaction of the intended purpose of the conversation, each path-based reward can be dynamically updated over time during the conversation. Based on this dynamically updated path-based reward, the look-ahead optimization scheme can be driven according to the utility (or preference) of each user on how to converse. Thus, it enables adaptive path planning. The following is an exemplary formula for path planning to optimize path selection based on look-ahead operations. In this exemplary formula, a* is the optimal path selected given multiple branch choices a, the current state s t and the parse graph pg 1...t , EU is the expected utility of the branch choice a, and R(st+1, a) represents the reward at state s t+1Reward for selecting a below. As can be seen, the optimization is recursive, which allows look-ahead at any depth.
[0135] a* = arg max EU(a|s t , pg 1...t )
[0136]
[0137] Combined with a state-based, utility-driven node plan, the dialogue system 100 according to this teaching can relate to the intended purpose of the underlying dialogue and dynamically control the conversation with the user based on the knowledge about the user accumulated in the past and the immediate observation of the user. Figure 8C Illustrated is the use of utility-driven dialogue management for a conversation with a student user based on a combination of node and path planning according to this teaching. That is, in a conversation with a student user, the dialogue agent conducts a conversation with the user via utility-driven dialogue management based on dynamic node and path selection based on a parameterized AOG.
[0138] In Figure 8C , the S-AOG 310 includes different nodes for each concept to be taught, with annotated rewards and / or conditions. The rewards associated with the nodes can be determined in advance based on knowledge about the user. For example, as shown, four nodes (310-2, 310-3, 310-4, and 310-5) can have lower rewards (represented as darker nodes), thus indicating, for example, that the student user has mastered the concepts of addition, subtraction, multiplication, and division. There is one node (i.e., a conversation can be arranged) with a high teaching reward, which is 310-1 regarding "fractions". Therefore, selecting one of the S-AOG nodes for a conversation is a reward-driven or utility-driven node plan.
[0139] Node 310-1 is shown as being associated with one or more T-AOGs 320, each T-AOG 320 corresponding to a dialogue strategy that governs the dialogue to teach the student the concept of "fraction". One of the T-AOGs (i.e., 320-1) can be selected to govern the dialogue session, and T-AOG 320-1 includes various steps such as 330, 340, 350, 360, 370, 380... T-AOG 320-1 can be parameterized with, for example, path-based rewards. During the dialogue, path-based rewards can be used to dynamically perform path planning to optimize the likelihood of achieving the goal of teaching the student to master the concept of "fraction". As shown, the highlighted nodes in 320-1 correspond to paths selected based on path planning, forming parse graphs, and representing dynamic traversals based on the actual dialogue. This is to illustrate that knowledge tracking during the dialogue enables the dialogue system 100 to continuously update the parameters in the parameterized AOG to reflect the utility / preferences learned from the dialogue, and such learned utility / preferences in turn enable the dialogue system 100 to adjust its path planning, making the dialogue more effective, engaging, and flexible.
[0140] The AOG for representing the agent's mental state in the information state 110 needs to be created with content before being used for human-machine dialogue. The process of creating the structure of the AOG and the content associated with the nodes / branches is called content creation. Traditionally, the content in the AOG is created by humans, which can be both time-consuming and tedious. Different AOGs can be created in a certain order. For example, the T-AOG for each node in the S-AOG can be created after that node in the S-AOG has been created. This is shown in Figure 9A which depicts an exemplary pattern of creating an AOG with created content according to an embodiment of the present teachings. As can be seen, the creation of the AOG includes creating the S-AOG, and then creating the T-AOGs associated with the nodes in the S-AOG. The present teachings disclose ways to create AOGs via automatic or semi-automatic means.
[0141] When creating an AOG, different creators can create different content. For example, a teacher can be required to create an AOG related to teaching. Different teachers can find different ways to teach a student a certain topic, such as addition and subtraction. Some may find it useful to teach addition first and then subtraction, while some may feel the opposite. Based on their personal experience, they can create different S-AOGs. Additionally, regarding a topic (e.g., "addition" corresponding to a node in the S-AOG), different creators can create different sequences of steps in the dialogue or T-AOG to interact with the student about that topic, thus creating a flexible way to convey the same topic. This is shown in Figure 6BIt shows that there are different ways to perform greetings. According to this teaching, different AOGs on the same topic can sometimes be integrated via graph matching (see Figure 6B and Figure 6E ). This enables the creation of a T-AOG with parameterized content while retaining the structure of the flow as a more concise representation of the parameterized T-AOG. While this can simplify the representation of the T-AOG without losing content, it does not make the content creation process more efficient. This teaching discloses different ways to create content in a more effective manner, including automatic and semi-automatic content creation processes.
[0142] Figure 9B FIG. depicts an exemplary high-level system diagram of a content creation system 900 for automatically creating an AOG via machine learning according to an embodiment of this teaching. Via this content creation system 900 of the AOG, the structures of the S-AOG and the associated T-AOG can be automatically created, and the content associated with the AOG can be automatically created, both via learning. In some embodiments, this automatically learned S-AOG and T-AOG can be further refined by a human, thus enabling a semi-automatic means of creating an AOG. In the illustrated embodiment, the content creation system 900 of the AOG includes a data-driven AOG learning engine 910, which is configured to learn not only the structure of the conversation (which corresponds to the S-AOG) but also the content (T-AOG) associated with different parts of the conversation structure. This learning is based on data from past conversations retrieved from a past conversation database 915. The content stored in the past conversation database 915 can be organized based on different criteria (such as topic, user demographics, feature profiles, etc.). The content from past conversations can be indexed with respect to different classifications so that appropriate content can be used for learning. By accessing the relevant indexed learning content, the data-driven AOG learning engine 910 can learn (both structurally and in terms of conversation content) to create an AOG suitable for a specific type of user.
[0143] To learn the AOG related to a specific item (e.g., tutoring), the data-driven AOG learning engine 910 can access relevant past conversation data about tutoring (e.g., via a project-based index 920) and learn the structure or S-AOG of the tutoring in terms of the conversation flow of sub-concepts and the speech content for each sub-concept involved (i.e., the T-AOG for each sub-concept). A tutoring session on any topic can have several common sub-conversations, and each sub-conversation can be for a sub-concept. Figure 6AAn example is shown in the form of S - AOG, where a tutoring session typically includes different sub - dialogues (nodes in the S - AOG), such as those for greeting (600), chatting (605), or reviewing (610), teaching the expected concept (615), testing (620), and evaluation (625). This general structure of the dialogue related to tutoring can be generally applicable to any tutoring session, regardless of the specific concept or topic to be tutored during the dialogue. This general structure forms the S - AOG for tutoring and can be learned from past dialogue data.
[0144] To learn the structure of the S - AOG associated with a project, the data - driven AOG learning engine 910 can call the topic - based classification model 930 to identify different sub - dialogues and classify them into different topics to derive the underlying structure. For example, past dialogues related to tutoring can be processed, and different sub - dialogue flows for different topics can be identified, and such sub - dialogues can either always exist (AND relationship) or alternatively occur (OR relationship). As Figure 6A shown, chatting (605) and reviewing (610) are connected by an OR relationship, and teaching (615) and testing (620) are two required sub - dialogues in tutoring and are related by an AND relationship. After evaluation (625), the next step can be any one of four possibilities, including returning to review (610), returning to testing (620), returning to teaching (615), or returning to chatting (605). These four possibilities are connected via an OR relationship.
[0145] To learn the T - AOG, the data - driven AOG learning engine 910 can classify different parts of each sub - dialogue into corresponding types based on the nature of the dialogue. For example, as Figure 6EAs shown, the sub-dialogue related to the greeting includes different parts, some parts related to the initial greeting (650), some parts related to the response to the initial greeting (655-1), and some parts related to the response to the response to the initial greeting (660-1). Such different parts can be identified by different means. For example, based on a voice-based method, the utterances from different parties can be identified in this way. Based on the content of the voices from different parties and a topic-based classification model 930, each part can be assigned a label representing the nature of the utterance. Different conversations can have parts with the same label but different content. For example, in response to an initial greeting (e.g., "Good morning"), in one conversation, the response can be "Thank you, and you?". One party in different conversations can respond to the same greeting in different ways, answering "Good morning to you too!". Both responses can be classified or labeled as responses to the initial greeting but have different content. Similarly, the response to the response (660-1) can also be identified from different parties making the response and classified as a response to the response. Such a response to the response to the greeting can also include different utterances (e.g., Figure 6E "Thank you" or "I'm fine" in Figure 6E . As seen in the example shown in
[0146] Figure 9C , each label in the T-AOG can be associated with alternative content that can be said to instantiate the label.
[0147] Figure 9D shows exemplary different types of item-based AOGs derived from machine learning according to an embodiment of the present teachings. Such AOGs (including S-AOG and T-AOG) can be derived by learning from past dialogue data in the manner disclosed herein. For example, based on past dialogue data, AOGs for tutoring mathematics, language,..., chemistry can be obtained from both the relationships between different sub-dialogues (S-AOG) under each AOG and the sequence of utterances in each sub-dialogue (T-AOG) (each step in the sequence is associated with alternative utterances).
[0147] Figure 9DFIG. 0 is a flowchart of an exemplary process of a content creation system for creating an AOG via machine learning according to an embodiment of the present teachings. In this illustrated embodiment, to create an AOG for a particular project (e.g., a tutoring session for teaching mathematics), past conversations related to the project (e.g., all past conversation data corresponding to a tutoring session for teaching a mathematics concept) are first accessed at 950 and used to identify sub-conversations (structures) in each past conversation at 955 based on a topic-based classification model 930, enabling an S-AOG or structure for the project to be obtained at 960 and each node in the S-AOG to be labeled at 965 according to the nature of the sub-conversation. For example, if a sub-conversation is related to teaching a student about the concept of fractions, the corresponding node in the S-AOG structure can be labeled as teaching.
[0148] For each node in the S-AOG linked to a sub-conversation, the past conversation data corresponding to the sub-conversation can then be analyzed at 970 to derive a T-AOG for the S-AOG node. In some embodiments, the content creation system 900 for the AOG can simply adopt certain sub-conversation content to form the T-AOG. In some embodiments, past conversation data from similar sub-conversations involving similar but different conversation content can be used to create different T-AOGs for the S-AOG node. In some embodiments, different T-AOGs learned from past conversation data can also be integrated (e.g., integrated via graph matching) to create one or more merged T-AOGs. In some cases, based on the different T-AOGs to be merged to generate an integrated T-AOG, the conversation content from the different T-AOGs can be used to generate parameterized content for the integrated T-AOG. To create the structure of the T-AOG, the content creation system 900 for the AOG can identify different parts of the sub-conversation (e.g., Figure 6E "initial greeting", "response to the initial greeting", and "response to the response" in ) at 975 based on the topic-based classification model 930. Each such part can be provided with parameterized conversation content that is generated based on the parameterized conversation content of similar parts of different T-AOGs for the same S-AOG node. The concept of a parameterized AOG is discussed herein with reference to Figure 6E - 6G FIG.
[0149] In addition to automatically creating an AOG (including S-AOG and / or T-AOG), an AOG can also be created semi-associatively with each S-AOG node, and the automatically generated AOG can also be inspected, confirmed, refined, or modified by a person via, for example, a user graphical interface. This can include adjusting the S-AOG and / or changing the conversation content in the T-AOG obtained via machine learning. Figure 10AIllustrated is an exemplary visual programming interface 1000 configured for authoring T-AOG content according to an embodiment of the present teachings. At the bottom of this exemplary visual programming interface 1000, it is indicated that the example is the content creation interface for a conversation related to the "greeting" user in a tutoring conversation session. There are questions (Q) and answers (A) (Q–1010-1, 1010-3,... and A–1010-2, 1010-4,...) with the authored text content. Associated with each text content, there is an illustrated "Edit" button to enable editing of the authored text content in the corresponding box.
[0150] In some embodiments, the authored text content 1010-1,... 1010-4,... can initially be created via learning and be displayed in the exemplary visual programming interface 1000 for potential processing. If the text content automatically authored by the machine is acceptable, a person can click the "Save" button 1025 to store the automatically authored text content associated with the T-AOG. A person can also modify the authored text context in the linked box via the "Edit" options (1020-1, 1020-2, 1020-3, 1020-4,...). After modification, "Save" can be clicked to save the modified conversation content associated with the T-AOG. Different people can save different versions of the modified conversation content of the T-AOG. For example, one person may find the text content authored via machine learning acceptable and then can save such automatically generated content for the robot agent to use in tutoring students in a "Addition" tutoring session. While another person may prefer the robot agent to teach the same "Addition" concept in a different way, so he / she may revise what the machine has learned from past conversation data, thus customizing the T-AOG in a different way and saving it accordingly to drive his / her robot agent.
[0151] In some embodiments, although the S-AOG can be automatically generated via learning, the generation of the T-AOG for the associated S-AOG nodes can be done manually, i.e., the exemplary visual programming interface 1000 may not initially display the automatically filled authored text content, but rather a person may need to enter the text content in each box. Even in the case of manually creating the T-AOG content, since the S-AOG is learned via machine learning, the semi-automatic AOG creation process is more efficient than a completely manual process. As discussed herein, in some embodiments, the conversation content of the T-AOG can be parameterized. Figure 10B Illustrated is an exemplary visual programming interface 1030 configured for authoring parameterized T-AOG according to an embodiment of the present teachings. It can be based on Figure 6FThe T-AOG in [the relevant context] provides this exemplary visual programming interface 1030. In some embodiments, Figure 6F the conversation flow / structure in [the relevant context] includes parameterized content and can learn from past conversation data via the content creation system 900 of the AOG. Then, the exemplary visual programming interface 1030 shown in Figure 10B [the relevant context] can be provided to one or more individuals to add selections for the parameterized content.
[0152] As Figure 10B seen, the exemplary visual programming interface 1030 can present different parts related to the T-AOG regarding the concept of testing "addition". Some parts can correspond to instruction parts such as 1040-1 and 1040-4 that only indicate a certain action to be performed (e.g., display (1040-1) something and / or say something (1040-4)). Some parts are editable, such as the underlined part for entering the values of variables X and Y (1040-2) or the content items in the brackets [], to specify, for example, the desired object (1040-3, 1040-4), comment (1040-6, 1040-8, and 1040-9). Some parts are used to provide conditions expressed in curly braces {}, such as 1040-5 and 1040-7. For example, if the question X+Y is presented to a student in a tutoring session and the student answers the question 1040-4 during the session, then the condition for issuing an affirmative comment is that the answer from the student is equal to X+Y, that is, the condition {[input]=X+Y} is satisfied. When [input] is not equal to X+Y, it is stipulated that the conversation is to evaluate whether the incorrect answer is due to not knowing (e.g., pure guess without knowing) or just being incorrect (knowing but saying it wrong). Such an evaluation can be performed instantaneously during the conversation based on probabilities estimated, for example, based on various observations related to the student. If it is evaluated as an incorrect answer, then alternative comments for [incorrect comment] 1040-8 can be created via this authoring tool. If it is evaluated as not knowing the answer, then comments for [incorrect comment] 1040-9 can be created and used to comment on the user's answer.
[0153] The editable parts can be edited, which can be achieved by selecting from a list of optional items via a drop-down menu or in an "edit" mode that allows a person to enter text content or modify an existing text string. For example, the drop-down menu associated with [obj] (1040-3) can be activated by right-clicking on 1040-3. When presenting a list of existing selectable items, a person can make a selection from the list. For example, for [obj1] and [obj2], their drop-down menus can be associated with a list of objects such as apples and oranges. Another editing mode is to enter a text string. For example, the values of variables X and Y can be entered by a person in the edit mode.
[0154] For some editable parts, they can be edited based on both existing selectable options (e.g., learned from past conversation data) and newly entered creative text content. Figure 10B An example of a content item [Positive Comment] 1040 - 6 is provided. When the associated edit button 1045 - 3 is clicked, an additional window 1046 can be popped up. The popped-up additional window 1046 includes pre-existing options for positive comments (each option can be associated with a selection button on the left) and an option to add more (by clicking the "More" button 1049 - 1) conversation content. In this example, there are two pre-existing options ("Well done!" and "That's correct. Well done!") and one of them is selected ("Well done!") and two new entries ("Great!" and "Excellent!") are entered. The button "Accept" 1049 - 2 can be clicked to save the selected and added options as alternative conversation content associated with [Positive Comment], i.e., any saved content item can be recognized in the conversation as a positive comment response to the student's answer about X + Y.
[0155] If the AOG structure is learned from past conversation data via machine learning, then this learned AOG can be directly used by the conversation system to communicate with human users, or they can be further modified or enriched via an exemplary visual programming interface 1030 based on the automatically generated AOG. The results from such a creative tool can also be used to generate program code to execute the underlying conversation specified by the AOG. That is, based on such generated AOGs (S - AOG and T - AOG), code can be automatically generated according to the conversation content embedded in the nodes of the AOG, which will follow the conversation flow specified by the AOG. Therefore, the exemplary visual programming interface 1030 can be regarded as a visual programming tool, and together with the content creation system 900 of the AOG, it can significantly enhance the process of designing and implementing machine agents for conversations.
[0156] Figure 10C It is a flowchart of an exemplary process for creating an AOG with creative conversation content according to an embodiment of the present teachings. Based on past conversation data, the content creation system 900 of the AOG learns the AOG at 1050. In the automatic mode determined at 1055, this machine-learned AOG can be directly used to generate code at 1085 for the machine agent to execute the underlying conversation without further content creation or editing in the semi-automatic mode. If the machine-learned AOG is to be further refined / modified / edited, then it can be done via semi-automatic means with creative tools, such as Figure 10A and 10BThe exemplary visual programming interface 1000 or the exemplary visual programming interface 1030 shown in []. To create or modify the dialogue content associated with the AOG, each machine-learned AOG can be displayed at 1065 for editing or creating content associated with the AOG. As discussed herein, in some embodiments, the learned AOG can be embedded with the dialogue content learned via learning, and such learned content can be used as a basis for further refinement. In some embodiments, the dialogue content can be recreated, as Figure 10A shown in []. During content creation in semi-automatic mode, when the modified or new dialogue content is received at 1070, they are appropriately stored together with the relevant AOG at 1075. If it is determined at 1080 that there are more AOGs to edit, the process proceeds to 1065 to display the next AOG. If all AOGs have been processed, a modified AOG based on the editing results is generated at 1085. Based on the modified / enhanced AOG, code for executing the dialogue specified by the AOG is generated at 1090 based on the created dialogue content.
[0157] Figure 11A Illustrates an exemplary code 1110 generated via visual programming according to an embodiment of the present teachings and the result of the code in presenting a scenario related to the S-AOG. The code in 1110 is generated to present a scenario with a set of objects, which is used to teach the concept of adding numbers to a child student in combination with the S-AOG related to math tutoring. The code 1110 can be generated via visual programming based on the dialogue content associated with the AOG and created via a semi-automatic content creation tool. As can be seen, the code 1110 is programmed to present a set of different types of objects (i.e., 1120-1 represents a pumpkin and 1130-1 represents a strawberry) and the quantity of each type to be presented (1 pumpkin and 2 strawberries). The execution of such code can then generate 1120-3 and 1130-3, which present different quantities of different types of products according to the created dialogue content. This demonstration is created to allow the dialogue to continue, aiming to achieve the intended goal of teaching the student to understand the concept of numbers and / or addition. Therefore, questions can be posed after such a created demonstration, and the machine agent can ask the student user questions about how many pumpkins, how many strawberries, and what the total number of fruits is in the picture.
[0158] Figure 11BIllustrated is an exemplary code generated based on the conversation content in a T-AOG obtained via semi-automatic content creation via visual programming according to an embodiment of the present teachings. The code 1150 shown in the figure implements a conversation represented by a T-AOG 1160 with AND and OR branches by traversing the T-AOG and adopting a specific path (or T-PG) based on the actual progress of the conversation. Such code is automatically generated based on the created T-AOG and the conversation content associated with each node in the T-AOG. For example, in a specific conversation, the machine agent may ask the user at 1130 whether he / she has tooth decay. If the answer to this question is "yes" or "don't know", then the machine agent traverses the conversation by following node 1140. Otherwise, the conversation continues with a diabetes-related question at 1150. If the user's answer to the diabetes question is "yes" / "don't know", then the path to 1170 is traversed, otherwise the path to 1160 is traversed. In this way, via automated AOG learning, or using semi-automatic content creation combined with visual programming, code for the T-AOG can be developed more efficiently.
[0159] As discussed herein, content creation can be done via a creation tool that allows a person to modify existing conversation content learned from past conversation data or enter new conversation content to enrich an existing AOG. In some embodiments, conversation content can also be created based on what the user says and does, so as to provide not only voice data and / or also the manner in which the voice is to be delivered. For example, instead of modifying or entering new text in an interface, a person can simply speak the conversation content to create the created content, and may also have certain expressions, facial, tone, and body movements to convey how the created content is to be delivered. That is, both voice content and metadata about the voice can be created based on a person's behavior. Adhering to such metadata can enable the delivery of conversation content with an intended emotion (e.g., whether angry, happy, or excited).
[0160] Figure 12ADepicts an exemplary high-level configuration of a system 1200 for content creation based on multimodal input from a user in accordance with an embodiment of the present teachings. In the illustrated embodiment, a person 1210 participates in creating content related to an AOG via different means. As discussed above, in some embodiments, the person 1210 can create text dialogue content via his / her computer / device by typing the created dialogue content via a visual programming interface 1205. In addition to this, the present teachings also allow the person 1210 to create content via other means, with instructions that describe the manner in which the content created during a user-machine dialogue will be delivered to the user. For example, a different means of creating content is for the person 1210 to create dialogue content by speaking (instead of typing) the content. In some embodiments, when speaking the created content, the person 1210 can also perform certain activities that can be transformed into instructions regarding the manner of delivering the dialogue content. For example, the person 1210 can speak the content in a particular tone with a particular volume, speed, and pitch, express a particular emotion with a particular facial expression, or make some body movements, all of which can be transformed into instructions such that the created content can be delivered in the manner demonstrated by the person 1210.
[0161] When using speech to create content, the utterance is captured by an audio sensor 1220 and then analyzed by an automatic speech recognizer (ASR) 1240 to convert the speech signal into text as the created content. At the same time, the acoustic signal can also be analyzed to extract acoustic features that can be used as acoustic-based instructions for rendering the created content. This information (features of the utterance rather than the content of the utterance) can be analyzed by an audio / visual (AV)-based instruction generator 1260 and used to generate, for example, acoustic-related rendering instructions. For example, the person 1210 can speak the content in a high pitch, at a fast speed, with anger or a high volume. If the person 1210 speaks the content with a facial expression, then such visual signals can be captured by a camera 1230 and then processed by the A / V-based instruction generator 1260 to convert the visual signals into expression instructions so that the created content can be delivered by a robotic agent with a particular facial expression. For example, the person 1210 can speak the content with a surprised facial expression.
[0162] In some embodiments, acoustic features and facial features can be combined to derive rendering instructions. The person 1210 can have a big smile on his / her face while speaking with an excited vocal characteristic. Both the acoustic and visual information can be captured simultaneously by the audio sensor 1220 and the camera 1230 and used to derive rendering instructions related to both acoustics and expression. In some applications, a robotic device can have a display as its face on its head, and then while the robot speaks a segment of dialogue content based on acoustic-related rendering instructions, a specified expression can be rendered based on the expression instructions.
[0163] In some cases, a robotic agent can have body parts that can be controlled based on instructions to perform certain poses, such as waving, tilting the head, making a fist, leaning forward, etc. The teachings disclosed herein also facilitate generating instructions related to body movement based on the actions of person 1210. Such instructions related to body movement can be used to control the robotic agent to perform certain body actions while rendering some authored dialogue content as part of the expression to be shown. As part of the content authoring content, instructions for such body movement(s) can be automatically generated. As Figure 12A shown, camera 1230 can capture the actions of person 1210, and movement instruction generator 1250 can analyze the body movement of person 1210 and generate instructions as metadata associated with the authored content.
[0164] Instructions generated based on the actions of person 1210 that are intended to direct the robot to achieve certain acoustic / visual / body characteristics while speaking can be referred to as A / V / P instructions. Figure 12B Illustrated are exemplary types of metadata that can be generated and stored together with authored dialogue content according to embodiments of the present teachings. Metadata or instructions associated with a segment of authored dialogue content can be for acoustic features, facial features... and body features. Acoustic features include speed, pitch, tone, and / or volume that are used to convert a segment of dialogue content in text form into its spoken form. Facial features can include a happy expression, a sympathetic expression, a concerned expression, etc. Body features can include raising an arm, making a fist, tilting the head,... or leaning the body, each of which can be further specified based on the body movement of person 1210, such as left, right, forward, and backward.
[0165] A / V / P instructions generated in this way based on the actions of person 1210 can be stored together with their corresponding dialogue content segments. For example, text content 1 can be associated with A / V / P instruction 1270-1; text content 2 can be associated with A / V / P instruction 1270-2; text content 3 can be associated with A / V / P instruction 1270-3; text content 4 can be associated with A / V / P instruction 1270-4;... and so on. In this way, whenever the robotic agent is to speak a particular segment of dialogue content, it can access the associated A / V / P instructions and then control the robot to speak the content with acoustic characteristics (e.g., tone, pitch, speed, volume, etc.) and specific facial expressions (e.g., smiling, frowning, or sad) and body features (e.g., leaning forward, pointing a finger to the sky, or jumping). In this way, dialogue content can be authored in an enriched manner in an efficient automatic or semi-automatic way.
[0166] Figure 12CIt is a flowchart of an exemplary process of a content creation system 1200 for creating content based on multimodal input from a user according to an embodiment of the present teachings. In content creation, the system 1200 first receives multimodal input from different sensors at 1205. Such multimodal sensors may include audio, visual, and other types of sensors. Based on the received audio signal, ASR is performed at 1215 to generate dialogue content created via speech. Additionally, various types of acoustic features (such as pitch, volume, tone, etc.) may be estimated at 1225 based on the received audio signal, and these acoustic features are used to generate acoustic-related rendering instructions associated with the created dialogue content segments at 1235. Meanwhile, based on the received visual signal, this visual input is analyzed at 1245 to extract different visual features. The visual features thus extracted can then be used to estimate facial expressions (if any) in order to generate relevant expression-related rendering instructions for the created dialogue content segments at 1255. Similarly, the extracted visual features can also be used to further estimate body movements (if any) performed by the person 1210 at 1265 in order to generate corresponding body movement-related rendering instructions at 1275. Then, at 1285, for the dialogue content created via speech, such automatically generated rendering instructions can be associated with the dialogue content for storage. If it is determined at 1290 that there is more dialogue content, then the process proceeds to the next segment of dialogue content until all dialogue content segments related to the underlying T-AOG are created.
[0167] Figure 13 It is an illustrative diagram of an exemplary mobile device architecture that can be used to implement a dedicated system for implementing the present teachings according to various embodiments. In this example, the user device on which the present teachings are implemented corresponds to a mobile device 1300, including but not limited to smartphones, tablets, music players, handheld game consoles, global positioning system (GPS) receivers, and wearable computing devices (e.g., glasses, wristwatches, etc.), or any other form of device. The mobile device 1300 may include one or more central processing units (“CPUs”) 1340, one or more graphics processing units (“GPUs”) 1330, a display 1320, a memory 1360, a communication platform 1310 (such as a wireless communication module), a storage device 1390, and one or more input / output (I / O) devices 1340. Any other suitable components, including but not limited to a system bus or controller (not shown), may also be included in the mobile device 1300. As Figure 13As shown, a mobile operating system 1370 (e.g., iOS, Android, Windows Phone, etc.) and one or more applications 1380 can be loaded from a storage device 1390 into a memory 1360 for execution by a CPU 1340. The application 1380 can include a browser or any other suitable mobile application for managing a conversation system on the mobile device 1300. User interaction can be implemented via an I / O device 1340 and provided to an automated conversation partner via a network.
[0168] To implement the various modules, units, and their functions described in this disclosure, a computer hardware platform can be used as the (one or more) hardware platform for one or more of the elements described herein. The hardware elements, operating systems, and programming languages of such computers are conventional in nature, and it is assumed that those skilled in the art are sufficiently familiar with these technologies to adapt those technologies to the appropriate settings as described herein. A computer with user interface elements can be used to implement a personal computer (PC) or other type of workstation or terminal device, but if appropriately programmed, the computer can also act as a server. It is believed that those skilled in the art are familiar with the structure, programming, and general operation of such computer equipment, and thus the drawings should be self-explanatory.
[0169] Figure 14 is an illustrative diagram of an exemplary computing device architecture that can be used to implement a dedicated system for implementing the teachings. This dedicated system incorporating the teachings has a functional block diagram illustration of a hardware platform that includes user interface elements. The computer can be a general-purpose computer or a special-purpose computer. Both can be used to implement the dedicated system of the teachings. This computer 1400 can be used to implement any component of a conversation or dialogue management system as described herein. For example, the conversation management system can be implemented on a computer such as computer 1400 via its hardware, software program, firmware, or a combination thereof. Although only one such computer is shown for convenience, computer functions related to the conversation management system described herein can be implemented in a distributed manner on multiple similar platforms to distribute the processing load.
[0170] For example, computer 1400 includes a COM port 1450 that is connected to a network to which it is attached to facilitate data communication. Computer 1400 also includes a central processing unit (CPU) 1420 in the form of one or more processors for executing program instructions. Exemplary computer platforms include an internal communication bus 1410, various forms of program storage devices and data storage devices (e.g., disk 1470, read-only memory (ROM) 1430, or random access memory (RAM) 1440) for various data files processed and / or transmitted by computer 1400, and program instructions that may be executed by CPU 1420. Computer 1400 also includes I / O components 1460 to support an input / output stream between the computer and other components therein, such as user interface elements 1480. Computer 1400 may also receive programming and data via network communication.
[0171] Thus, as outlined above, aspects of the methods of dialogue management and / or other processes may be implemented in programming. The program aspects of the technology may be thought of as a "product" or "article of manufacture", typically in the form of executable code and / or associated data carried or embodied in a machine-readable medium. Tangible, non-transitory "storage device" type media include any or all memory or other storage devices for a computer, processor, etc., or associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which may provide storage at any time for software programming.
[0172] All or part of the software can sometimes be transmitted through a network such as the Internet or various other telecommunications networks. For example, such communication can enable the software to be loaded from one computer or processor into another, e.g., related to conversation management. Thus, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices, over wired and optical landlines networks, and through various air links. Physical elements that carry such waves (such as wired or wireless links, optical links, etc.) can also be considered media that carry software. As used herein, unless restricted to tangible "storage" media, terms such as computer or machine "readable media" refer to any medium that participates in providing instructions to a processor for execution.
[0173] Thus, machine-readable media can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media includes, for example, optical or magnetic disks, such as any storage device in (one or more) any computer, which can be used to implement the system or any of its components, as shown in the figures. Volatile storage media includes dynamic memory, such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables; copper wire and fiber optics, including the wires that form a bus within a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic medium, CD-ROM, DVD or DVD-ROM, any other optical medium, punched cards, paper tape, any other physical storage medium with a pattern of holes, RAM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, a cable or link transporting such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media can involve transporting one or more sequences of one or more instructions to a physical processor for execution.
[0174] Those skilled in the art will recognize that the present teachings are subject to various modifications and / or enhancements. For example, while the implementations of the various components described above can be implemented in a hardware device, it can also be implemented as a pure software solution - for example, installed on an existing server. Additionally, the fraud network detection techniques disclosed herein can be implemented as firmware, a firmware / software combination, a firmware / hardware combination, or a hardware / firmware / software combination.
[0175] While what has been described above is considered to constitute the present teachings and / or other examples, it should be understood that various modifications can be made thereto and the subject matter disclosed herein can be implemented in various forms and examples, and the teachings can be applied to many applications, only some of which are described herein. The following claims are intended to claim any and all applications, modifications, and variations that fall within the true scope of the present teachings.
Claims
1. A method for adaptive dialogue management, the method being implemented on at least one machine including at least one processor, a memory, and a communication platform capable of connecting to a network, the method comprising: Receiving a language understanding result and an evaluation of the language understanding result, wherein the language understanding result is derived from an utterance of a user participating in a dialogue on a topic, the dialogue being governed by a dialogue strategy, and the evaluation is obtained for an expected result represented in the dialogue strategy; Determining a plurality of probabilities based on the language understanding result and the associated evaluation; Updating a first parameter set associated with the dialogue strategy based on the plurality of probabilities, wherein the first parameter set parameterizes the user with respect to the dialogue strategy and characterizes the effectiveness of the dialogue with the user under the dialogue strategy; Updating a second parameter set associated with the representation of the topic based on the plurality of probabilities, wherein the second parameter set represents a dynamic evaluation of the user's mastery of the topic; the first parameter set and the second parameter set characterize the user's utility with respect to the topic; The utility includes: A state reward, the state reward indicating the level of reward for the dialogue with the user on the topic and being associated with the representation of the topic; and One or more path rewards, each path reward being associated with one of the alternative dialogue paths embedded in the dialogue strategy and representing the effectiveness of the dialogue with the user along that dialogue path; Determining a response to the user based on the user's utility, the utility being dynamically updated based on knowledge tracked with respect to the dialogue strategy for the topic; adaptively determining to continue the dialogue in a utility-driven dialogue plan, the utility-driven dialogue plan including a dialogue node plan and a dialogue path plan; the dialogue node plan refers to selecting a node in the spatial AND-OR graph S-AOG for continuing the dialogue session; the dialogue path plan refers to selecting a path for the dialogue in the temporal AND-OR graph T-AOG.
2. The method according to claim 1, wherein: The dialogue strategy represents alternative ways of having the dialogue with the user on the topic; and The expected result represents an answer from the user, the answer being in response to a statement presented to the user according to the dialogue strategy.
3. The method according to claim 1, wherein The plurality of probabilities includes: A know affirmative probability, the know affirmative probability indicating the likelihood that the user knows the expected result regardless of whether the language understanding result is the same as the expected result; A know negative probability, the know negative probability indicating the likelihood that the user does not know the expected result regardless of whether the language understanding result is the same as the expected result; and A guess probability, the guess probability indicating the likelihood that the user guesses the language understanding result.
4. A machine-readable and non-transitory medium having recorded thereon information for adaptive dialogue management, wherein when the information is read by the machine, the machine is caused to perform the following operations: Receive the language understanding result and the evaluation of the language understanding result, where The language understanding result is derived from the utterances of a user participating in a conversation on a topic, the conversation being governed by a conversation policy, and the evaluation being obtained for the expected result represented in the conversation policy; Determine a plurality of probabilities based on the language understanding result and the associated evaluation; Update a first set of parameters associated with the conversation policy based on the plurality of probabilities, wherein the first set of parameters parameterizes the user with respect to the conversation policy and characterizes the effectiveness of the conversation with the user under the conversation policy; The information also causes the machine, when read by the machine: Update a second set of parameters associated with the representation of the topic based on the plurality of probabilities, wherein the second set of parameters represents a dynamic evaluation of the user's mastery of the topic; The first set of parameters and the second set of parameters characterize the utility of the user with respect to the topic; The utility includes: A state reward that indicates the level of reward for conversing with the user about the topic and is associated with the representation of the topic; and One or more path rewards, each path reward being associated with one of the alternative conversation paths embedded in the conversation policy and representing the effectiveness of the conversation with the user along that conversation path; The information also causes the machine, when read by the machine, the utility being dynamically updated based on knowledge tracked with respect to the conversation policy for the topic: Determine a response to the user based on the user's utility; wherein, adaptively determining to continue the conversation in a utility-driven conversation plan, the utility-driven conversation plan including a conversation node plan and a conversation path plan; The conversation node plan refers to selecting a node in the spatial AND-OR graph S-AOG for continuing the conversation session; The conversation path plan refers to selecting a path for having a conversation in the temporal AND-OR graph T-AOG.
5. The medium according to claim 4, wherein: The conversation policy represents alternative ways of having the conversation with the user about the topic; and The expected result represents an answer from the user that responds to a statement presented to the user according to the conversation policy.
6. The medium according to claim 4, wherein The plurality of probabilities includes: A know affirmative probability that indicates the likelihood that the user knows the expected result, regardless of whether the language understanding result is the same as the expected result; A know negative probability that indicates the likelihood that the user does not know the expected result, regardless of whether the language understanding result is the same as the expected result; and A guess probability that indicates the likelihood that the user guesses the language understanding result.
7. A system for adaptive conversation management, comprising: A knowledge tracking unit configured to receive a language understanding result and an evaluation of the language understanding result, wherein the language understanding result is derived from the utterances of a user participating in a conversation on a topic, the conversation being governed by a conversation policy, and the evaluation being obtained for the expected result represented in the conversation policy; Multiple probability estimators, configured to determine multiple probabilities based on the language understanding result and the associated evaluation; An information state updater, configured to update a first parameter set associated with the dialogue strategy based on the multiple probabilities, wherein the first parameter set parameterizes the user with respect to the dialogue strategy and characterizes the effectiveness of the dialogue with the user under the dialogue strategy; The information state updater is further configured to: update a second parameter set associated with the representation of the topic based on the multiple probabilities, wherein the second parameter set represents a dynamic evaluation of the user's mastery of the topic; the first parameter set and the second parameter set characterize the user's utility; The utility includes: A state reward, which indicates the level of reward for the dialogue with the user on the topic and is associated with the representation of the topic; and One or more path rewards, each path reward is associated with one of the alternative dialogue paths embedded in the dialogue strategy and represents the effectiveness of the dialogue with the user along that dialogue path; It further includes: determining a response to the user based on the user's utility, the utility being dynamically updated based on the knowledge tracked with respect to the dialogue strategy for the topic; adaptively determining to continue the dialogue in a utility-driven dialogue plan, the utility-driven dialogue plan including a dialogue node plan and a dialogue path plan; the dialogue node plan refers to selecting a node in the spatial AND-OR graph S-AOG for continuing the dialogue session; the dialogue path plan refers to selecting a path for the dialogue in the temporal AND-OR graph T-AOG.
8. The system according to claim 7, wherein: The dialogue strategy represents alternative ways of having the dialogue with the user on the topic; and The expected result represents an answer from the user, the answer being in response to a statement presented to the user according to the dialogue strategy.
9. The system according to claim 7, the multiple probability estimators include: A know affirmative probability estimator, configured to estimate a know affirmative probability, which indicates the likelihood that the user knows the expected result regardless of whether the language understanding result is the same as the expected result; A know negative probability estimator, configured to estimate a know negative probability, which indicates the likelihood that the user does not know the expected result regardless of whether the language understanding result is the same as the expected result; and A guess probability estimator, configured to estimate a guess probability, which indicates the likelihood that the user guesses the language understanding result.
Citation Information
Patent Citations
System and method for applying probability distribution models to dialog systems in the troubleshooting domain
US20090112598A1
Systems and methods for interactive dynamic learning diagnostics and feedback
US20190130511A1