Systems and methods for context-aware human-machine conversations for personalization and multimodality
Through multimodal information processing and context perception technology, the problem that traditional dialogue systems cannot adapt to different user conversation modes is solved, personalized and adaptive dialogue management is realized, and user interaction experience is improved.
Patent Information
- Application Number
- CN202080054000.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-17
- Filing Date
- 2020-06-09
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2040-06-09
AI Technical Summary
Traditional computer-assisted dialogue systems cannot effectively adapt to the conversation modes of different human users, resulting in the problem of stimulation or loss of interest during the conversation, and existing systems are difficult to achieve adaptive knowledge representation and understanding of human emotional state in dynamic information.
Multimodal information processing technology is adopted to analyze voice, visual and environmental information through a multimodal information processor, and combine it with surrounding knowledge trackers and spoken comprehension engine to realize context-aware dialogue management, and dynamically update the dialogue content to meet users' personalized needs.
It improves the adaptability and attractiveness of the dialogue system, can better understand user intentions and emotional states, generate personalized responses, and improve interactive experience.
Smart Images

Figure CN114270337B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to the following patent applications: U.S. Provisional Patent Application 62 / 862296, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503589); U.S. Provisional Patent Application 62 / 862,253, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503572); U.S. Provisional Patent Application 62 / 862257, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503574); U.S. Provisional Patent Application 62 / 862261, filed on June 19, 2019 (Attorney Docket No.: 047437 - 0503575); U.S. Provisional Patent Application 62 / 862264, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503578); U.S. Provisional Patent Application 62 / 862265, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503581); U.S. Provisional Patent Application 62 / 862273, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503579); U.S. Provisional Patent Application 62 / 862275, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503580); U.S. Provisional Patent Application 62 / 862279, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503584); U.S. Provisional Patent Application 62 / 862282, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503585); U.S. Provisional Patent Application 62 / 862286, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503586); U.S. Provisional Patent Application 62 / 862290, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503587), and U.S. Provisional Patent Application 62 / 862296, filed on June 17, 2019 (Attorney Docket No.: 047437 - 0503589), the contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] This teaching generally relates to computers. More specifically, this teaching relates to human - machine dialogue management. Background Art
[0004] With the progress of artificial intelligence technology and the explosive growth of Internet-based communications due to ubiquitous Internet connectivity, computer-aided dialogue systems have become increasingly popular. For example, more and more call centers are deploying automated dialogue robots to handle customer calls. Hotels are starting to install various kiosks that can answer questions from tourists or guests. In recent years, automated human-machine communication in other fields has also become increasingly popular.
[0005] Traditional computer-aided dialogue systems are usually pre-programmed with certain dialogue content, such as questions and answers based on well-known conversation patterns in related fields. Unfortunately, some conversation patterns may be suitable for some human users but may not be suitable for others. In addition, human users may go off-topic during the conversation, and continuing with a fixed conversation pattern without considering what the user says may cause irritation or loss of interest, which is not desirable.
[0006] When planning a conversation, human designers usually need to manually create the content of the conversation based on known knowledge, which is time-consuming and tedious. Considering the need to create different conversation patterns, even more labor is required. When creating dialogue content, any deviation from the designed conversation pattern may need to be noted and used to determine how to continue the conversation. Previous dialogue systems have not effectively solved such problems.
[0007] With the latest developments in the field of AI, observed dynamic information can be adaptively incorporated into learning and used to guide the progress of human-machine interaction sessions. How to develop a knowledge representation that can incorporate dynamic information in different dimensions and sometimes in different modalities is a challenging problem. Since this knowledge representation is the basis for the dynamic conversation process between humans and machines, it needs to be fully configured to support adaptive conversation in a relevant manner.
[0008] In order to communicate with humans, an automated dialogue system may need to achieve different levels of understanding of what humans say linguistically, what the semantics of what is said are, sometimes the emotional state of the person, and the mutual causal relationship between what is said and the conversation environment. Traditional computer-aided dialogue systems are not sufficient to solve such problems.
[0009] Therefore, methods and systems are needed to address such limitations. Summary of the Invention
[0010] The teachings disclosed herein relate to methods, systems, and programming for advertising. More specifically, the teachings relate to methods, systems, and programming related to exploring the sources of advertising and their utilization.
[0011] In one example, a method implemented on a machine having at least one processor, a memory, and a communication platform capable of connecting to a network for human-machine conversation. Human-machine conversation. Receive an utterance about a topic in a conversation scenario from a user participating in the human-machine conversation. Obtain and analyze multimodal surround information related to the human-machine conversation to track the multimodal context of the human-machine conversation. Based on the tracked multimodal context, perform an operation of spoken language understanding of the utterance in a context-aware manner to determine the semantics of the utterance.
[0012] In different examples, a system for human-machine conversation includes various functional modules for performing the method of human-machine conversation. Such a human-machine conversation system includes a surround knowledge tracker and a spoken language understanding (SLU) engine. The surround knowledge tracker is configured to obtain multimodal surround information related to the human-machine conversation and analyze the multimodal surround information to track the multimodal context of the human-machine conversation. The SLU engine is configured to receive an utterance about a topic in a conversation scenario from a user participating in the human-machine conversation and personalize the spoken language understanding (SLU) of the utterance in a context-aware manner based on the tracked multimodal context to determine the semantics of the utterance.
[0013] Other concepts relate to software for implementing the present teachings. A software product according to this concept includes at least one machine-readable non-transitory medium and information carried by the medium. The information carried by the medium can be executable program code data, parameters associated with the executable program code, and / or information related to a user, a request, content, or other additional information.
[0014] In one example, a machine-readable non-transitory and tangible medium having data for human-machine conversation recorded thereon, wherein when the medium is read by a machine, the machine performs a series of steps to execute a method for user-machine communication. Human-machine conversation. Receive an utterance about a topic in a conversation scenario from a user participating in the human-machine conversation. Obtain and analyze multimodal surround information related to the human-machine conversation to track the multimodal context of the human-machine conversation. Based on the tracked multimodal context, perform an operation of spoken language understanding of the utterance in a context-aware manner to determine the semantics of the utterance.
[0015] Additional advantages and novel features will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the following and the accompanying drawings, or may be learned by the production or operation of examples. The advantages of the present teachings may be realized and obtained by practicing or using various aspects of the methods, tools, and combinations set forth in the detailed examples discussed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The methods, systems, and / or programs described herein are further described in accordance with exemplary embodiments. These exemplary embodiments are described in detail with reference to the accompanying drawings. These embodiments are non-limiting exemplary embodiments, where like reference numerals in several views of the drawings represent similar structures, and where:
[0017] Figure 1A An exemplary configuration of a dialogue system centered on an information state that captures dynamic information observed during a dialogue in accordance with an embodiment of the present teachings is depicted;
[0018] Figure 1B is a flowchart of an exemplary process of a dialogue system that uses an information state that captures dynamic information observed during a dialogue in accordance with an embodiment of the present teachings;
[0019] Figure 2A An exemplary construction of an information state in accordance with an embodiment of the present teachings is depicted;
[0020] Figure 2B Illustrates a representation of how different estimated mindsets are connected in a dialogue with a robot tutor that teaches addition of fractions in accordance with an embodiment of the present teachings;
[0021] Figure 2C Shows an exemplary relationship between an estimated mindset of an agent, a shared mindset, and an estimated mindset of a user in an information state in accordance with an embodiment of the present teachings;
[0022] Figure 3A Shows an exemplary relationship between different types of And-Or-Graphs (AOGs) for representing the estimated mindsets of the parties involved in a dialogue in accordance with an embodiment of the present teachings;
[0023] Figure 3B Depicts an exemplary association between a Spatial AOG (S-AOG) and a Temporal AOG (T-AOG) in an information state in accordance with an embodiment of the present teachings;
[0024] Figure 3C Illustrates an exemplary S-AOG and its associated T-AOG in accordance with an embodiment of the present teachings;
[0025] Figure 3D Illustrates an exemplary relationship between S-AOG, T-AOG, and C-AOG according to an embodiment of the present teachings;
[0026] Figure 4A Illustrates an exemplary S-AOG according to an embodiment of the present teachings, which partially represents the mindset of an agent for teaching different mathematical concepts;
[0027] Figure 4B Illustrates an exemplary T-AOG according to an embodiment of the present teachings, which represents a dialogue strategy partially associated with the mindset of an agent teaching the concept of fractions;
[0028] Figure 4C Shows exemplary dialogue content for teaching concepts associated with fractions according to an embodiment of the present teachings;
[0029] Figure 5A Illustrates an exemplary temporal parsed graph (T-PG) within a T-AOG according to an embodiment of the present teachings, which represents the shared mindset between a user and a machine;
[0030] Figure 5B Illustrates a part of a dialogue between a machine and a person along a dialogue path according to an embodiment of the present teachings, where the dialogue path represents the current representation of the shared mindset;
[0031] Figure 5C Depicts an exemplary S-AOG according to an embodiment of the present teachings, which has nodes parameterized with measurements related to the level of mastery of different underlying concepts to represent the mindset of a user;
[0032] Figure 5D Shows exemplary types of user personality traits according to an embodiment of the present teachings, which can be estimated based on observations from a dialogue;
[0033] Figure 6A Depicts a general S-AOG for a tutoring dialogue according to an embodiment of the present teachings;
[0034] Figure 6B Depicts a specific T-AOG for a dialogue about greetings according to an embodiment of the present teachings;
[0035] Figure 6C Shows different types of parameterization alternatives for different types of AOG according to an embodiment of the present teachings;
[0036] Figure 6DIllustrated is an S-AOG according to an embodiment of the present teachings, which has different nodes parameterized by rewards that are updated based on dynamic observations from a conversation;
[0037] Figure 6E Illustrated is an exemplary T-AOG generated according to an embodiment of the present teachings by merging different graphs via graph matching using parameterized content;
[0038] Figure 6F Illustrated is an exemplary T-AOG according to an embodiment of the present teachings, which has parameterized content associated with nodes;
[0039] Figure 6G Illustrated is an exemplary T-AOG according to an embodiment of the present teachings, where different paths traverse different nodes, and these different paths are parameterized by rewards that are updated based on dynamic observations from a conversation;
[0040] Figure 7A Depicted is a high-level system diagram of a knowledge tracking unit according to an embodiment of the present teachings;
[0041] Figure 7B Illustrated is how knowledge tracking according to an embodiment of the present teachings enables adaptive dialogue management;
[0042] Figure 7C Is a flowchart of an exemplary process of a knowledge tracking unit according to an embodiment of the present teachings;
[0043] Figure 8A Shows an example of a utility-driven tutoring (node) plan for an S-AOG according to an embodiment of the present teachings;
[0044] Figure 8B Illustrated is an example of a utility-driven path plan for a T-AOG according to an embodiment of the present teachings;
[0045] Figure 8C Illustrated is a dynamic state in utility-driven adaptive dialogue management based on a parameterized AOG according to an embodiment of the present teachings;
[0046] Figure 9A-9B Shows a scheme for enhancing spoken language understanding in a human-machine conversation by automated enrichment of parameterized AOG content according to an embodiment of the present teachings;
[0047] Figure 9C Illustrated is an exemplary way of generating enriched training data for parameterized AOG content according to an embodiment of the present teachings;
[0048] Figure 10ADepicts an exemplary high-level system diagram for enhancing ASR / NLU by training a model based on automatically generated enriched AOG content according to an embodiment of the present teachings;
[0049] Figure 10B Is a flowchart of an exemplary process for enhancing ASR / NLU by training a model based on automatically generated enriched AOG content according to an embodiment of the present teachings;
[0050] Figure 11A Depicts an exemplary high-level system diagram for context-aware spoken language understanding based on surrounding knowledge tracked during a conversation according to an embodiment of the present teachings;
[0051] Figure 11B Illustrates an exemplary type of surrounding knowledge according to an embodiment of the present teachings, which is to be tracked to facilitate context-aware spoken language understanding;
[0052] Figure 12A Illustrates tracking a personal profile based on the conversation that occurs in a conversation according to an embodiment of the present teachings;
[0053] Figure 12B Illustrates tracking a personal profile based on visual observations during a conversation according to an embodiment of the present teachings;
[0054] Figure 12C Shows an exemplary partial personal profile that is tracked during a conversation based on multi-modal input information obtained from the conversation scene according to an embodiment of the present teachings;
[0055] Figure 12D Shows an exemplary event knowledge representation constructed based on the conversation according to an embodiment of the present teachings;
[0056] Figure 13A Illustrates a characteristic group for classifying a user for an adaptive conversation plan according to an embodiment of the present teachings;
[0057] Figure 13B Illustrates an exemplary content / structure of a personalized user profile according to an embodiment of the present teachings;
[0058] Figure 13C Shows an example of establishing a user's group characteristics and their use to facilitate an adaptive conversation plan according to an embodiment of the present teachings;
[0059] Figure 14A Depicts an exemplary high-level system diagram for tracking individual speech-related characteristics and their user profiles to facilitate an adaptive conversation plan according to an embodiment of the present teachings;
[0060] Figure 14B is a flowchart of an exemplary process for tracking individual speech-related characteristics and their user profiles to facilitate adaptive dialogue planning according to an embodiment of the present teachings;
[0061] Figure 15A provides an exemplary structure for constructing a representation of events observed during a dialogue according to an embodiment of the present teachings;
[0062] Figure 15B-15C illustrates an example of tracking event-centric knowledge as dialogue context based on observations from a dialogue scenario according to an embodiment of the present teachings;
[0063] Figure 15D-15E illustrates another example of tracking event-centric knowledge as dialogue context based on observations from a dialogue scenario according to an embodiment of the present teachings;
[0064] Figure 16A depicts an exemplary high-level system diagram for personalized context-aware dialogue management according to an embodiment of the present teachings;
[0065] Figure 16B is a flowchart of an exemplary process for personalized context-aware dialogue management according to an embodiment of the present teachings;
[0066] Figure 16C depicts an exemplary high-level system diagram of an NLG engine and a TTS engine for generating context-aware and personalized audio responses according to an embodiment of the present teachings;
[0067] Figure 16D is a flowchart of an exemplary process of an NLG engine according to an embodiment of the present teachings;
[0068] Figure 16E is a flowchart of an exemplary process of a TTS engine according to an embodiment of the present teachings;
[0069] Figure 17A depicts an exemplary high-level system diagram for adaptive personalized tutoring based on dynamic tracking and feedback according to an embodiment of the present teachings;
[0070] Figure 17B illustrates an exemplary method by which a robotic agent tutor can tutor a student according to an embodiment of the present teachings;
[0071] Figure 17C provides exemplary aspects of a student user that a grader can dynamically observe according to an embodiment of the present teachings;
[0072] Figure 17DProvides examples of standard acoustic / viseme features of speech;
[0073] Figure 17E Shows an example of an adaptive tutoring program designed based on a user's acoustic / viseme features relative to the underlying speech's acoustic / viseme features, according to an embodiment of the present teachings;
[0074] Figure 17F Shows an example of tutoring content for a user based on the deviation of the user's viseme features from the viseme features of the underlying speech, according to an embodiment of the present teachings;
[0075] Figure 17G Is a flowchart of an exemplary process of adaptive personalized tutoring based on dynamic tracking and feedback, according to an embodiment of the present teachings;
[0076] Figure 18 Is an explanatory diagram of an exemplary mobile device architecture that can be used to implement a dedicated system for practicing the present teachings, according to various embodiments; and
[0077] Figure 19 Is an explanatory diagram of an exemplary computing device architecture that can be used to implement a dedicated system for practicing the present teachings, according to various embodiments. Detailed Description
[0078] In the following detailed description, numerous specific details are set forth by way of example in order to facilitate a thorough understanding of the relevant teachings. However, it will be apparent to those skilled in the art that the present teachings may be practiced without these details. In other instances, well-known methods, procedures, components, and / or circuitry have been described at a relatively high level without detail in order to avoid unnecessarily obscuring aspects of the present teachings.
[0079] The present teachings are directed to addressing the deficiencies of traditional human-machine dialogue systems and providing methods and systems that enable rich representation of multimodal information from a conversation environment to allow a machine to have an improved perception of the context and environment of the conversation content and better adapt to the conversation by enhancing interaction with the user. Based on such a representation, the present teachings further disclose different modalities for creating such a representation and authoring the content of a conversation within such a representation. Additionally, to allow for adjustment of these representations based on the dynamics occurring during a conversation, the present teachings also disclose mechanisms for tracking the dynamics of a conversation and updating the representation accordingly, and then the machine uses these dynamics and the representation to conduct a conversation in a utility-driven manner to achieve maximized results.
[0080] Figure 1AFIG. 0 depicts an exemplary configuration of a dialogue system 100 according to an embodiment of the present teachings. The dialogue system 100 is centered around an information state 110 that captures dynamic information observed during a conversation. The dialogue system 100 includes a multimodal information processor 120, an automatic speech recognition (ASR) engine 130, a natural language understanding (NLU) engine 140, a dialogue manager (DM) 150, a natural language generation (NLG) engine 160, and a text-to-speech (TTS) engine 170. The system 100 interacts with a user 180 to conduct a conversation.
[0081] During a conversation, multimodal information is collected from the environment (including from the user 180), which captures surrounding information of the conversation environment, the user's speech, and (facial or body) expressions, etc. The multimodal information thus collected is analyzed by the multimodal information processor 120 to extract relevant features of different modalities in order to estimate different characteristics of the user and the environment. For example, a voice signal can be analyzed to determine voice-related features such as speaking speed, pitch, or even accent. Visual signals related to the user can also be analyzed to extract, for example, facial features or body postures, etc., in order to determine the user's expression. By combining acoustic features and visual features, the multimodal information analyzer 120 may also be able to infer the user's emotional state. For example, a high pitch, fast speaking combined with an angry facial expression may indicate that the user is unhappy. In some embodiments, the observed user activities can also be analyzed to better understand the user. For example, if the user points to or walks towards a specific object, then it can reveal what the user is referring to in his / her speech. Such multimodal information can provide useful context to understand the user's intent. The multimodal information processor 120 can continuously analyze the multimodal information and store such analyzed information in the information state 110, and then the different components in the system 100 use the analyzed information to facilitate decisions related to dialogue management.
[0082] In operation, voice information from user 180 is sent to ASR engine 130 to perform speech recognition. Speech recognition may include identifying the language spoken by user 180 and the words spoken. To understand the semantics of what the user said, the results from ASR engine 130 are further processed by NLU engine 140. This understanding may depend not only on the words spoken, but also on other information (such as the expression and posture of user 180) and / or other context information (such as what has been said previously). Based on the understanding of the user's utterance, dialogue manager 150 determines how to respond to the user. Then, the determined response can be generated in text form via NLG engine 160 and further transformed from text form into a voice signal via TTS engine 170. Then, the output of TTS engine 170 can be delivered to user 180 as a response to the user's utterance. Through this back-and-forth interaction, the process for the machine dialogue system continues to conduct a conversation with user 180.
[0083] As Figure 1A seen, the components in system 100 are connected to information state 110, which, as discussed herein, captures the dynamics surrounding the conversation and provides relevant and rich context information that can be used to facilitate speech recognition (ASR), language understanding (NLU), and various conversation-related determinations, including what is an appropriate response (DM), what language features to apply to the text response (NLG), and how to transform the text response into voice form (TTS) (e.g., what accent). As discussed herein, information state 110 may represent the conversation-related dynamics obtained based on multimodal information, which is related to user 180 or related to the surroundings of the conversation.
[0084] After receiving multimodal information from the conversation scenario (regarding the user or regarding the conversation environment), multimodal information processor 170 analyzes the information and characterizes the conversation surrounding environment in different dimensions, e.g., acoustic characteristics (e.g., the user's pitch, speed, accent), visual characteristics (e.g., the user's facial expression, objects in the environment), physical characteristics (e.g., the user waving or pointing at an object in the environment), the estimated user's emotion and / or mindset, and / or the user's preferences or intentions. This information can then be stored in information state 110.
[0085] The rich media context information stored in the information state 110 can facilitate different components to play their respective roles, enabling the conversation to proceed in an adaptive, more engaging, and more effective manner with respect to the intended goal. For example, the rich context information can improve the understanding of the user 180's utterance based on what is observed in the conversation scenario, evaluate the performance of the user 180 and / or estimate the user's utility (or preference) according to the intended goal of the conversation, determine how to respond to the user 180's utterance based on the estimated emotional state of the user, and deliver the response in the most appropriate manner considered based on the understanding of the user, etc. For example, if accent information represented in the form of sound (e.g., a special way of pronouncing certain phonemes) and visual form (e.g., the user's special visual elements) is captured in the information state regarding the user, then the ASR engine 130 can utilize this information to determine the words spoken by the user. Similarly, the NLU engine 140 can also utilize the rich context information to determine the semantics referred to by the user. For example, if the user points to a computer placed on the table (visual information) and says "I like this", then the NLU engine 140 can combine the output of the ASR engine 130 (i.e., "I like this") and the visual information that the user points to the computer in the room to understand that the user's "this" refers to the computer. As another example, if the user 180 repeatedly makes mistakes during a tutoring session and, at the same time, it is estimated based on the tone of voice and / or the user's facial expression (e.g., they are determined based on multimodal information) that the user is annoyed, then instead of continuing to push forward with the tutoring content, the DM 140 can decide to temporarily change the topic based on the user's known interests (e.g., talking about Lego games) in order to continue to engage the user. The decision to distract the user can be determined based on, for example, the utility of what has worked in the past regarding the user (e.g., intermittently distracting the user has worked in the past) and what has not worked (e.g., continuing to pressure the user to do better).
[0086] Figure 1B is a flowchart of an exemplary process of the dialogue system 100 according to an embodiment of the present teachings, where the information state 110 captures the dynamic information observed during the conversation. As Figure 1BAs can be seen, the process is an iterative one. At 105, multimodal information is received and then analyzed by the multi-information processor 170 at 125. As discussed herein, multimodal information includes information related to the user 180 and / or information related to the dialogue context. Multimodal information related to the user may include the user's utterances and / or visual observations of the user, such as body posture and / or facial expression. Information related to the dialogue context may include information related to the environment, such as the objects present, the spatial / temporal relationship between the user and such observed objects (e.g., the user stands in front of a table), and / or the dynamic relationship between the user's activities and the observed objects (e.g., the user walks towards the table and points to the computer on the table). Then, the understanding of the multimodal information captured from the dialogue scenario can be used to facilitate other tasks in the dialogue system 100.
[0087] Based on the information stored in the information state 110 (representing the past state) and the analysis results from the multi-modal information processor 170 (regarding the current state), the ASR engine 120 and the NLU engine 130 respectively perform speech recognition at 125 to determine the words spoken by the user and the language understanding based on the recognized words. ASR and NLU can be performed based on the current information state 110 and the analysis results from the multi-modal information processor 170.
[0088] Based on the results of the multimodal information analysis and language understanding (i.e., what the user said or meant), the change in the dialogue state is tracked at 135, and such change is used to update the information state 110 accordingly at 145 to facilitate subsequent processing. To execute the dialogue, the DM 140 determines a response at 155 based on the dialogue tree designed for the underlying dialogue, the output of the NLU engine 130 (the understanding of the dialogue utterance), and the information stored in the information state 110. Once the response is determined, the NLG engine 150 generates the response based on the information state 110, for example, in its text form. When the response is determined, there may be different ways to say it. The NLG engine 150 can generate the response at 165 in a style based on the user's preference or a style known to be more suitable for the specific user in the current dialogue. For example, if the user answers a question incorrectly, there may be different ways to point out that the answer is incorrect. For a specific user in the current dialogue, if the user is known to be sensitive and easily frustrated, a more gentle way can be used to tell the user that his / her answer is incorrect to generate the response. For example, instead of saying "This is wrong", the NLG engine 150 can generate a text response of "It's not entirely correct".
[0089] The text response generated by the NLG engine 150 can then be rendered into an audible form, e.g., an audio signal form, by the TTS engine 160 at 175. Although standard or common TTS techniques can be used to perform TTS, the present teachings disclose that the response generated by the NLG engine 150 can be further personalized based on the information stored in the information state 110. For example, if it is known that a slower talking speed or a softer talking manner is more effective for the user, then the generated response can be rendered by the TTS engine 160 at 175 into an audible form, e.g., having a lower speed and pitch. Another example is to render the response with an accent consistent with the known accent of the student based on the personalized information of the user in the information state 110. The rendered response can then be delivered to the user as a response to the user's utterance at 185. After the response to the user, the dialogue system 100 then tracks additional changes in the dialogue and updates the information state 110 accordingly at 195.
[0090] Figure 2A Depicts an exemplary construction of the information state representation 110 according to an embodiment of the present teachings. Non-limitingly, the information state 110 includes a representation for an estimated mindset. As shown, different representations can be estimated to represent, for example, the mindset 200 of the agent, the mindset 220 of the user, and the shared mindset 210, along with other information recorded therein. The mindset 200 of the agent can refer to the (one or more) intended goals to be achieved by the dialogue agent (machine) in a particular dialogue. The shared mindset 210 can refer to a representation of the current dialogue situation, which is a combination of the agent's execution of the intended agenda based on the agent's mindset 200 and the actual performance of the user. The mindset 220 of the user can refer to a representation of the agent's estimate of the state of the student with respect to the intended purpose of the dialogue based on the shared mindset or the performance of the user. For example, if the current task of the agent is to teach a student user the concept of fractions in mathematics (which can include sub-concepts for building an understanding of fractions), then the mindset of the user can include an estimated level of mastery of the user over various related concepts. Such an estimate can be derived based on an assessment of the student's performance at different stages of tutoring such related concepts.
[0091] Figure 2BIllustrated is how such representations of different mindsets are connected in an example where a robotic tutor 205 teaches a concept 215 related to fraction addition to a student user 180 according to an embodiment of the present teachings. As can be seen, the robotic agent 205 interacts with the student user 180 via multimodal interaction. The robotic agent 205 can start tutoring based on an initial representation of the agent's mindset 200 (e.g., a lesson on fraction addition that can be represented as an AOG). During the tutoring, the student user 180 can answer questions from the robotic tutor 205 and such answers to the questions form a specific dialogue path, enabling an estimation of the representation of the shared mindset 210. Based on the user's answers, the user's performance is evaluated, and a representation of the user's mindset 220 is estimated with respect to different aspects, e.g., whether the student has mastered the taught concept and what kind of dialogue style works for this particular student.
[0092] As Figure 2A seen, the representation of the estimated mindset is based on some graph-related forms (including but not limited to the spatio-temporal-causal AND-OR graph STC-AOG 230, the STC parse graph (STC-PG) 240), and can be used in combination with other types of information stored in the information state, such as the dialogue history 250, the dialogue context 260, event-centered knowledge 270, the common sense model 280,... and the user profile 290. These different types of information can belong to multiple modalities and constitute different aspects of the dynamics of each dialogue for each user. Thus, the information state 110 captures the general information of various dialogues as well as the personalized information about each user and each dialogue, and interconnects them to facilitate different components in the dialogue system 100 to perform corresponding tasks in a more adaptive, personalized, and engaging manner.
[0093] Figure 2C Illustrated is an exemplary relationship among the agent's mindset 200, the shared mindset 210, and the user's mindset 220 represented in the information state 110 according to an embodiment of the present teachings. As discussed herein, the shared mindset 210 represents the state of the dialogue achieved via the interaction between the agent and the user, and is a combination of the agent's intended intention (according to the agent's mindset) and the user's performance in following the agent's intended agenda. Based on the shared mindset 210, the dynamics of the dialogue can be traced with respect to what the agent can achieve and what the user can achieve up to that point. [[ID=||10]] [[ID=||11]]
[0094] Tracking this dynamic knowledge enables the system to estimate what the user has achieved up to that point, and in what ways the student user has grasped which concepts or sub - concepts (i.e., which dialogue paths have worked and which might not work). Based on what the student user has achieved so far, the user's mindset 220 can be inferred or estimated, which will be used to determine how the agent can further facilitate the adjustment or update of the dialogue strategy in order to achieve the desired goal or adjust the agent's mindset to suit the user. The process of adjusting the agent's mindset enables the derivation of the updated agent's mindset 200. Based on the dialogue history, the dialogue system 100 learns the user's preferences or what is more effective (utility) for the user. This information, once incorporated into the information state, will be used to adjust the dialogue strategy via a utility - driven (or preference - driven) dialogue plan. The updated dialogue strategy drives the next step in the dialogue, which in turn leads to a response from the user and subsequent updates to the shared mindset, the user's mindset, and the agent's mindset. This process iterates so that the agent can continue to adjust the dialogue strategy based on the dynamic information state.
[0095] According to this teaching, different mindsets are represented based on, for example, STC - AOG and STC - PG. Figure 3A An exemplary relationship between different types of AND - OR graphs (AOGs) for representing the estimated mindsets of the parties involved in a dialogue according to an embodiment of this teaching is shown. An AOG is a graph with AND (conjunction) branches and OR (disjunction) branches. Branches associated with a node in the AOG and connected by an AND relationship represent tasks that need to be traversed in their entirety. Branches emerging from a node in the AOG and connected by an OR relationship represent tasks that can be selectively traversed. As discussed herein, STC - AOG includes an S - AOG corresponding to a spatial AOG, a T - AOG corresponding to a temporal AOG, and a C - AOG corresponding to a causal AOG. According to this teaching, an S - AOG is a graph that includes nodes, each of which can correspond to a topic to be covered in the dialogue. A T - AOG is a graph that includes nodes, each of which can correspond to a temporal action to be taken. Each T - AOG can be associated with a topic or node in the S - AOG, i.e., represents steps to be performed during a dialogue about the topic corresponding to the S - AOG node. A C - AOG is a graph that includes nodes, each of which can be linked to a node in the T - AOG and the corresponding node in the S - AOG, thus representing the action that occurs in the T - AOG and the causal impact of that action on the node in the corresponding S - AOG.
[0096] Figure 3BDepicts an exemplary relationship between nodes in an S-AOG according to an embodiment of the present teachings and nodes of an associated T-AOG represented in information state 110. In this illustration, each K-node corresponds to a node in the S-AOG and represents a skill or topic to be taught in the conversation. The assessment regarding each K-node can include "mastered" or "not yet mastered", for example, the corresponding probabilities P(T) and 1 - P(T), i.e., P(T) represents the transition probability from the not yet mastered to the mastered state. P(L0) represents the probability of prior learning skills or prior knowledge regarding the topic, i.e., the likelihood that the student has already mastered the concept before the tutoring session begins. To teach the skill / concept associated with each K-node, the robot tutor can pose multiple questions according to the T-AOG associated with the K-node, and then the student will answer each question. Each question is shown as a Q-node and the student's answer is represented as an A-node in Figure 3B as seen in
[0097] During the conversation between the agent and the user, the student's answer can be a correct answer A(c) or an incorrect answer A(w), as seen in Figure 3B Based on each answer received from the user, additional probabilities are determined based on various knowledge or observations collected during the conversation, for example. For instance, if the user provides a correct answer (A(c)), then the probability P(G) that the answer was a guess can be determined, which represents the likelihood that the student did not know the correct answer but guessed correctly. Conversely, 1 - P(G) is the probability that the user knew the correct answer and answered correctly. For an incorrect or wrong answer A(w), the probability P(S) can be determined, which represents the likelihood that the student gave an incorrect answer but actually knew the concept. Based on P(S), the probability 1 - P(S) can be estimated, which represents the likelihood that the student gave an incorrect answer because they did not know the concept. Such probabilities can be calculated for each node along the path experienced based on the actual conversation and can be used to estimate when the student has mastered the concept and to estimate what might work and what might not work in teaching this particular student regarding each specific topic.
[0098] Figure 3CIllustrated is an exemplary S-AOG and its associated T-AOG according to an embodiment of the present teachings. In the example S-AOG 310 for guiding the concept of fractions, each node corresponds to a topic or concept to be taught to a student user during a conversation. For example, S-AOG 310 includes a node P0 or 310-1 representing the concept of fractions, a node P1 or 310-2 representing the concept of division, a node P2 or 310-3 representing the concept of multiplication, a node P3 or 310-4 representing the concept of addition, and a node P4 or 310-5 representing the concept of subtraction. In this example, the different nodes in S-AOG 310 are related. For example, to master the concept of fractions, at least some of the other concepts of addition, subtraction, multiplication, and division need to be mastered first. To teach the concept (e.g., fractions) represented by an S-AOG node, the agent may need to perform a series of steps or processes during a conversation session with the student user. Such a process or series of steps corresponds to a T-AOG. In some embodiments, for each node in the S-AOG, there may be multiple T-AOGs, and each T-AOG may represent a different way of teaching the student and may be invoked in a personalized manner. As shown, the S-AOG node 310-1 has multiple T-AOGs 320, one of which is shown as 320-1, which corresponds to a series of time steps such as question / answer 330, 340, 350, 360, 370, 380... etc. In each tutoring session for teaching the concept of fractions, the choice of which T-AOG to use can vary and can be determined based on various considerations (e.g., the user in the session (personalized), the degree of mastery of the current concept (e.g., P(L0)), etc.).
[0099] The representation capture of a dialogue based on STC-AOG captures entities / objects / concepts (S-AOG) related to the dialogue, possible actions (T-AOG) observed during the dialogue, and the impact of each of these actions on these entities / objects / concepts (C-AOG). The actual dialogue activities that occur during the dialogue (voice) cause traversal of the corresponding graphical representation or STC-AOG, resulting in a parse graph (PG) corresponding to the traversed part of the STC-AOG. In some embodiments, the S-AOG can model the spatial decomposition of the objects and scenarios of the dialogue. In some embodiments, the S-AOG can model the decomposition of concepts and sub-concepts as discussed herein. In some embodiments, the T-AOG can model the temporal decomposition of events / sub-events / actions that can be performed or have occurred in a dialogue related to certain entities / objects / concepts represented in the corresponding S-AOG. The C-AOG can model the decomposition of the events represented in the T-AOG and their causal relationships with the corresponding entities / objects / concepts represented in the S-AOG. That is, the C-AOG describes the changes to the nodes in the S-AOG caused by the events / actions taken in the dialogue and represented in the T-AOG. This information is about different aspects of the dialogue and is captured in the information state 110. That is, the information state 110 represents the dynamics of the dialogue between the user and the dialogue agent. This is illustrated in Figure 3D as follows.
[0100] As discussed herein, based on the actual dialogue session, the specific path traversed based on the conversation results in different types of corresponding parse graphs (PGs). For example, it can be applied to the S-AOG to produce an S-PG, to the T-AOG to produce a T-PG, and to the C-AOG to produce a C-PG. That is, based on the STG-AOG, the actual dialogue results in a dynamic STC-PG that at least partially represents the different mental states of the parties participating in the dialogue session. To illustrate this, Figure 4A-4C an exemplary S-AOG / T-AOG is shown, which is associated with the agent's mind for teaching fraction-related concepts; Figure 5A-5B an exemplary representation of the shared mental state is provided via the T-PG, which is generated based on the dialogue in a specific tutoring session; Figure 6A-6B an exemplary representation of the user's mental state in terms of the estimated mastery of different concepts taught in the dialogue with the dialogue agent is shown.
[0101] Figure 4AAn exemplary representation of the agent's mindset regarding fraction tutoring according to an embodiment of the present teachings is shown. As discussed herein, the representation of the agent's mindset can reflect what the agent expects or is designed to cover in a conversation. The agent's mindset can be adjusted during a conversation session based on the user's performance / behavior such that the representation of the agent's mindset captures such dynamics or adjustments. As Figure 4A shown, the exemplary representation of the agent's mindset includes various nodes, each representing a sub-concept related to the concept of fractions. For example, there are sub-concepts related to: "Understanding Fractions" 400, "Comparing Fractions" 405, "Understanding Equivalent Fractions" 410, "Expanding and Simplifying Equivalent Fractions" 415, "Finding Factor Pairs" 420, "Applying the Properties of Multiplication / Division" 425, "Adding Fractions" 430, "Finding the LCM" 435, "Solving for Unknowns in Multiplication / Division" 440, "Multiplication and Division within 100" 445, "Simplifying Improper Fractions" 450, "Understanding Improper Fractions" 455, and "Addition and Subtraction" 460. These sub-concepts can form the landscape of fractions, and some sub-concepts may need to be taught before other sub-concepts. For example, "Understanding Improper Fractions" 455 may need to be covered before "Simplifying Improper Fractions" 450, and "Addition and Subtraction" 460 may need to be mastered before "Multiplication and Division within 100" 445, and so on.
[0102] Figure 4B An exemplary T-AOG according to an embodiment of the present teachings is illustrated, which represents the agent's mindset when teaching concepts related to fractions. As discussed herein, the T-AOG includes various steps associated with a conversation, some of which relate to what the agent says, some of which relate to what the user responds, and some of which correspond to certain evaluations of the conversation performed by the agent. There are branches in the T-AOG that represent decisions. For example, at 470, the corresponding action is for the agent to highlight the numerator and denominator boxes, which, for example, can be after teaching a student what the numerator and denominator are. After link 480, the agent advances to 490 to request user input, for example, requesting the user to tell the agent which of the highlighted ones is the denominator. Based on the answer received from the student, the agent follows two links combined by an OR (plus sign), where each link represents a path taken by the user. For example, if the user correctly answers which one is the denominator, then the agent advances to 490-3, for example, further asking the user to evaluate the denominator. If the user answers incorrectly, then the agent advances to 490-4 to provide a hint about the denominator to the user, and then returns to 490 along link 490-2, again requesting user input about which one is the denominator.
[0103] If the agent asks the user to evaluate the denominator at 490-3, then there are two associated outcomes, one being the wrong answer and the other being the correct answer. The former leads to 490-5, where the agent indicates to the user that the answer is incorrect and then follows the link 490-1 back to 490, asking the user for input again. If the answer is correct, then the agent follows another path forward to 490-6, letting the user know that he / she is correct and continuing along that path to further set the denominator and clear the highlighting at 490-7 and 490-8 respectively. As can be seen, the steps at 490 represent the time actions planned by the agent and related to teaching the concept of the denominator, which are related to the Figure 4A S-AOG in which the concept representing the agent's plan to teach the student the concept of fractions is shown. Thus, they together form a part of the representation of the agent's mental state. Figure 4C Illustrates exemplary dialogue content created for teaching concepts associated with fractions according to an embodiment of the present teaching. By using a similar dialogue strategy, the conversation is intended to be executed in a question-and-answer flow.
[0104] Figure 5A Illustrates an exemplary representation of a shared mental state in the form of a T-PG according to an embodiment of the present teaching (corresponding to the path in the T-AOG in Figure 4B . The highlighted steps form a specific path taken by the actions performed by the dialogue agent in the conversation based on the answers from the user. Compared with the T-AOG shown in Figure 4B , Figure 5A shows the T-PG of the individual highlighted steps (e.g., 470, 510, 520, 530, 540, 550...) along the highlighted path. Figure 5A The T-PG shown in represents the instantiated path traversed based on the actions of both the agent and the user, and thus represents the shared mental state. Figure 5B Illustrates a part of the created dialogue content between the agent and the user according to an embodiment of the present teaching, based on which the representation of the shared mental state can be obtained. As discussed herein, the representation of the shared mental state can be derived based on the flow of the conversation, which forms a specific path or T-PG traversed along the existing T-AOG.
[0105] As discussed herein, during the conversation, the dialogue agent estimates the mental state of the user participating in the conversation based on the observation of the conversation with the user, such that both the conversation and the representation of the estimated mental state of the user are adjusted based on the dynamics of the conversation. For example, in order to determine how to proceed with the conversation, the agent may need to evaluate or estimate the user's mastery of a particular topic based on the observation of the user. As regarding Figure 3BAs discussed, the estimate can be probabilistic. Based on such probabilities, the agent can infer the current level of mastery of the concept and determine how to proceed further in the conversation. For example, if the estimated level of mastery is insufficient, then continue to instruct on the current topic, or if the estimated user's mastery of the current concept is sufficient, then move on to other concepts. The agent can evaluate periodically during the conversation and annotate the PG (parameterized) during this process to facilitate the decision of the next move when traversing the graph. Such an annotated or parameterized S-AOG can produce an S-PG, that is, for example, indicating which nodes in the S-AOG have been sufficiently covered and which have not been sufficiently covered. Figure 5C depicts an exemplary S-PG of a corresponding S-AOG according to an embodiment of the present teachings, which represents the estimated state of mind of the user. The underlying S-AOG is shown in Figure 4A In this illustrated example, during the conversation, each node in this S-AOG is evaluated based on the conversation and parameterized or annotated based on such evaluation. As shown in Figure 5C , the nodes representing different sub-concepts related to the fraction are annotated with corresponding different parameters (these parameters indicate, for example, the level of mastery of the corresponding nodes).
[0106] As shown in Figure 5C , the nodes in the initial S-AOG ( Figure 4A ) are now annotated in Figure 5C with different weights, each weight indicating the evaluated level of mastery of the sub-concept of that corresponding node. As can be seen, the nodes in Figure 5C are presented in different shades, which are determined according to the weights representing different levels of mastery of the underlying sub-concepts. For example, the nodes that are now dotted can correspond to those sub-concepts that have been mastered and thus do not require further traversal. The nodes 560 and 565 (corresponding to "understanding fractions" and "understanding improper fractions") can correspond to sub-concepts that have not reached the required level of mastery. All the nodes connected to these two nodes between these two nodes (for example, mastered and unmastered) can be considered as the reasons why the user has not mastered the concepts of fractions and improper fractions.
[0107] Such an estimated level of mastery of the corresponding nodes in the original S-AOG results in an annotated S-PG, which represents the estimated state of mind of the user, and the estimated state of mind of the user will indicate the degree of understanding of the concepts associated with such nodes. This provides a basis for the dialogue agent to understand the relevant context of the user, for example, what the user has understood about what has been taught and what the user still has questions about. As can be seen, the representation of the state of mind of the user is dynamically estimated based on, for example, the performance and activities of the user during the ongoing conversation. In addition to estimating the level of mastery of the concepts associated with different nodes to understand the state of mind of the user, the contextual observations and information about the user that can be collected during the ongoing conversation can also be used to estimate other characteristics or behavioral metrics of the user as part of understanding the state of mind of the user. Figure 5D Illustrated are exemplary types of user personality traits that can be estimated based on information observed during a conversation according to an embodiment of the present teachings. As shown, during a conversation with a user, based on observations of the user's behavior or expressions (whether verbal, visual, or physical), the agent can estimate various characteristics of the user in different dimensions, such as whether the user is extroverted, how mature the user is, whether the user is naughty, whether the user is easily excited, whether the user is generally cheerful, how confident or secure the user feels about himself / herself, whether the user is reliable, meticulous, etc., via multimodal information processing (e.g., by multimodal information processor 170). Once such information is estimated, it forms a profile of the user, which can affect the dialogue system 100 to determine how to adjust the dialogue strategy of the dialogue system 100 when needed and in what manner the agent of the dialogue system 100 should converse with the user.
[0108] Both the S-AOG and the T-AOG can have certain structures that are organized based on, for example, the topic, concept, or flow of the conversation. Figure 6A Depicted is an exemplary general structure of an S-AOG related to a tutoring conversation according to an embodiment of the present teachings. As Figure 6AThis general structure shown is not specific to a topic, but can be used to teach any topic. Exemplary structures include the different stages involved in a tutoring conversation, which are represented as different nodes in the S-AOG. As shown, node 600 is for conversations related to greetings, node 605 is for chatting about, for example, the weather or health, node 610 is for conversations to review previously learned knowledge (e.g., as a basis for teaching the intended topic), node 615 is for teaching the intended topic, node 620 is for testing the student user on the taught topic, and node 625 is for conversations to evaluate the student user's mastery of the taught topic based on the test. The different nodes can be connected in a way that covers different flows between underlying sub-conversations, but the specific flow in each conversation can be determined dynamically based on the situation. Some branches coming out of a node can be related via an AND relationship, and some branches coming out of a node can be related via an OR relationship.
[0109] As Figure 6A seen, a conversation related to tutoring can start with a greeting conversation 600, such as "Good morning", "Good afternoon", or "Good evening". There are three branches coming out of node 600, including going to node 605 for a short chat, going to node 610 for reviewing previously learned knowledge, and going to node 615 for starting teaching directly. These three branches are ORed together, i.e., the conversation agent can proceed to follow any one of these three branches. After the chat session 605, there are also three branches, one going to the teaching node 615, one going to the testing node 620, and one going to the review node 610. The review node 610 also has two branches, one going to the teaching node 615 and the other going to the testing node 620 (the prior knowledge or prior mastery level of the topic of the student can be tested first before teaching). In this illustrated embodiment, the teaching and testing nodes are required conversations, such that the branches from nodes 605 and 610 to the teaching and testing nodes 615 and 620 are related via AND.
[0110] Teaching and testing can be iterative, as indicated by the two-way arrows between nodes 615 and 620. As needed, either the teaching node 615 or the testing node 620 can proceed to the evaluation node 625. That is, the evaluation can be performed based on the teaching results from node 615 or the testing results from node 620. Based on the evaluation results, the conversation can proceed to one of three alternative paths (related by OR), including teaching 615 (reviewing the concepts again), testing 620 (retesting), or review 610 (reinforcing the user's understanding of some concepts), or even chatting 605 (e.g., if the user is found to be frustrated, then the dialogue system 100 can switch topics to continue engaging the user rather than losing the user). This general S-AOG for tutoring-related conversations is provided by way of illustration and not limitation. The S-AOG for tutoring can be derived according to any logical flow required by the application.
[0111] As Figure 6A seen, each node is itself a conversation and, as discussed herein, can be associated with one or more T-AOGs, each T-AOG representing a conversation flow for an intended topic. Figure 6B Depicts an exemplary T-AOG with conversation content for a greeting authored for the S-AOG node 600 according to an embodiment of the present teachings. The T-AOG can be defined as the conversation strategy. Following the steps defined in the T-AOG is to execute a strategy for achieving certain intended purposes. In Figure 6B each, the content in each rectangular box represents what the agent is to say, and the content in the ellipse represents what the user responds with. As seen, in the Figure 6B T-AOG for greeting shown, the agent first says one of three alternative greetings, namely, good morning 630-1, good afternoon 630-2, and good evening 630-3. The user's response to such a greeting can vary. For example, the user can repeat what the agent said (i.e., good morning, good afternoon, or good evening). Some people will repeat and then add "you too" 635-1. Some people will say "thank you, and you?" at 635-2. Some people will say both 635-1 and 635-2. Some people can just remain silent 635-3. There can be other alternative ways to respond to the agent's greeting. After receiving the response from the user, the dialogue agent can then answer the user's response. For each alternative response from the user, each answer can correspond to the user's response. This is illustrated by the content at 640-1, 640-2, and 640-3 in Figure 6B each.
[0112] Figure 6B The T-AOG shown in each can cover multiple T-AOGs. For example, Figure 6B630-1, 635-2, and 640-2 in Figure 6B can form a T-AOG for greeting. Similarly, 630-1, 635-1, 640-1 can correspond to another T-AOG for greeting; 630-2, 635-1, 640-1 can form another T-AOG; 630-1, 635-3, 640-3 form a different T-AOG; 630-2, 635-3, and 640-3 can form another different T-AOG, and so on. Although different, these alternative T-AOGs all have a substantially similar structure and common content. This commonality can be used to generate a simplified T-AOG that has flexible content associated with each node. This can be achieved via, for example, graph matching. For example, the different T-AOGs related to greeting mentioned above, although having different creative content regarding greeting, all have a similar structure, that is, an initial greeting plus a response from the user and plus a response to the user's response to the greeting. In this sense, Figure 6B the T-AOG in
[0113] may not correspond to the most simplified general T-AOG for greeting.
[0113] To facilitate flexible conversation content and enable the dialogue system 100 to adjust the dialogue in a personalized manner, the AOG can be parameterized. According to different embodiments of the present teachings, such parameterization can be applied to both S-AOG and T-AOG with respect to both the parameters associated with the nodes in the AOG and the parameters associated with the links between different nodes. Figure 6C Illustrates different exemplary types of parameterization according to an embodiment of the present teachings. As shown, the parameterized AOG includes a parameterized S-AOG and T-AOG. For the parameterized S-AOG, each of its nodes can be parameterized with a reward, which, for example, represents the reward obtained by covering the topic or topic / concept associated with the node. In the context of tutoring, the higher the reward associated with a node in the S-AOG, the greater the value for the agent to teach the concept associated with that node to the student user. Conversely, if the student user is already familiar with the concept associated with a node in the S-AOG (e.g., has already mastered the concept), then the lower the reward assigned to that node, because there is no further benefit in teaching the associated concept to the student. This reward associated with the node can be dynamically updated during the process of the tutoring dialogue. This is shown in Figure 6D where the S-AOG 310 has nodes associated with relevant mathematical concepts related to fractions. As can be seen, each node representing the relevant concept is parameterized with a reward that is estimated to indicate whether it is rewarding to teach the concept to the student.
[0114] Each node in the S-AOG can have different branches, and each branch leads to another node associated with a different topic. Such branches can also be associated with parameters such as the probability of taking the corresponding branch, as Figure 6C shown. Figure 6D Also illustrated in Figure 6D is the parameterization of paths in the AOG. Teaching fractions may require building knowledge starting from addition and subtraction, followed by multiplication and division. Along each connection between different concepts, there is a probability of going from one to the other. For example, as shown in the figure, from the "Addition" node 310-4 to the "Subtraction" node 310-5, the parameterized probability P a,s can indicate the likelihood of successfully teaching a student to understand the concept of "subtraction" if the concept of "addition" is taught first. Conversely, the probability P s,a can indicate the likelihood of successfully teaching a student to understand addition if subtraction is taught first. As another example, from "Addition" to "Multiplication" / "Division" are parameterized with probabilities P a,m and P a,d respectively. Similarly, from "Subtraction" to "Multiplication" / "Division" are also parameterized with probabilities P s,m and P s,d respectively. With such probabilities, the dialogue agent can maximize the probability of successfully teaching the intended concept by choosing an optimized path in an order that may work better. Such probabilities can also be dynamically updated based on, for example, observations from the dialogue. In this way, the optimal process of teaching a student can be adjusted in real time based on individual circumstances.
[0115] Parameterization can also be applied to the T-AOG, as Figure 6C indicated. As discussed herein, the T-AOG represents a dialogue strategy for a specific topic. Each node in the T-AOG represents a specific step in the dialogue, which is often related to what the dialogue agent is going to say to the user, or what the user is going to respond to the dialogue agent, or the evaluation of the transition. As discussed herein, it often happens that the same thing can be said in different ways, and any way of saying it should be considered to convey the same thing. Based on this observation, the content associated with the nodes in the T-AOG can be parameterized. According to an embodiment of the present teachings, this is shown in Figure 6E As Figure 6B shown, there are different ways to execute a greeting dialogue. Even for such a simple topic, there can be many different ways to express almost the same thing. The content of the greeting dialogue can be parameterized in a more simplified T-AOG. Figure 6E shows Figure 6BThe exemplary parameterized T-AOG corresponding to the T-AOG shown. The initial greeting is now parameterized as "[___] Good!" 650, where the content in the brackets is parameterized with possible instances "Morning", "Afternoon", and "Evening". The user's response to the initial greeting is now divided into two cases, one with a verbal response 655-1 and the other without a verbal response or silence 655-2. The verbal response 655-1 can be parameterized with different content selections in response to the initial greeting, as shown in the braces associated with 655-1. That is, any content included in the parameterized set 655-1 can be recognized as a possible answer from the user to the initial greeting from the agent. Similarly, in response to the user's answer, the content of such a response from the agent at 660-1 can also be parameterized as a set of all possible responses. In the case of user silence, the response content 660-2 of the agent can be parameterized similarly.
[0116] Figure 6F Another example of parameterizing the content associated with a node in the T-AOG is illustrated. This example relates to a T-AOG for testing students on the concept of "addition". As shown, the T-AOG for such a test can include the following steps: posing a question (665), asking the student for an answer (667), the student providing an answer (670-1 or 675-1), responding to the user's answer (670-2 or 675-2), and then evaluating the reward associated with the S-AOG node for "addition" (677). For Figure 6F For each node of the T-AOG in, the content associated with it is parameterized. For example, for step 665, the parameters involved include X, Y, Oi, where X and Y are numbers and Oi refers to an object of type i. By instantiating specific values of these parameters, many questions can be formed. In Figure 6F In this example in, the first step at 665 of the test is to present X objects of type 1 ("o1") and Y objects of type 2 ("o2"), where X and Y are instantiated with numbers (3 and 4) and "o1" and "o2" can be instantiated with types of objects (such as apples and oranges). Based on this instantiation of the parameters, specific test questions can be generated. In Figure 6FAt the 667th position of the T-AOG in Chinese, in order to test students, the dialogue agent will ask the user what the sum of X objects of type "o1" and Y objects of type "o2" is. When X, Y, o1, and o2 are instantiated with specific values, such as X = 3, Y = 4, o1 = apple, and o2 = orange, the text question will be presented as "3 apples, 4 oranges" (or even its picture), and the test question can be asked by instantiating the parameterized question to ask for the sum of X + Y. For example, "How many fruits are there?" or "Can you tell me the total number of fruits?" In this way, flexible test questions can be generated under the general and parameterized T-AOG. Similarly, the parameterized test questions can also facilitate the generation of the expected correct answers. In this example, since X and Y are instantiated as 3 and 4 respectively, the expected correct answer for the summation test question can be dynamically generated as X + Y = 7. Then this dynamically generated expected correct answer can be used to evaluate the answer from the student user in response to this question. In this way, the T-AOG can be parameterized with a simpler graph structure, while enabling the dialogue agent to flexibly configure different dialogue contents in the parameterized framework to perform the expected tasks.
[0117] As discussed in this article, when the parameterized content is instantiated, the dialogue agent can also dynamically derive the basis for the evaluation used in the test. In this example, the expected correct answer 7 is formed based on the instantiation of X = 3 and Y = 4. When an answer is received, this answer can be classified as an answer of not knowing at 670-1 (for example, when the user does not respond at all or the answer does not contain a number) or an answer with a number at 675-1 (correct or incorrect number). The response to the answer can also be parameterized. Both the answer of not knowing or the incorrect answer can be considered as non-correct answers, and the parameterized response can be used to respond to this non-correct answer at 670-2.
[0118] The response to the non-correct answer can also be classified into different cases, and in each case, the response content can be parameterized using the appropriate content suitable for that classification. For example, when the non-correct answer is an incorrect total, the response to it can be parameterized to address the incorrect answer, which may be due to an error or a guess. If the non-correct answer is because the user simply does not know (for example, does not answer at all), then the response can be parameterized to directly target that situation using appropriate response alternatives. Similarly, the response to the correct answer can also be parameterized to address the fact that it is indeed the correct answer or an estimated lucky guess.
[0119] As Figure 6FAs shown, after the dialogue agent responds to the answer (or lack thereof) from the user in different situations, T-AOG includes steps at 677 for evaluating the current reward for mastering the "addition" concept. After this evaluation, the process can return to 665 to test the student on more questions. If the evaluation reveals that the student does not fully understand the concept, then the process can also continue with "teaching", or if the student is considered to have mastered the concept, then the process exits. In some cases, the process can also encounter an anomaly, for example, if it is detected that the student is simply unable to complete the assignment, then the system can consider temporarily switching topics, such as regarding Figure 6A as discussed.
[0120] Furthermore, since T-AOG corresponds to a dialogue strategy (which indicates alternative possible flows of the conversation), the actual conversation can traverse a part of the T-AOG by following a specific path in the T-AOG. Different users in different dialogue sessions or the same user can produce different paths embedded in the same T-AOG. This information can be useful for allowing the dialogue system 100 to personalize the dialogue by parameterizing the links along different paths for different users, and such parameterized paths can indicate what works and what does not work for each user. For example, for each link between two nodes in the T-AOG, the reward for that link can be estimated based on the performance of each student user's understanding of the underlying concept being taught. This path-centric reward can be calculated based on the probabilities associated with the different branches of each node along that path. Figure 6G illustrates a T-AOG associated with a user according to an embodiment of the present teachings, having different paths between different nodes, the different paths being parameterized with rewards updated based on dynamic information observed during the conversation with the user. In this exemplary parameterized T-AOG (similar to Figure 6F the T-AOG presented in 11 , 12 , 11 ), after presenting X objects 1 and Y objects 2 to the student user at 680, the agent asks the user at 685 for the total number of X+Y. Based on previous teaching or testing of the same user, there can be an estimated likelihood of how the student will proceed in this round, i.e., for each possible outcome (690-1, 680-2, and 690-3), there are associated rewards R
[0121] If the answer from the student is incorrect (690-2), then there can be different ways to respond to it, e.g., 695-1, 695-2, and 695-3. Based on the user's past experience or known personality (also estimated in a personalized manner), there can be different reward scores R 22 , R 23 and R 24 . For example, if it is known that the user is sensitive and does better in an encouraging or positive manner, then the reward associated with response 695-2 can be the highest. In this case, the dialogue system 100 can choose to utilize response 695-2 to respond to the incorrect answer, which is more positive in terms of encouragement. For example, the dialogue agent can say "You're almost there. Think again." Different users may prefer not to be told they are wrong, in which case, for that user, the reward R 23 and R 24 compared to the reward R 22 linked to response 695-1 can be the highest. Such reward scores associated with alternative paths of the T-AOG are personalized based on knowledge of the specific user and / or past interactions with the user. By configuring the AOG by leveraging parameters for both nodes and paths, the dialogue system 100 can dynamically configure and update the parameters during each conversation to personalize the AOG, and thus conduct these conversations in a flexible (content is parameterized), personalized (parameters are calculated based on information for personalization), and therefore more efficient manner.
[0122] As discussed herein and shown in Figure 2A , the information state 110 is represented based not only on the AOG but also on various types of information (such as the dialogue context 260, the dialogue history 250, the user profile 290, event-centered knowledge 270, and some commonsense models 280). The representation of different mental states 200-220 is determined based on the dynamically updated AOG and other information from 250-290. For example, while the AOG is used to represent different mental states, their corresponding PGs (the results of traversing the AOG based on the dialogue) are generated based on the actual traversal (nodes and paths) in the AOG and the dynamic information collected during the dialogue. For example, the values of the parameters associated with the nodes / links in the AOG can be dynamically estimated based on the ongoing dialogue, the dialogue history, the dialogue context, the user profile, the events that occur during the dialogue, etc. Given this, to update the information state 110, different types of information (such as knowledge about events, the surrounding environment, the characteristics of the user, activities, etc.) can be tracked, and then this tracked knowledge can be used to update different parameters and ultimately update the information state 110.
[0123] As discussed herein, AOG / PG is used to represent different mental states, including the mental state of a robotic agent (designed according to what is expected to be accomplished), the representation of a shared mental state between the robotic agent and the user (derived based on the actual conversation that takes place), and the mental state of the user (estimated based on the conversation that takes place and the user's performance in the conversation). When AOG and PG are parameterized, the values of the parameters associated with the nodes and links can be evaluated based on, for example, information related to the conversation, the user's performance, the user's characteristics, and optionally (one or more) events that occur during the conversation. Based on such dynamic information, the representation of such mental states can be updated over time based on changing circumstances during the conversation.
[0124] Figure 7A FIG. depicts a high-level system diagram of a knowledge tracking unit 700 according to an embodiment of the present teachings, the knowledge tracking unit 700 being used to track information and update the rewards associated with the nodes / paths in AOG / PG. As discussed herein, the nodes in AOG can be parameterized with state-related rewards respectively, and the paths in PG can also be parameterized with path-related rewards or utilities respectively. The rewards / utilities associated with the states or nodes in AOG can include rewards / utilities representing the degree of mastery of the concepts associated with the nodes. The higher the degree of mastery of the concept of an AOG node, the lower the state reward / utility associated with that node, i.e., the reward / utility for teaching a concept that has already been mastered is quite low. The rewards / utilities are personalized and obtained based on an assessment of the user's performance, and such assessment can be continuously performed instantaneously or periodically during the conversation.
[0125] As discussed herein, each S-AOG node associated with a concept (e.g., to be taught in tutoring) can have one or more T-AOGs, and each T-AOG can correspond to a specific way of teaching the concept. The parse path or PG is formed based on the nodes and links in the T-AOGs traversed during the conversation. The rewards / utilities associated with the paths in T-AOG or T-PG can represent the likelihood that this path will lead to a successful tutoring session or successful mastery of the concept. Given this, the better the user's performance evaluated while traversing the path, the higher the path reward / utility associated with that path. Such path-related rewards / utilities can also be determined based on the performance of multiple users, which statistically indicates which teaching style works better for a group of users. When determining which branch path to take to continue the conversation, such estimated rewards / utilities along different branch paths can be particularly helpful in the conversation session and can guide the conversation agent to select the path that statistically has a better chance of leading to better performance (i.e., faster reaching the degree of mastery of the concept).
[0126] Figure 7AThe illustrated embodiment shown in FIG. 0 is directed to tracking status and path rewards during a conversation. In this illustrated embodiment, the rewards associated with nodes and paths are determined based on different probabilities estimated based on conversation-based dynamic situations. For example, for a node in the S-AOG associated with a concept such as "addition", its reward is a state-based reward that represents whether there is a return or reward in teaching a particular user the concept of "addition". For each student user registered to learn mathematics from the robot agent, the reward value for each node in the S-AOG on a mathematical concept is adaptively calculated. The reward for each node in such an S-AOG (e.g., for the mathematical concept "addition") can be assigned an initial reward value, and the reward value can continue to change as the user engages in a conversation indicated by the associated T-AOG (conversation flow regarding the "addition" concept). During the conversation prescribed by the T-AOG, the robot agent can ask the user questions, and then the user answers the questions. The answers from the user can be continuously evaluated, and the probability that the user is learning or making progress can be estimated. Such probabilities estimated when traversing the T-AOG can be used to estimate the rewards associated with the nodes in the S-AOG (i.e., the nodes representing the concept "addition"), which indicates whether the user has mastered the concept. That is, the reward value associated with the node representing the concept is updated during the conversation. If the teaching is successful, the reward can be reduced to a low value, which indicates that there is no further value or reward in teaching the student this concept because the student has mastered the concept. As can be seen, such state-based rewards are personalized because they are calculated based on each user's performance in the conversation.
[0127] There are also rewards associated with different paths in the T - AOG. The T - AOG includes different nodes, and each node can have multiple branches, where these branches represent alternative pathways. The selection of different branches leads to different traversals of the underlying T - AOG, and each traversal produces a T - PG. In a tutoring application, to track the effectiveness of tutoring, at each node of the T - AOG (traversed during a conversation), different branches can be associated with corresponding measurements that can indicate the likelihood of achieving the desired goal when the corresponding branch is selected. The higher the measurement associated with a branch, the more likely it is to lead to a path that meets the desired purpose. However, optimizing the selection of branches coming out of each individual node may not result in an overall optimal path. In some embodiments, instead of optimizing the individual selection of branches at each node, optimization can be performed on a path - by - path basis, i.e., the optimization is performed with respect to paths (of a particular length). In operation, this path - based optimization can be implemented as a look - ahead operation, i.e., what is the best choice at the current branch of the current node when considering the next K choices along the path. This look - ahead operation selects branches based on a composite measurement along each possible path, which is determined based on the measurements accumulated on the links starting from the current node along each possible path. The length of the look - ahead can vary and can be determined based on application needs. The composite measurement associated with all alternative paths (originating from the current node) can be referred to as the path - based reward. Then, the branch starting from the current node can be selected by maximizing the path - based rewards for all possible traversals starting from the current node.
[0128] The rewards along the paths of the T - AOG can be determined based on multiple probabilities determined based on the observed performance of the user during the conversation. For example, at the current node in the T - AOG, the dialogue agent can pose a question to the student user and then receive an answer in response to the question from the user, where the answer corresponds to a branch starting from the node for that question in the T - AOG. Then, the measurement associated with the reward for that branch can be estimated based on probabilities. Such measurements and path - based rewards are personalized as they are calculated based on the personal information observed from a conversation involving a particular user. The measurements associated with different branches along the T - AOG paths (associated with S - AOG nodes) can be used to estimate the rewards of the S - AOG nodes regarding the student's mastery level. These rewards (including node - based rewards and path - based rewards) can constitute the user's "utility" or preference, and can be used by the robotic agent to adaptively determine how to continue the conversation in a utility - driven dialogue plan. This is in Figure 7BAs shown, this figure shows how knowledge can be tracked so that the dialogue system 100 can instantaneously adapt relevant knowledge based on a "shared mind" (which represents an actual conversation) and use this tracked knowledge to dynamically update models (parameters in a parameterized AOG, e.g., rewards for the S-AOG regarding the student / user's mastery of basic concepts in the mind, and / or rewards for different paths in the T-AOG in the agent's mind), and these models can then be used (by the agent) to execute a utility-driven dialogue plan according to the dynamics of the conversation with a specific user.
[0129] Return reference Figure 7A , to perform knowledge tracking and update of the information state 110 based on the tracked knowledge, the knowledge tracking unit 700 includes an initial knowledge probability estimator 710, a knowledge affirmative probability estimator 720, a knowledge negative probability estimator 730, a guess probability estimator 740, a state reward estimator 760, a path reward estimator 750, and an information state updater 770. Figure 7C is a flowchart of an exemplary process of the knowledge tracking unit 700 according to an embodiment of the present teachings. In operation, the initial knowledge probability of nodes in the relevant AOG representation can be estimated first at 705. This can include the initial knowledge probability of each relevant node in the S-AOG and each branch of each T-AOG associated with the S-AOG node.
[0130] Using the estimated initial probabilities, the dialogue agent can have a conversation with the user about a specific topic, which is represented by the relevant S-AOG node and the specific T-AOG for that S-AOG node, and the associated probabilities have been initialized. To initiate the conversation, the robotic agent starts the conversation by following the T-AOG. When the user responds to the robotic agent, the NLU engine 120 can analyze the response and generate a language understanding output. In some embodiments, to understand the user's utterance, the NLU engine 120 can also perform language understanding based on information outside the utterance (e.g., information from the multimodal information analyzer 702). For example, the user may say "This is a robotic toy" while pointing at a toy on the table. To understand the semantics of this utterance (i.e., what "this" means), the multimodal information analyzer 702 can analyze the audio and visual information to combine clues in different modalities to facilitate the NUL engine 120 in understanding the user's meaning and outputting the user's response, which has an evaluation of the correctness of the response based on, e.g., the T-AOG.
[0131] When the knowledge tracking unit 700 receives a response from a user with the assessment at 715, different modules can be invoked to estimate corresponding probabilities based on the received input in order to track knowledge based on what has occurred in the conversation. For example, if the user's response corresponds to the correct answer, then the know-certain probability estimator 720 can be invoked to determine the probability associated with definitely knowing the correct answer; the know-uncertain probability estimator 730 can be invoked to estimate the probability associated with not knowing the answer; the guess probability estimator 740 can be invoked to determine the probability of assessing that the user is merely making a guess. If the user's response corresponds to an incorrect answer, then the know-certain probability estimator 720 can also determine the probability associated with definitely knowing but the user making a mistake; the know-uncertain probability estimator 730 can estimate the probability associated with not knowing and the user still answering incorrectly; the guess probability estimator 740 can determine the probability that the answer is merely a guess. These steps are performed at 725, 735, and 745 respectively.
[0132] As discussed herein, for T-AOG, when a user interacts with a dialogue agent, such interactions form a parse graph that continues to grow as the conversation progresses. An example is shown in Figure 5A Given the parse graph or history of the interaction between the robot agent and the user, the probability that the user is aware of the underlying concept can be adaptively updated based on the estimated probabilities. In some embodiments, the probability of initially being aware (or knowing) a concept at time t+1 can be updated based on the observations. In some embodiments, it can be calculated based on the following formula:
[0133]
[0134]
[0135] where P(L t+1 |obs = correct) represents the probability of initially knowing at time t+1 given the observed correct answer, P(L t+1 |obs = wrong) represents the probability of initially knowing at time t+1 given the observed wrong answer, P(L t ) is the probability of initially knowing at time t, P(S) is the probability of a slip, and P(G) represents the probability of a guess. Thus, using the probabilities estimated based on the observations of the conversation, the probability of prior knowledge can be dynamically updated, as shown in the examples herein. Then, this prior knowledge probability associated with the nodes in S-AOG can be used by the state reward estimator 760 at 755 in Figure 7C to calculate the state-based reward or node-based reward associated with the nodes in S-AOG, which represent the user's mastery of the relevant skills associated with the concept nodes.
[0136] Based on probabilities calculated for different branches of each node along the PG path in the T-AOG (e.g., some corresponding to the correct answer and some to the wrong answer), a path reward estimator 750 can calculate a path-based reward for each path at 765 in Figure 7C . Based on this estimated state-based reward and path-based reward, an information state updater 770 can then continue to update the parameters of the parameterized AOG in the information state 110 at 775. When the parameters associated with the AOG in the information state 110 are updated, the updated parameterized AOG can then be used to control the dialogue based on the user's utility (preference).
[0137] In some embodiments, different parameters for parameterizing the AOG can be learned based on observations and / or calculated probabilities. In some implementations, unsupervised learning methods can be employed to learn such model parameters. This includes, for example, knowledge tracing parameters and / or utility / reward parameters. Such learning can be performed online or offline. Below, an exemplary learning scheme is provided:
[0138]
[0139] (α1(j)=π j b j (o1), j ∈ [1, N]
[0140]
[0141] β T (i) = 1, i ∈ [1, N]
[0142]
[0143]
[0144]
[0145]
[0146] Figure 8A-8B Depicts a utility-driven dialogue plan based on dynamically calculated AOG parameters according to an embodiment of the present teachings. The utility-driven dialogue plan can include a dialogue node plan and a dialogue path plan. The former can refer to selecting a node in the S-AOG for continuing the dialogue session. The latter can refer to selecting a path in the T-AOG for conducting the dialogue. Figure 8A Shows an example of a utility-driven tutoring plan for a parameterized S-AOG according to an embodiment of the present teachings. Figure 8BShows an example of utility-driven path planning in a parametric T-AOG according to an embodiment of the present teachings.
[0147] For node planning, as Figure 6D shown, an exemplary S-AOG 310 is used to teach various mathematical concepts and each node corresponds to a concept. In Figure 8A different nodes are shown, the nodes have reward-related parameters associated with them, and some of the nodes can be parameterized with conditions formulated based on the rewards of the connected nodes. As Figure 8A seen, node 310-4 is used to teach the concept "addition", node 310-5 is used to teach the concept "subtraction",... and so on. Each node is parameterized with, for example, an indication of the reward for teaching the concept, which is related to, for example, the current level of mastery of the concept. The rewards associated with some of the nodes in the S-AOG 310 are expressed as a function of the reward parameters from their connected nodes.
[0148] Some concepts may need to be taught subject to requirements or conditions (e.g., prerequisites) that the user has already mastered some other concepts. For example, in order to teach a student the "division" concept, it may be required that the user has already mastered the concepts of "addition" and "subtraction". This can be indicated by requirement 820, which is expressed as R d = F d (R a , R s ), where the reward R d associated with node 310-3 is a function F a of R s , R d , R a , R s represent the rewards associated with node 310-4 regarding "addition" and node 310-5 regarding "subtraction", respectively. For example, an exemplary condition for teaching the "division" concept 310-3 can be that its reward level must be high enough (i.e., the user has not yet mastered the "division" concept) and the reward R a for "addition" (310-4) and the reward R s for "subtraction" (310-5) must be low enough (i.e., the user has already mastered the prerequisite concepts regarding "addition" and "subtraction"). The mathematical formula of the function F d can be designed according to the application needs to meet these conditions.
[0149] A node-based plan can be set such that a dialogue (T-AOG) associated with a node conditional on a certain reward criterion in the S-AOG may not be scheduled until the reward condition associated with that node is satisfied. In this way, initially, when the user knows no concepts, the only unconditional nodes that can be scheduled are 310-4 and 310-5. During a dialogue for "addition" or "subtraction", the associated reward (R a or R s ) can be continuously updated and propagated to nodes 310-2 and 310-3 such that R m or R d is also updated according to F m or F d . At a certain point when the user has mastered the concepts of "addition" and "subtraction", the rewards R a and R s become low enough that the dialogues associated with nodes 310-4 and 310-5 do not need to be scheduled. At the same time, the low R a and R s can be substituted into F m or F d such that the conditions associated with nodes 310-2 and 310-3 can now be satisfied to make 310-2 and 310-3 active because R m or R d can now become high enough that they are ready to be selected to execute a dialogue on the topics of multiplication and division. When this occurs, the associated T-AOG can be used to initiate a dialogue for teaching the corresponding concept.
[0150] This also applies to the nodes for "fractions". It can be required that the user has mastered the concepts of "multiplication" and "division" (the rewards for 310-2 and 310-3 are low enough), and the rewards for the node "fractions" become reasonably high accordingly. In this way, the state-based rewards associated with the nodes in the S-AOG can be used to dynamically control how to traverse between the nodes in the S-AOG in a personalized manner, for example, in a way that adapts to the situation related to each individual. That is, in an actual dialogue with different users, the traversal can be adaptively controlled in a personalized manner based on the observation of the actual dialogue situation. For example, in Figure 8ADepending on the stage of teaching, different nodes can have different rewards at different times. As shown in the figure, node 310-4 regarding "addition" is the darkest, indicating, for example, the lowest reward value, which may indicate that the user has mastered the concept of "addition". Node 310-5 regarding "subtraction" has an intermediate reward value, for example indicating that the user has not yet mastered the concept but is close. Nodes 310-1, 310-2, and 310-3 are light-colored, thus indicating, for example, high levels of reward values, which represent that the user has not yet mastered the corresponding concepts.
[0151] The path-related or path-based rewards associated with the paths in the T-AOG can also be dynamically calculated based on the observation of the actual conversation and can also be used to adjust how to traverse the T-AOG (how to select branches) during the conversation. Figure 8B An example of a utility-driven path plan for the T-AOG according to an embodiment of the present teachings is illustrated. As shown, when traversing the T-AOG, at each moment, for example, at time t, after receiving an answer from the user, the robot agent needs to determine how to respond. During moments 1,..., t, the conversation traverses the parse graphs pg1...t, and the states traversed are s1, s2,... s t . To respond, from state s t There can be multiple branches leading to the next state s t+1 .
[0152] To determine which branch to take, a look-ahead operation can be performed along the alternative paths based on the path-based rewards. For example, to look ahead one step, the rewards associated with the alternative branches originating from s t (taking one step forward) can be considered, and the branch representing the best path-based reward can be selected. To look ahead two steps, consider the rewards associated with each branch in the first set of alternative branches originating from s t and the rewards associated with each secondary alternative branch (originating from each branch in the first set of alternative branches), and the branch resulting in the best path-based reward is selected as the next step. Deeper look-ahead can also be implemented based on the same principle. Figure 8B The example shown in is a scheme implementing two-step look-ahead, that is, at time t, the scope of look-ahead includes multiple paths at t + 1 and each path among the multiple paths originating from each path at t + 1 at t + 2. Then, the branch is selected via look-ahead to optimize the path-based reward.
[0153] Path-based rewards associated with branches can be initialized first and then updated during a conversation. In some embodiments, an initial path-based reward can be calculated based on previous conversations indicated by a user. In some embodiments, such initial path-based rewards can also be calculated based on previous conversations of multiple users with similar situations. Then, based on how each branch choice leads to satisfaction of the expected purpose of the conversation, each path-based reward can be dynamically updated over time during the conversation. Based on such dynamically updated path-based rewards, a look-ahead optimization scheme can be driven according to the utility (or preference) of each user on how to conduct the conversation. Thus, it enables adaptive path planning. The following is an exemplary formula for path planning that optimizes path selection based on look-ahead operations. In this exemplary formula, a* is the optimal selected path given multiple branch choices a, the current state st, and the parse graph pg1…t, EU is the expected utility of the branch choice a, and R(st+1,a) represents the reward for selecting a in state s t+1 As can be seen, the optimization is recursive, which allows look-ahead at any depth.
[0154] a * = arg max EU(a|s t , pg1...t)
[0155]
[0156] Combined with state-based, utility-driven node planning, the dialogue system 100 according to this teaching is capable of relating to the expected purpose of the underlying conversation and dynamically controlling the conversation with the user based on the knowledge about the user accumulated in the past and the immediate observation of the user. Figure 8C Illustrates the use of utility-driven dialogue management for a conversation with a student user based on a combination of node and path planning according to an embodiment of this teaching. That is, in a conversation with a student user, the dialogue agent conducts the conversation with the user via utility-driven dialogue management based on dynamic nodes and path selection based on a parameterized AOG.
[0157] In Figure 8CIn it, the S-AOG 310 includes different nodes for the respective concepts to be taught, with annotated rewards and / or conditions. Rewards associated with the nodes can be determined in advance based on knowledge about the user. For example, as shown, four nodes (310-2, 310-3, 310-4, and 310-5) can have lower rewards (represented as darker nodes), thus indicating, for example, that a student user has mastered the concepts of addition, subtraction, multiplication, and division. There is a node (i.e., 310-1) about "fractions" that has a high teaching reward. Therefore, selecting one of the S-AOG nodes for a conversation is a reward-driven or utility-driven node plan.
[0158] Node 310-1 is shown as being associated with one or more T-AOGs 320, each T-AOG 320 corresponding to a dialogue strategy that dominates the conversation to teach the student the concept of "fractions". One of the T-AOGs (i.e., 320-1) can be selected to dominate the conversation session, and the T-AOG 320-1 includes various steps such as 330, 340, 350, 360, 370, 380... The T-AOG 320-1 can be parameterized with, for example, path-based rewards. During the conversation, path-based rewards can be used to dynamically perform path planning to optimize the likelihood of achieving the goal of teaching the student to master the concept of "fractions". As shown, the highlighted nodes in 320-1 correspond to the paths selected based on path planning, forming an analysis graph, representing the dynamic traversal based on the actual conversation. This is to illustrate that knowledge tracking during the conversation enables the dialogue system 100 to continuously update the parameters in the parameterized AOG to reflect the utility / preferences learned from the conversation, and such learned utility / preferences in turn enable the dialogue system 100 to adjust its path planning, making the conversation more effective, more engaging, and more flexible.
[0159] As Figure 6F shown, parameterized content associated with different nodes can be used to create the T-AOG. The parameterized content associated with each node represents what is expected to be said / heard during the conversation. The more alternative content included in the parameterized content associated with each node, the more flexible the dialogue the parameterized T-AOG can support. This alternative content for each node can be created manually, semi-automatically, or automatically. This created alternative content can also be used as training data to facilitate effective recognition to improve the adaptive ability of dialogue management. Figure 9A-9B A scheme for enhancing speech understanding in human-machine conversations by automatically enriching the parameterized content for the AOG according to an embodiment of the present teaching is shown. Figure 9A Presented is as Figure 6F seen in the T-AOG, except thatFigure 9A In addition to each node in being now associated with one or more alternative content sets that the dialogue manager can use to conduct a dialogue. As discussed herein, such alternative parameterized dialogue content sets can also be used to train ASR and / or NLU models to understand the utterances to be spoken.
[0160] The training data set associated with the node represents the parameterized content of the node. For example, as Figure 9A shown, node 665 is associated with two training data sets [T N and [T O , where the former is for numbers (X and Y can be the data items included therein), and the latter is for object o1 or o2 (e.g., apples, oranges, pears, etc.); node 667 is associated with the training data set [T I for query statements; node 670-1 is associated with the training data set [T NKA for when the user does not know the answer; node 675-1 is associated with the training data set [T N for when the user's answer has numbers; node 670-2 is associated with the training data set [T RNCA for responses to incorrect answers (which is an alternative response to not knowing the answer or an alternative response to a wrong answer); node 670-1 is associated with the training data set [T RCA for responses to correct answers; node 675-2 is associated with the training data set [T NKA for responses to the user not knowing the answer; node 910 is associated with the training data set [T RNK for responses to the user not knowing the answer; node 920 is associated with the training data set [T RI for responses to incorrect answers from the user; node 930 is associated with the training data set [T RC for responses to correct answers from the user; node 940 is associated with the training data set [TR m for responses to wrong answers from the user; and node 950 is associated with the training data set [T RG for responses to the user not knowing the answer.
[0161] Figure 9B Illustrates exemplary training data sets associated with different nodes of the T-AOG in accordance with an embodiment of the present teachings. As shown, for example, [T Figure 9A can be a set of any unit or multi-digit numbers, [T N can include names of alternative objects, [T O can include alternative ways of asking for a total of two numbers; [T I can include alternative ways of asking for a total of two numbers; [TNKA may include alternative ways in which the user says "I don't know"; [T RC may include alternative ways of responding to a correct answer; [T RNK may include alternative ways of responding to a user's lack of knowledge of the answer; [TR m may include alternative ways of responding to a wrong answer from the user; [T RG may include alternative ways of responding to a guessed answer from the user; [T RI may include alternative ways of responding to an incorrect answer from the user, which may include alternative responses to the wrong answer [TR m or alternative responses to the guessed answer [T RG ; [T RNCA may include alternative ways of responding to a non - correct answer from the user, which may include alternative responses to the lack of knowledge of the answer [T RNK or alternative responses to the incorrect answer [T RI ; and [T RCA may include alternative ways of responding to a correct answer from the user, which may include alternative responses to the correct answer [T RC or alternative responses to the guessed answer [T RG . The training dataset for responses (e.g., [T RC , [T RNK , ……, [T RCA ) can be used to generate responses of the robotic agent to the user's answers. This can be supplementary for ASR / NLU to understand the utterance.
[0162] Such enriched training datasets associated with different nodes in the T - AOG can significantly improve the ability of the robotic agent to understand different ways of expressing the same answer from the user and the flexibility to generate responses to the user in different situations. The enriched training datasets can be automatically generated in a bootstrap manner. Figure 9CIllustrated are exemplary types of language variants (alternatives) for generating enriched training data for nodes with parameterized content according to embodiments of the present teachings. Language variants can be due to different spoken languages, alternative expressions..., different accents, or even in combination with different acoustic characteristics (e.g., pitch, volume, speed, etc.) specified. Regarding alternative expressions, for each text string, there can be linguistically and semantically equivalent expressions based on, for example, synonyms or slang. To enhance the capabilities of a robotic agent, it is desirable to have more alternative content associated with each parameterized node. However, manual generation can be time-consuming, tedious, and expensive, and thus efficient methods are needed to create such alternative content. The present teachings disclose methods for automatically generating an enriched training data set associated with nodes in a T-AOG based on, for example, authored content associated with each node. That is, by using the already authored content associated with each node as a basis, the disclosed methods automatically generate alternative content for the authored content. The initial authored content can then be combined with such automatically generated alternative content as a training data set for the parameterized content associated with the node.
[0163] Figure 10AFIG. 1000 is an exemplary high-level system diagram depicting an exemplary system for automatically generating enriched AOG content according to an embodiment of the present teachings and its use for training an ASR / NLU model to achieve enhanced performance. System 1000 is provided for generating an enriched training data set for different nodes in a T-AOG and then training an ASR and an NLU model based on such an enriched training data set to obtain enhanced ASR and NLU models. In this illustrated embodiment, system 1000 includes a parameterized AOG retriever 1010, a parameterized AOG training data generator 1020, an ASR model training engine 1035, and an NLU model training engine 1040. As discussed herein, an AOG is used to represent different aspects in managing human-robot conversations. Each S-AOG may include a set of nodes, each node being related to a specific topic regarding the subject matter. Each S-AOG node for a particular topic may be associated with one or more T-AOGs, each T-AOG corresponding to a conversation strategy regarding that topic and including the back-and-forth conversation between a user and a robot agent regarding that topic. As discussed herein, a T-AOG may also have multiple nodes, each node representing one or more alternative content items spoken by a party (agent or user) in the conversation. A T-AOG with multiple nodes may be created with parameterized content associated with each node. The parameterized content of the nodes may initially be populated with some authored content generated when the T-AOG is created. Such initially parameterized content may be extended or enriched. According to the present teachings, the goal is to enrich the parameterized content associated with each node to include other alternatives so that conversations can be carried out in a more flexible and enriched manner.
[0164] In operation, the initially parameterized content associated with the TAG nodes may be used as a basis for generating a set of enriched parameterized content, which is then used as training data to train the ASR and / or NLU model to enhance the ability of the robot agent to perform more flexible conversations. Figure 10B FIG. is a flowchart of an exemplary process for obtaining enhanced ASR / NLU models based on automatically generated enriched AOG content according to an embodiment of the present teachings. At 1055, the parameterized AOG retriever 1010 first selects a topic-based template (AOG), such as an S-AOG at 1055, from the storage device 1005 and then retrieves, at 1060, the T-AOGs associated with each node in the selected S-AOG, which have an initial set of parameterized content associated therewith. The parameterized content of such retrieved T-AOG nodes is then sent to the parameterized AOG training data generator 1020 to generate enriched parameterized content.
[0165] To enrich the content created for T-AOG nodes, the parameterized AOG training data generator 1020 accesses, at 1065, the content created for each T-AOG node and uses it as a basis for generating enriched content. The parameterized AOG training data generator 1020 accesses the model in 1015 to obtain known language variants and generates, at 1070, enriched training data based on the initial content created for each T-AOG node. The enriched training data thus generated is then used to create an enriched parameterized AOG, which is stored in the enriched parameterized AOG storage device 1025. At the same time, the enriched set of parameterized content serves as enriched training data and is saved in the enriched training data storage device 1030.
[0166] Based on this enriched training data, the ASR model training engine 1035 uses, at 1075, the training data 1030 bootstrapped from the initial content created based on the language variant model 1015 to train the ASR model 1045. As discussed herein, a language variant can be a model in terms of speech style (language, accent, etc.) and speech content (different ways of saying the same thing). The ASR model 1045 obtained based on the enriched training data can then be used by the ASR to recognize utterances of the user's speech content in the speech style captured in the enriched set of parameterized content. The exported ASR model is then stored, at 1080, in the storage device 1045 for future use by the ASR engine 130( Figure 1A )
[0167] Similarly, to make full use of the enriched training data, the NLU model training engine 1040 uses, at 1085, the training data 1030 bootstrapped from the initial content created based on the language variant model 1015 to train the NLU model 1050. As discussed herein, a language variant can be a model in terms of speech style (language, accent, etc.) and speech content (different ways of saying the same thing). The NLU model 1050 obtained based on the enriched training data 1030 can then be used by the NLU engine 140( Figure 1A ) to understand the meaning of the user's speech based on the ASR results from the ASR engine 130( Figure 1A )
[0168] By developing an enriched, parameterized set of content associated with nodes of a T-AOG, it generates augmented dialogue strategies. When such an enriched, parameterized set of content can be automatically generated (without human activity) and used to further enhance machine-based speech recognition and understanding, it results in a more effective and efficient human-machine dialogue. According to this teaching, in addition to enhancing the human-machine dialogue by automatically generating augmented dialogue strategies, the process can be further enhanced by exploring multi-modal context information during spoken language understanding.
[0169] Figure 11A Depicts an exemplary high-level system diagram for context-aware spoken language understanding based on ambient knowledge tracked during a conversation, according to an embodiment of this teaching. In the illustrated embodiment, spoken language understanding includes both automatic speech recognition (which recognizes the spoken words) and natural language understanding (understanding the semantics of the speech based on the spoken words). Traditionally, spoken language understanding can utilize speech context information to understand the semantics of an utterance. This traditional context information utilizes linguistic context, e.g., words / phrases spoken before or after. For example, in the sentence "Bob bought a bike and he went for a ride", "he" refers to Bob (semantics), and this semantic ambiguity is resolved based on the linguistic context.
[0170] In human-machine interaction or conversation, different types of context can be utilized and used to more deeply understand the surrounding environment to achieve more appropriate interaction. Such different types of context can include personal characteristics (related to the user) observed during the interaction / conversation, language, vision, environment, events, preferences, and the surrounding environment. According to this teaching, different types of context can be fully utilized to assist in understanding the situation encountered. For example, if a user says, "What are these things?" without linguistic context, it is difficult, if not impossible, to know what "these things" refers to unless some additional information is accessible and used to disambiguate. In this case, if the captured visual information reveals that the user is pointing at a pile of fruits on a table, then it is now possible to understand what the user means by "these things". Different types of context can facilitate ASR in recognizing the spoken words and NLU in understanding the meaning of the recognized sentence / words. To utilize different types of context, different types of sensors can be deployed at the dialogue scene to continuously monitor the surrounding environment, collect relevant sensor data, extract features, estimate characteristics related to different events and activities, determine spatial relationships between objects, and store this knowledge learned through observation. This continuously acquired multi-modal context information can then be stored in different databases in the information state 110 (e.g., dialogue history database 250, rich media dialogue context database 260, event-centered knowledge database 270) and utilized during the ASR and / or NLU processes to improve human-machine interaction / conversation.
[0171] As shown Figure 11A Figure 11A
[0172]
[0172]
[0173] When performing ASR, ambiguity may arise. For example, ambiguity may arise regarding phonemes. While acoustic information can reach the limit of disambiguation, visual information can be used to assist in this determination, for example, by identifying visemes (which correspond to phonemes) based on visual observation of the speaker's mouth movements. Based on the phonemes (or visemes, or both), the ASR engine 130 can then identify the spoken word based on a vocabulary 1120, which can specify not only the words but also the phonemic components of each word. In some cases, some words may have the same pronunciation. During a human-computer conversation, determining the exact word spoken can be helpful. To disambiguate in such situations, visual information can be used. For example, in Chinese, "he" and "she" have the same pronunciation. To identify who "he" or "she" refers to, visual information can be used to see if any visual clues can be used to disambiguate. For example, when a speaker refers to "he" or "she" in speech, they may be pointing to a person, and this information can be used to estimate who the speaker is referring to.
[0174] The ASR engine 130 outputs one or more estimated word sequences spoken by the speaker, each sequence being associated with, for example, a number of probabilities, each probability representing the likelihood of the identified word sequence. Such output can be fed to the NLU engine 140, where the semantics of the word sequences are estimated. In some embodiments, multiple output sequences (candidates) can be processed by the NLU engine 140, and the most likely understanding of the speech can be selected. In some embodiments, the ASR engine 130 can select a single word sequence with the highest probability to send to the NLU engine 140 for further processing to determine the semantics. The semantics of the word sequence can be determined based on the language model 1130, the NLU model 1050, and multimodal information from different sources (such as the conversation history 250, the conversation context 260, the event knowledge 270, and the profile 290 from the speaker).
[0175] In some embodiments of the present teachings, for human-machine conversations, the NLU model 1150 can be trained based on the created content associated with different AOGs. As discussed herein, the created content (text, acoustics, or both) associated with an AOG can be automatically enriched based on, for example, a language variant model, and then such enriched text training data can be used to train the NLU model 1150 to derive a model that can support a wider range of speech content. While traditional natural language understanding (NLU) is based on audio signals, the present teachings provide enhanced natural language understanding based on different types of context information obtained in a multi-modal domain. The aim is to enhance the ability of the NLU engine 140 to resolve ambiguities in different situations based on relevant multi-modal information. For example, if a user says, "See what's on this," there can be different ways to figure out what "this" refers to. The traditional use of speech context is not sufficient to resolve such ambiguity. According to the present teachings, additional context information from other modalities can be explored to resolve different ambiguities through language understanding that is sensitive to multi-modal context.
[0176] As Figure 11A shown, the information stored in the dialogue history 250, the rich media dialogue context 260, the event knowledge 270, and the speaker's profile 290 can be used for language understanding. Taking the above example, when a user says "See what's on this," to figure out what "this" is, the visual information captured from the dialogue scenario can be analyzed to identify clues that can be used to resolve the ambiguity. For example, the user may be standing in front of an object and pointing at it. For example, the user can be standing next to, for example, a table and pointing at, for example, a computer on the table. Based on this knowledge learned from the visual data, a representation can be generated in the rich media dialogue context (260) that reveals the user is pointing at the computer on the table near the user. Using this visual observation, the NLU engine 140 can estimate that "this" in the speech refers to the computer on the table and the user wants to know what is shown on the computer screen.
[0177] As another example, when asked what the user likes to do on his / her birthday this year, the user may respond "the same thing I did last year". This response is not clear as to what "the same thing I did last year" is. There are two ambiguities in this response from the user. First, what is the time range of last year? Second, what did the user do within the estimated time range last year? Given the language context, it can be estimated that the time range is the user's birthday last year. However, from the language context of the current conversation, it may not be possible to infer or figure out what the user means by "the things I did last year". In such a case, the rich media conversation context according to some embodiments of the present teachings can enable the NLU engine 140 to explore information stored in the event knowledge 270 about events related to the user within the time range around his / her last birthday. For example, it can be recorded that the user went to Boston for his / her birthday last year. Such information can help the NLU engine 140 disambiguate the meaning of "the same thing I did last year". The events of last year can be recorded as events or logs in, for example, the user's file storage device 290. In this way, by making full use of the rich media context, the NLU engine 140 can then understand that the user means to visit Boston again on his / her birthday this year. Thus, the understanding can then help the dialogue manager 150 (see FIG. 1) then determine how to respond to the user in the conversation. Therefore, by exploring the rich media context observed over time and in the conversation scenario, the spoken language understanding unit 1100 can enhance its spoken language understanding ability in both ASR (recognizing the spoken words) and NLU (the semantics of what is said).
[0178] Figure 11A The surrounding knowledge 270 can broadly include information in different aspects. Figure 11B Illustrated are exemplary types of surrounding knowledge to be tracked to facilitate context-aware spoken language understanding according to embodiments of the present teachings. As shown, the surrounding knowledge can include, but is not limited to, observations of the environment (e.g., whether it is sunny), events that have occurred (both in the past and in the current conversation), objects observed in the conversation scenario, and the spatial relationships between objects, …, observed activities (acoustic or visual), and / or statements made by the user. Such observations can be made via sensors of multiple modalities and can then be used to estimate or infer the user's preferences or profile, and then the dialogue manager 150 can use them to determine a response based on both the inferred user profile and the surrounding circumstances.
[0179] Figure 12AIllustrated is an example of tracking a personal profile based on statements made during a conversation or dialogue according to an embodiment of the present teachings. As shown, there is a robotic agent (duck talking agent) 1210 that is having a conversation with a user (child) 1220. During the dialogue, the robotic agent 1210 asks the user 1220 where he was born. The user answers that he was born in Chicago. This voice information is analyzed, and the voice-based information (audio) is tracked and analyzed to extract useful profiling information. For example, based on such a conversation, the user's profile is updated with an additional graph (or sub-graph) 1230 that links a node 1230-1 representing "me" (or the user) to another node 1230-2 representing "Chicago", with a link 1230-3 annotated as "born in". Although a simple example, the user profile can be gradually updated continuously based on the tracked audio information.
[0180] Figure 12BIllustrated is an example of tracking a personal profile during a vision-based conversation according to an embodiment of the present teachings. As shown, if a vision sensor captures a conversation scene as shown in 1240, which includes a boy (presumably the user participating in the conversation) and objects / fixtures in the conversation scene (e.g., a small bookshelf, a lamp on the bookshelf, another lamp on the floor, a window, a painting hanging on the wall that says "Michael Jordan"). Such a vision scene can be collected and analyzed so that relevant objects can be extracted, spatial relationships can be inferred, …, it is observed that the boy is gazing at the Michael Jordan poster. By analyzing the objects in the scene, knowledge can be extracted and some conclusions can be estimated. For example, based on this visual representation of the observed scene, an estimated conclusion might be that the boy likes Michael Jordan. Another estimated conclusion could be that the boy and the poster coexist in the conversation scene. Other different factors can also affect the estimation of what the visual representation might mean. For example, if the scene is the boy's own room, then the fact that he has such a poster on the wall can give a conclusion with a higher likelihood that he likes Michael Jordan. If the boy is in a scene in his friend's room or at school or other places, then the weight of the conclusion that the boy and the poster coexist is higher. In either case, based on the visual observation, a graph (or sub-graph) can be created representing the relationship between the observed boy and the poster, where a node 1250-1 representing the boy is linked to another node 1250-2 representing Michael Jordan by a link 1250-3 representing the relationship between the two. In this illustrated example, the link represents "likes". If its probability is higher, then the link can also be a "coexist" relationship. In some embodiments, the link 1250-3 can also be annotated with different relationships (e.g., "likes" and "coexist"), and each relationship can be associated with a probability. In this way, the monitored visual information can be used to continuously update the surrounding information. The dialogue manager 150 can use such observations to determine how to conduct the conversation. For example, if it is observed that the boy in the scene is looking at the poster and seems to be distracted from the focus of the conversation (e.g., tutoring about math), then the dialogue manager 150 can decide to ask the user "Do you like Michael Jordan?" to increase the user's engagement.
[0181] Information tracked based on sensor information in different modalities can be combined to create an integrated representation of the surrounding environment. Figure 12C Illustrated is an exemplary partial personal profile updated during a conversation based on multi-modal input information obtained from a conversation scene according to an embodiment of the present teachings. In this example, the integrated graphical representation 1260 is derived based on the tracked audio information Figure 12A in FIG. 1230 and the tracked visual information Figure 12BThe combination of FIG. 1250 in. There are other types of knowledge that can be tracked during a conversation and used to update the surrounding knowledge 270 and / or the user profile 290. In some embodiments, users can be classified into different groups based on the tracked information, and characteristics associated with each group can also be continuously established based on the tracked information from users in these groups. As will be discussed below, such group characteristics can be used by the dialogue manager to adaptively adjust the dialogue strategy to make the dialogue more effective.
[0182] Figure 12D An exemplary personal knowledge representation constructed based on conversations according to an embodiment of the present teachings is shown. In a dialogue session (or multiple dialogue sessions), information obtained from such conversations can be used to continuously build profile knowledge about the user, and then the dialogue manager 150 can rely on this profile knowledge to guide the dialogue. For example, in the user profile, the user's birthday can be established. If during the conversation the user mentions certain trips on his birthdays in different years, then such knowledge can be represented in a diagram as Figure 12D shown in. In this example, on the birthday in 2016, the user took a boat trip to Alaska, and on the same day in 2018, the user flew to Las Vegas to celebrate his birthday. This knowledge can be explicitly communicated by the user or can be inferred by the system from the conversation. For example, the user mentioned the dates of the trips without explicitly stating that these trips were to celebrate his birthday. With knowledge of the user's birthday, the system can infer that these trips were to celebrate the user's birthday.
[0183] As can be seen from Figure 12DAs can be seen, such knowledge can be represented as Figure 1270, where the user is represented as the entity "I" 1270-1, and it is linked to different knowledge fragments about the user. For example, via the link "Birthday" 1270-2, it points to the representation of the birthday 1270-3 with the specific date "10 / 2 / 1988" of the user. Based on two trips mentioned by the user, two destination representations 1270-5 and 1270-7 are provided, where 1270-7 represents Alaska and 1270-5 represents Las Vegas. To link the user to these two destinations, there are two links 1270-4 and 1270-8 that link the user to these two destinations, and each link has an annotation of the corresponding travel period on the link. Since these travel periods coincide with the user's birthday of 10 / 2, additional representations can be provided to link the user's birthday to these two trips. For this purpose, two additional links 1270-6 and 1270-9 can be provided that link the birthday representation 1270-3 to the two destination representations 1270-5 and 1270-7. On links 1270-6 and 1270-9, the year of the trip is indicated with an annotation of the calculated age of the user at the time of the trip. Through this representation, the user's birthday and some events associated with the user's birthday can be identified, that is, the user took two trips on his / her birthday, one by boat to Alaska in 2016 and the other by plane to Las Vegas in 2018 when the user was 28 years old and 30 years old respectively (calculated from the birthday and the year of the trip).
[0184] The robotic agent can later utilize the representation of certain knowledge tracked about the user to determine what to say in certain situations. For example, if one day the robotic agent is talking to the user and notices that this day is the user's birthday, then the robotic agent can greet the user with "Happy Birthday!" Additionally, to further engage the user, the robotic agent can say "You went to Las Vegas on your birthday last year. Where do you plan to go this year?" Such personalized and context-aware conversations can enhance user engagement and improve the intimacy between the user and the robotic agent.
[0185] Figure 13A An exemplary user group formed based on observations about the user to facilitate adaptive dialogue management according to an embodiment of the present teachings is illustrated. Figure 13AThe user groups in [description] are for illustration and not limitation, and any other grouping can be added and is possible. As shown, users can be classified into different groups, such as social groups (e.g., Facebook chat or interest groups), ethnic groups (e.g., Asian group or Hispanic group), age groups that can include youth groups (e.g., toddler group or teenage group) and adult groups (e.g., senior group and working professional group), gender groups..., and various possible professional groups. Each group of users can share some common characteristics that can be used to build a group profile. Such a group profile with characteristics shared by group members can be related to the dialogue plan or control. For example, users in some subgroups (e.g., Asian group) can share some common characteristics, such as an accent when speaking English. Such common characteristics of a group of users can be used, for example, in the dialogue plan. For example, when tutoring English, understanding that a user belongs to a specific ethnic group with a well-known accent can facilitate the dialogue manager 150 to plan a tutoring session using certain measures designed to address accent-related issues to form correct pronunciation. As another example, if a user is in the teenage group, then when the dialogue manager 150 expects to better engage the teenage user in the conversation, the dialogue manager 150 can explore some popular interests common to teenagers that are known.
[0186] In addition to the group profile, the dialogue system 100 can also accumulate knowledge or a profile of individual users based on immediate observations from the dialogue or information about each user from other sources. Figure 12A-12C Some examples of building some aspects of a user based on multimodal information collected during a dialogue are shown. Figure 13B An exemplary content / construction of a personalized user profile according to an embodiment of the present teachings is illustrated. The personal profile of a user can include the user's demographic information (e.g., ethnicity, age, gender, etc.), known preferences, language (e.g., mother tongue, second language, etc.), learning schedule (e.g., tutoring sessions on different topics),..., and performance (e.g., an assessment of the performance recorded related to different tutoring sessions on different topics). Some of the profile information can be declared (e.g., demographics) and some can be observed or learned via communication. For example, for information related to the user's second language (e.g., English), information about his / her proficiency, accent, etc. can be recorded and can be used, for example, by the dialogue manager 150 to plan or control the dialogue. Details related to some of the personal characteristics can also be collected and recorded. Figure 13B One example shown in [description] is a detailed characterization of the accent of the user's second language. For each second language that the user is learning, the accent can be described in terms of both acoustic features (e.g., the acoustic characteristics of the user's pronunciation) and viseme features (e.g., the visual characteristics of the user's mouth movement when speaking). Such a specific characterization of a person's accent in a language can be utilized in dialogue control.
[0187] Figure 13C Shows an exemplary visualization of the accent profile distribution related to different ethnic groups according to an embodiment of the present teachings. A set of features can be used to characterize an accent. For example, visemes can be used to represent features related to an accent, i.e., using visual features related to mouth movements to characterize an accent. An accent can also be characterized based on acoustic features. A person's accent can be represented by a vector with instantiated values of these features, and such a vector can be projected as a point into a high-dimensional space in a coordinate system. Figure 13C Shown therein is a 3D projection of accent vectors of different ethnic groups in a 3D coordinate system. In this illustration, each accent feature vector of a particular user is projected as a point, and Figure 13C the different circles in represent the boundaries of the projected accent feature vectors of users corresponding to different ethnic groups. Each circle includes one or more points therein, and the variation in the accents of different members of the same group corresponds to the extent or shape of the circle.
[0188] As Figure 13C shown therein, there are exemplary four sets of accent distributions A1 1310, A2 1320, A3 1330, and A4 1340, corresponding to the accent characterizations of English-speaking people of ethnic groups of Chinese, Japanese, American, and French. Each circle represents the distribution range of the accent profiles of group members, and the accent profile of the group can be derived based on multiple projected points (member accent vectors) representing the accent profiles of group members. For example, the group accent profile vector can be derived by, for example, averaging the member accent vectors in each dimension, or by using, for example, the centroid of each group distribution as the group profile. As can be seen in the figure, Figure 13C each circle in has a centroid (the center of the circle) representing the group accent profile, which is projected, for example, from the group accent feature vectors.
[0189] When a user participates in a human-machine conversation, the user can be observed, and multimodal information can be captured from the conversation scenario to estimate different types of information to be used for an adaptive conversation strategy. Based on the observed multimodal information, a robot agent can dynamically establish and use a user profile to adaptively control how to talk to the user. An exemplary type of information to be estimated is the user's accent. Accent information can be used to improve the quality of the conversation and the effectiveness of tutoring languages. With the estimated accent information, the robot agent can adjust its language understanding ability to determine the meaning of the user's utterance. If the robot agent is used to teach a user to learn a certain language (e.g., English), then this estimated accent information can be explored to determine a dynamic teaching plan in an adaptive manner.
[0190] As Figure 13C shown, the group profiles can be used to determine which group a new user may belong to. For example, in Figure 13CAmong them, the point 1350 may represent the accent vector of a new user. Without knowing the race declared by the new user, the distance between the point 1350 and the distributions of different racial groups can be used to estimate the race of the new user. For example, as shown in the figure, the distance between the point 1350 and the centroid of group A1 is 1360; the distance between the point 1350 and the centroid of group A2 is 1370; the distance between the point 1350 and the centroid of group A3 is 1380; and the distance between the point 1350 and the centroid of group A4 is 1390. In different embodiments, the distance can also be evaluated between the point 1350 and the nearest point of each distribution.
[0191] Once the distances to each group are determined, the new user can be estimated to belong to the racial group that has the shortest distance 1370 between the point 1350 and the centroid of that estimated racial group. In the example shown, since the vector point of the projection of the new user 1350 is closest to group A2, it can be estimated that the new user is a member of group A2 (for example, the Japanese group). If there is a high confidence in the estimation of the new user's race, then the group profile (centroid) can also be updated. In some embodiments, the estimated race and the representative accent profile of that racial group (i.e., the centroid of A2) can be used for an adaptive tutoring program. Figure 17B An example is shown in -17E, which shows how known accent information (e.g., visemes) of a user can be used to adaptively conduct a conversation to tutor a student in learning a language.
[0192] Figure 14A An exemplary high-level diagram of a system for tracking individual speech-related characteristics to update a user profile according to an embodiment of the present teachings is depicted. As discussed herein, the user profile can be dynamically updated based on information observed during a human-machine conversation to characterize the underlying user, and then the user profile can be used to guide a robotic agent to adjust the conversation with the user according to the observed characteristics of the user. The exemplary system diagram 1400 is for establishing and updating the user's accent information regarding a certain language (e.g., English). In the illustrated embodiment, the system 1400 receives audio / visual information captured when the user 1402 is speaking, analyzes the received audio / visual information to extract speech-related acoustic and visual features (viseme features), classifies the features to estimate the characteristics of the user's speech, and accordingly updates the user profile stored in 290.
[0193] The system 1400 includes an acoustic feature extractor 1410, a viseme feature extractor 1420, an acoustic feature updater 1440, and a feature-based accent group classifier 1450. Figure 14BFIG. 0 is a flowchart of an exemplary process of a system 1400 for tracking an individual's speech-related characteristics and their user profile in accordance with an embodiment of the present teachings. When a user 1402 is speaking (e.g., in a language that a robotic agent is tutoring), an audio signal corresponding to the user's utterance and a visual signal capturing the user's mouth movements during speech are obtained. When such obtained multimodal information is received at 1405 by acoustic and phoneme feature extractors 1410 and 1420, the audio signal is analyzed at 1415 by the acoustic feature extractor 1410 to estimate acoustic features associated with the speech of the user 1402. These acoustic features are identified based on, for example, an acoustic accent profile 1470, which characterizes different accents based on acoustic characteristics identified from speech signals.
[0194] At 1425, the received visual signal is analyzed by the phoneme feature extractor 1425 to estimate or determine phoneme features related to the user's mouth movements during speech. Based on various language-based phoneme models 1420, the phoneme features are identified based on, for example, visual images of the user's mouth movements during speech. To classify the user's accent into an appropriate accent group (see Figure 13A-13C ), the feature-based accent group classifier 1450 receives the acoustic features from the acoustic feature extractor 1410 and the phoneme features from the phoneme feature extractor 1420, and classifies the user's accent into an accent group at 1445 according to a language-based accent group profile stored in 1460. Once classified, the user profile in 290 associated with the user 1402 is updated accordingly to incorporate the estimated accent classification. As discussed herein, this estimated accent information can later be used, for example, by a robotic agent to adaptively plan a conversation.
[0195] In addition to the user's language characteristics, multimodal information ( Figure 2A 250 - 290 in ) that can be used to facilitate adaptive dialogue management also includes activities that occur, events that occur, and so on. Knowledge about events can assist a robotic agent in improving dialogue engagement. An event can refer to something that has happened in the past or something observed during a conversation, whether done or mentioned by someone during the conversation. For example, a user in a dialogue scene can walk towards an object in the scene, or the user can mention that he will fly to Las Vegas in June to celebrate his birthday.
[0196] Figure 15A An exemplary abstract structure representing an event in accordance with an embodiment of the present teachings is provided. As Figure 15AAs shown, an event includes an action performed on an object. The action involved in the event can be described by a verb, and the object can be anything the action is directed at. Regarding a conversation, the action can be performed by an entity in the conversation scene, and the object can be any object in the scene that can be acted upon. For example, as Figure 15A shown, the event can be someone walking towards an object (such as a blackboard, a table... or a computer); someone pointing at an object (such as a blackboard, a table or a computer on the table);..., and someone lifting an object, etc.
[0197] Figure 15B-15C illustrates an example of tracking event - centered knowledge as conversation context during a conversation based on observations from a conversation scene according to an embodiment of the present teachings. As discussed herein, knowledge about an activity or event (whether it occurs during a conversation or not) can enable a robotic agent to enhance its performance. Event - centered knowledge can be used to assist the robotic agent in understanding the meaning or intention of speech from a human user. For example, a user having a conversation with a robotic agent can walk towards a blackboard, then point at it and ask, "What is this?" This is illustrated in Figure 15B In this case, with only acoustic information, it is usually difficult, if not impossible, for the robotic agent to understand what the user is saying. However, other types of information from the conversation scene can provide useful clues and can be combined with the speech recognition result (the spoken words) to estimate what the user is referring to.
[0198] According to the present teachings, the conversation scene can be tracked, and information can be analyzed to detect various objects present in the conversation scene, the spatial relationships between such objects, the action(s) performed by the user, and the impact of such actions on the objects in the scene, including the dynamic change of the spatial relationships of the objects caused by such actions, etc. For example, as Figure 15B shown, it is observed via one or more visual sensors that the user in the scene walks towards the blackboard, raises a hand, and points at the blackboard when the user says, "What is this?" Such observations can be used to dynamically construct a [as Figure 15CThe events of the exemplary representation shown in (the user walks towards and points to the blackboard in the scene) are used to describe the events visually observed in the dialogue scene. In this example, the event includes two actions. One (Action 1) is walking towards the blackboard, and the other (Action 2) is pointing to the blackboard. Each action involved can be annotated with a time T indicating when the action is performed (see T1 for Action 1 and T2 for Action 2). In this way, the sequence of actions in the event can be associated with a timeline associated with different actions. For example, the user may first walk towards (Action 1 at T1) the blackboard and then point to (Action 2 at T2) the blackboard when the user says "What's this?". The timing of each action can be compared with the timing of the utterance, and the temporal correspondence can also assist in understanding the meaning of the utterance.
[0199] The information collected in multiple modalities facilitates tracking / learning knowledge about what is happening in the dialogue scene, which can be described in a dynamically constructed event representation. When traditional language models and context cannot effectively resolve certain ambiguities, this dynamically constructed event knowledge can play an important role in helping the robot agent estimate the semantics of the word sequence recognized (by ASR) through NLU. Figure 15B and 15C Taking the example shown in , based on the knowledge of the event that the user walks towards the blackboard and points to the blackboard while saying "What's this?", the robot agent can at least draw the following conclusions: The user's "this" refers to the blackboard, either what is on the blackboard or what the blackboard is. This understanding is more useful than what traditional methods can achieve (i.e., not knowing what the user means by "What's this?"). In some cases, if more information is available, then the robot agent may be able to further narrow down the user's intention. For example, if the user points to the computer screen and asks "What's this?" and the robot agent has just shown the user a picture, then the robot agent can determine an appropriate response to the user's question, such as asking "Do you mean the picture I just showed on the computer?". Therefore, rich media context information (including visually observed events, actions, etc.) can assist the robot agent in designing adaptive dialogue strategies to continuously engage the user in the human-robot dialogue session.
[0200] Figure 15D-15EShows another example of tracking multimodal information according to an embodiment of the present teachings to identify event - centered knowledge to facilitate spoken language understanding in a human - machine dialogue. In this example, the user in the dialogue can say "I like to try this". Although ASR can process the utterance and identify the words "I", "like", "try", and "this", without more non - traditional context information, the robot agent cannot determine what the user's "this" refers to. In this example, the visual information being simultaneously monitored in the dialogue scene can be analyzed, and the robot agent can use the events detected when the user utters the words to estimate the meaning of the utterance. In this example, as Figure 15D shown, it can be observed via visual means that the user who utters these words reaches out a hand to a laptop computer on the table (action 1), opens (action 2) the laptop computer, and then starts typing on the keyboard (action 3). Such event knowledge learned from the visual information can be used to construct an event representation as Figure 15E shown, where the event is represented as including three actions (reach out, open, and type), each action associated with a timing (T1, T2, and T3 respectively) and the object that each action is directed at. In this example, the sequence of actions in the event can be identified according to the order of the times associated with the actions. By combining such visual observations with the recognized words "I like to try this", the robot agent can infer that the user's "this" might refer to something the user likes to do on the laptop computer. As a response, the robot agent can design a response strategy by asking the user "What do you want to try on the computer?" so as to stay relevant to what the user is doing / thinking, thereby achieving better engagement with the user.
[0201] As discussed herein, a robot agent according to an embodiment of the present teachings conducts a dialogue, which aims to achieve personalized context - aware human - machine dialogue. It is personalized because the robot agent according to the present teachings utilizes user profile information and dynamic personal information updates ( Figure 12A-12C example shown) to understand and respond to (as described below) the user. It is context - aware because it utilizes event knowledge dynamically estimated from information acquired via different modalities in real - time ( Figure 15A-15E example shown) or previously established to facilitate the robot agent's understanding of the dialogue context. This personalized and context - aware dialogue control enables the robot agent to not only better understand the user but also generate more relevant, engaging, and appropriate responses for dynamic scenarios.
[0202] Figure 16ADepicts an exemplary high-level system diagram for personalized context-aware dialogue management (PCADM) according to an embodiment of the present teachings. In this embodiment, the information state 100 is centered around PCADM and includes rich media context information and personalized information established / updated based on sensor information across different modalities. As shown, PCADM includes a multimodal information processing component, a knowledge tracking / updating component, a component for estimating / updating the minds of different parties, a dialogue manager 150, and a component responsible for generating a deliverable response (a response determined by the dialogue manager 150), all of which use a personalized and context-aware manner based on various information dynamically updated in the information state 110.
[0203] In Figure 16A the illustrated embodiment shown, the multimodal information processing component includes, for example, an SLU engine 1100 (speech understanding, including both ASR 130 and an NLU engine 140, see FIG. 11), a visual information recognizer 1600, and a self-motion detector 1610. The knowledge update component may include, for example, a surrounding knowledge tracker 1620, a user profile update engine 1630, and a personalized common sense updater 1670. The component for estimating / updating the minds of the parties includes, for example, an agent mind update engine 1660, a shared mind monitoring engine 1640, and a user mind estimation engine 1650. The component for generating a deliverable response includes, for example, an NLG engine 160 and a TTS engine 170 (see FIG. 1).
[0204] The dialogue manager 150 manages the dialogue based on the speech understanding of the user's utterance from the SLU engine 1100, the relevant dialogue tree (e.g., a specific T-AOG governing the underlying dialogue), and different types of data from the information state 110 (e.g., user profile, dialogue history, rich media context, estimated different minds, event knowledge,..., common sense), which enables the dialogue manager 150 to determine a response to the user's utterance in a personalized and context-aware manner. As discussed herein, the information state 110 is dynamically established / updated by various components (such as 1620, 1630, 1640, 1650, 1660, and 1670). For example, after receiving a signal related to the user's speech (which may include, for example, audio and visual of mouth movement), the SLU engine 1100 performs speech understanding. In some embodiments, in order to understand the user's utterance, information from the information state 110 may be explored during both speech recognition (which is used to determine the spoken words) and speech understanding (which is used to understand the meaning of the utterance). For example, if the information state for the user participating in the dialogue records the user's known accent or visemes, then such information can be used to identify the spoken words. As regarding Figure 15A-15EAs discussed, the event knowledge about the user observed in the dialogue scenario can also be used to resolve certain ambiguities. This event knowledge can be derived by analyzing visual input by the visual information recognizer 1600 and the surrounding knowledge tracker 1620, and then the event representation can be stored in the information state 110 and accessed by the SLU engine 110 to understand the semantics of the user's utterance.
[0205] Visual and other types of observations (e.g., tactile input from the user) can also be monitored and analyzed to derive different context information either individually or in combination with audio information. For example, the chirping of birds (sound) and green trees (vision) can be combined to infer that it is an outdoor scene. The self-movement of the user detected via, for example, tactile information can be combined with visual information to infer changes in the spatial relationship between the user and the objects present in the scene. The facial expression of the user can also be recognized from visual information and can be used to estimate the user's mood or intention. Such an estimate can also be stored in the information state 110 and then used, for example, by the user's mind estimation engine 1650 to estimate the user's mind, or by the dialogue manager 150 to determine what to do to continue engaging the user in the dialogue.
[0206] As discussed herein, the dialogue between the robotic agent and the user can be driven by the AOG (or the agent's mind), where the AOG represents the expected topic and has an expected conversation flow with specific authored dialogue content. During the dialogue, depending on the recognized user utterance, certain dialogue paths in the AOG can be recognized by the shared mind monitoring engine 1640, and the information related to the shared mindset can be estimated and used to update the shared mindset representation recorded in the information state 110. When estimating the shared mindset, the shared mind monitoring engine 1640 can also utilize the rich media context developed based on the multimodal sensing input. In addition, based on the estimated shared mind, the user mind estimation engine 1650 can further estimate the user's mind and then update the user mind recorded in the information state 110.
[0207] Utilizing the rich media surrounding context, the semantics of the user's utterance from the SLU engine 1100, and the user profile 290, the dialogue manager 150 determines a personalized and context-aware response to the user based on the updated information state 110. This response can be determined based on the understanding of what the user said (semantics), the user's emotional or mindset state, the user's intention, the user's preferences, the dialogue strategy specified by the relevant AOG, and the estimated level of user engagement. In addition to determining the content of the response to the user, the robotic agent according to this teaching can further personalize the response by generating personalized content for the response in a manner suitable for a specific user. For example, as regarding Figure 9A-9BAs discussed, the response can come from a categorized topic that can have parameterized content (e.g., a response to an incorrect answer). Given this, specific content associated with the parameterized content can be selected for a particular user based on knowledge about the user (e.g., the user's preferences or the user's emotional state). This is implemented by the NLG engine 160. Refer to Figure 16C-16D More details related to the NLG engine 160 and the TTS engine 170 are provided.
[0208] Figure 16B is a flowchart of an exemplary process for personalized context-aware dialogue management according to an embodiment of the present teachings. When a multimodal input is received at 1605, Figure 16A different components in analyze information in various domains. For example, an acoustic signal related to the user's speech and / or ambient sound can be received; the SLU engine 1100 can analyze the audio signal to understand what the user is saying. Visual signals capturing the user's physical appearance and movement (e.g., mouth and body movement) can also be received, and the visual signals can be analyzed by the visual information recognizer 1600 to detect, for example, different objects, the user's facial features, body movement, etc. Such multimodal data analysis results from 1100 and 1600 can then be utilized by other different components to derive a higher level of understanding of the surrounding environment, preferences, and / or emotions. For example, the surrounding knowledge tracker 1620 can track, at 1625, the dynamic spatial relationships between different objects, evaluate the user's mood or intention, etc. Then, this tracked surrounding situation can be used by the user profile update engine 1630 to evaluate, for example, the user's characteristics (such as observed preferences), and update the user profile in the information state 110 at 1635 based on the observation and analysis results.
[0209] Based on the tracked surrounding information (e.g., the tracked user's movement, events, and environment), the rich media context stored in the information state 110 can be updated at 1645. The SLU engine 1100 can then utilize the updated user profile and rich media context to perform personalized context-aware spoken language understanding at 1655, which includes but is not limited to identifying the words spoken by the user (e.g., identifying based on accent information related to the user) and the semantics of the spoken words (e.g., identifying based on visual or other cues revealed in other modalities). Based on the understanding of the user's speech, the dialogue manager 150 then determines, at 1665, the response to the user that is considered appropriate given the context and the user's known characteristics. For example, the dialogue manager 150 can determine that when the user answers a question incorrectly, a response pointing this out to the user will be delivered to the user.
[0210] To ensure that such a response is personalized, the NLG engine 160 can select one of multiple alternative responses associated with a node having a specified purpose (see Figure 9A-9B ) to generate a user-specific personalized response at 1675. The selection can be based on personal information stored in the information state 110 (e.g., the user has a sensitive personality, previously answered a similar question incorrectly, and currently appears frustrated). For example, given that the user is known to be sensitive, easily frustrated, and has made mistakes repeatedly, the NLG engine 160 can generate a response that is intended to be gentle to avoid further frustrating the user. To deliver the personalized response to the user, the response can also be rendered by the TTS engine 170 in a personalized and context-aware manner. For example, if the user is known to have a southern accent and is currently in a noisy environment (e.g., as specified in the information state 110), then the TTS engine 170 can render the response with a southern accent and a higher volume at 1685. Once the response is delivered to the user, the process returns to step 1605 to process the next round of the conversation in a personalized and context-aware manner.
[0211] Figure 16C Depicts an exemplary high-level system diagram of the NLG engine 160 and the TTS engine 170 in accordance with an embodiment of the present teachings to produce context-aware and personalized audio responses. In this illustration, both the NLG engine 160 and the TTS engine 170 can adjust the response based on the tracked information stored in the information state 110. First, the NLG engine 160 can generate a response based on the text response from the dialogue manager 150, where the modification or adjustment is determined based on information related to the user and according to the known dialogue context. In addition to this, the TTS engine 170 can further adjust the (already adjusted) response in its rendered delivery form in a personalized and context-aware manner.
[0212] The NLG engine 160 includes a response initializer 1602, a response enhancer 1608, a response adjuster 1612, and an adaptive response generator 1616. Figure 16D Is a flowchart of an exemplary process of the NLG engine 160 in accordance with an embodiment of the present teachings. In operation, the response initializer 1602 first receives a text response from the dialogue manager 150 at 1632, and then initializes the response at 1634 by, for example, selecting an appropriate response from, for example, parameterized content associated with a specific node in the dialogue strategy. For example, assume that as Figure 9A shown, the dialogue strategy for a dialogue teaching the student the concept of "addition" is used to regulate a tutoring dialogue with the user. During this dialogue, a question about adding X and Y is presented to the user. When the user answers the question correctly, the dialogue manager 150 follows Figure 9AThe path in the dialogue strategy reaches node 675-2, and it is determined that the response comes from the parameterized content associated with node 675-2. As Figure 9A seen, in this case, the correct answer from the user can be because the user understood the question and answered correctly or due to a lucky guess. The dialogue manager 150 can use the parameterized content at node 675-2 as the pool from which to select a response. In this case, the parameterized content (associated with node 675-2) set for the current situation is [T RCA . As Figure 9B shown, [T RCA is defined as the combination of [T RC and [T RG , where the set [T RC is used for responses to correct answers based on a correct understanding of the taught content, while the set [T RG is used for responses to correct answers provided based on a guess. The determination of whether the user's answer is a guess or an actual correct answer can be based on the probabilities P(L t+1 |obs = correct) and P(G) for the probability estimation discussed herein. Based on these probabilities, the response initializer 1602 can use which of the two parameterized content sets as the pool from which a response can be further selected. In some embodiments, a response can be selected from the group selected based on the probabilities as described above. For example, if it is more likely that the correct answer is provided because the user truly grasped the taught concept, then the selection pool for the response is [T RC . Once selected, a response will be selected from the parameterized content set [T RC .
[0213] When generating an appropriate response, it can be selected or generated in a syntactic (based on syntactic and semantic models 1604 and 1606), knowledgeable (based on topic knowledge model 1616 and common sense model 280), humanized (based on user profile 290, the agent's mindset 200, and the user's mindset 220 estimated based on the dialogue situation), and intelligent (based on topic control according to STC-AOG 230, dialogue context 260, dialogue history 250, and event-centered knowledge 270 established based on observations of relevant conversations) manner. The output of the response initializer 1602 (e.g., the selection of a parameterized content set) can then be sent to the response enhancer 1608, which can further narrow the selection based on, for example, expertise on the topic (represented by the topic knowledge model 1614), common sense knowledge (represented by the common sense model 280), etc. Knowledge based on relevant topics and the common sense model are retrieved at 1636 and used to enhance response selection / generation at 1638. The topic knowledge model 1614 can include a knowledge graph representing human knowledge in certain topics, which can be used to control the generation of responses. The common sense model 280 can include any representation (such as subject-predicate-object triples) to model human common sense. These models can be used to ensure that the NLG engine 160 will produce meaningful sentences or sentences consistent with known facts.
[0214] To further enhance response generation, an enhanced version of the initially generated response or a further selection from a pool of alternative responses (e.g., a correct answer for a math problem can be selected from a parameterized content set for response) can then be sent to the response adjuster 1612, where the response is adjusted in a context-aware and personalized manner. For this purpose, the response adjuster can operate based on multiple considerations by accessing, at 1642, from, for example, the dialogue history (250), dialogue context (260), previously acquired event knowledge (270), any user preferences (290), and the estimated mindsets of the agent and the user (200 and 220) and adjusting the response accordingly at 1646. For example, if it is known that the user is shy and sensitive (based on the previous dialogue history), then a mild response to an incorrect answer can be generated or selected (from the existing content set). If there are some known relevant past events, e.g., the user's birthday trips in the past few years (see Figure 12D ), when the dialogue just occurs before the user's birthday, the robot agent can ask the user "Where are you going to travel for your birthday this year?" to better engage the user. As another example, if it is observed (dialogue context 260) that the user is holding a Lego toy while talking to the robot agent, then the robot agent can ask "Do you like Lego?" to have a more interesting and engaging conversation. Responses adapted to the user's familiarity or preferences may better attract the user.
[0215] As discussed herein, during a conversation, the robotic agent continuously updates its estimates of the agent's mental state 200, the shared mental state 210, and the user's mental state 220 based on what it observes. Responses to user utterances can also be adjusted based on the estimated mental states. For example, if the user correctly answers multiple questions about a mathematical concept (e.g., fractions), then the estimate of the user's mental state can indicate that the user has mastered the concept. In such a case, among the alternative responses to the last correct answer (e.g., "Well done!", "Great job!",..., or "Well done. You've mastered this concept. Do you want to move on to a new concept?"), the robotic agent can choose to confirm the user's achievement and provide a response to move forward.
[0216] The adjusted response from the response adjuster 1612 is then sent to the adaptive response generator 1616 to generate a context-aware and personalized response at 1648. To ensure that the robotic agent speaks grammatically and semantically correct sentences, the adaptive response generator 1616 generates a response according to the syntactic model 1604 and the semantic model 1606. This generated adaptive response is then sent to the TTS engine 170 at 1652 to generate a reproduction suitable for the user's preferences.
[0217] Figure 16E FIG. is a flow chart of an exemplary process for the TTS engine 170 according to an embodiment of the present teachings. When the adaptive response processor 1618 receives the adaptive response in text form from the NLG engine 160 at 1654, it processes the received response at 1656. To render the text response, the text response needs to be rendered into an acoustic form via, for example, text-to-speech conversion or TTS. To generate the audio form of the response (the utterances of the robotic agent), the adaptive TTS feature analyzer 1622 retrieves relevant information from the user profile 290 at 1658. The relevant information can include, for example, the user's age group and / or ethnic group, the user's known accent, gender, etc., which can indicate a preferred way to convert the text response into an audio (speech) form. For example, based on the known user ethnicity and the representative accent of that ethnic group (e.g., the characteristic acoustic features of the centroid of the distribution representing the average accent of that ethnic group), the text response can be converted into a speech signal with the distinctive acoustic features of that ethnic group.
[0218] If at 1662 it is determined that no such preference-related information about the user is available, then the adaptive TTS feature analyzer 1622 invokes the text-to-speech converter 1624 to convert the adaptive text response into a voice form at 1668, for example, based on a standard TTS configuration stored in, for example, the TTS feature configuration 1626. If any preferences exist, then the adaptive TTS feature analyzer 1622 analyzes the information from the user profile at 1664 to identify specific preferences and retrieves specific TTS conversion configuration parameters from the TTS feature configuration 1626 at 1666 to convert the text response into a voice form that exhibits the specific voice characteristics of the user's preferences. Based on this retrieved rendering parameter, the text-to-speech converter 1624 converts the received adaptive text response into an audio signal in the voice form of the adaptive text response representing the style specified by the user's preferences at 1668. Then, at 1672, the generated audio signal for the voice form of the adaptive text response is sent to the robotic agent for rendering (or for responding to the user).
[0219] As disclosed herein, by leveraging adaptively tracked multimodal ambient information, spoken language understanding in a conversation can be personalized in a context-aware manner, and the conversation itself or the way the conversation is conducted (the back-and-forth exchange between a machine and a human) can be adaptively configured based on the dynamics of the conversation, and these exchanges can be delivered through personalized and content-sensitive determined style choices. For example, a robotic agent for tutoring a student in a second language (e.g., English) can conduct a course based on an understanding of the user. If the student user belongs to a known ethnic group that is generally known to have a specific accent profile, then the tutoring can be conducted taking this into account, and a lesson plan can be generated for that specific accent profile to specifically overcome that accent to ensure that the student will develop the correct pronunciation.
[0220] As discussed herein, by tracking different types of information related to the context of a conversation based on multimodal information, responses to the conversation user are designed in a personalized and context-aware manner. Another consideration in deriving a response in a conversation is the robotic agent's assessment of the user's performance. Such an assessment can not only guide how to conduct the conversation with the user but can also be used to adaptively adjust the teaching plan during the conversation. Figure 17AFIG. depicts an exemplary high-level diagram of a machine tutoring system 1700 for adaptive personalized tutoring via dynamic feedback in accordance with an embodiment of the present teachings. The illustrated machine tutoring system 1700 includes a tutor 1710 supported by a lesson plan execution unit 1770, a communication understanding unit 1720 for understanding what the student user 1702 says, a grader unit 1730 for evaluating the performance of the student user 1702, and a lesson plan adjustment unit 1750. A lesson plan for the student user can be designed based on knowledge about the user from the user profile 290. For example, if it is known that the student belongs to a certain ethnic group that may have a corresponding accent profile regarding the language to be tutored, then the lesson plan can be designed based on the known accent profile to teach the student the language.
[0221] The four units in the system 1770 form a feedback loop, making it possible to continuously adjust the lesson plan during the tutoring process. After determining an initial lesson plan 1760 based on the curriculum 1740 and the user profile 290, during the tutoring process, the communication understanding unit 1720 analyzes the student's communication, and the results are sent to the grader unit 1730 to evaluate the user's performance. Such an evaluation can be carried out with respect to the curriculum 1740 that specifies the expected performance, and the expected performance can be evaluated according to the expected grades specified in the curriculum 1740. For example, when tutoring a student in English, there can be a series of tutoring sessions, some for pronunciation, some for vocabulary, some for grammar, some for reading aloud, and some for composition. Each tutoring session can target a specific goal according to a series of goals related to mastering a certain level of English. For each session, there can be specific content covered by the goals to be achieved, so that the grader unit 1730 can evaluate the user's performance based on the specific goals to be achieved. For example, when teaching a student how to read aloud with correct pronunciation, the content can target different words with different phonemes for reading aloud. The acoustic signals of the words spoken by the user and the visual information about the mouth movements when speaking the words can be recorded, and the acoustic and visual characteristics of the user reading these words can be analyzed and compared with the standard corresponding acoustic and visual profiles as part of the evaluation.
[0222] Then, the evaluation results are used by the lesson plan adjustment unit 1750 to determine whether the initial lesson plan needs to be adjusted, and if so, how to change the lesson plan according to the student's performance and the expected goals of the curriculum 1740. The adjustment can be based on the deviation between the observed acoustic / visual characteristics and the standard acoustic / visual profiles. The adjustment to the lesson plan is then used to update the lesson plan 1760, so that the revised teaching activities can be executed by the lesson plan execution unit 1770. At the same time, the observed performance and its evaluation can also be sent to the user profile updater 1780 to continuously update the user profile 290. As described hereinFigure 17A The different units in Figure 16A-16B can be part of personalized and context-aware dialogue management as shown in Figure 17A is to show the feedback nature of the closed-loop system 1700 in dynamically adjusting the teaching plan during a dialogue session.
[0223] An example of adjusting a teaching plan based on a user's profile is to have a personalized teaching plan to teach each student the correct pronunciation in a language. As is known, pronunciation can be measured by both acoustic (phonemes) and visual (visemes). As referred to herein Figure 13B-14B discussed, accent profiles of different ethnic groups and individual users can be established via audio and video information. For an individual user, an accent profile can be established based on audio / video information about how the user reads certain language materials. For an ethnic group, the accent profile of the group can be designed based on the accent profiles of its members. The deviation between the user's accent profile of a language and the representative accent profile of the group speaking that language can be used to design a teaching plan to correct the user's accent.
[0224] Figure 17B illustrates an exemplary method by which a robotic agent tutor according to an embodiment of the present teachings can be used to tutor a student. As shown, the robotic tutor can design different ways to tutor a student in learning a language like a human. For example, the tutor can tutor the student via acoustic, visual, or text means. Acoustically, the tutor can explain to the student what is the correct and incorrect way of pronouncing words or phonemes, and can acoustically show the comparison between the correct and incorrect ways of pronunciation. The tutor can also give a visual explanation to the student in terms of the movement of the lips / mouth during pronunciation. In some cases, the tutor can also show an audio track animation to the student while producing the sound so that the student can follow to pronounce correctly. The tutor can also provide text passages for the student to read aloud so that the student can pronounce the words correctly according to the explanation.
[0225] In human - interaction - based tutoring, a robotic tutor can rely on dynamic observations of a student's performance (either auditory or visual) in order to selectively choose a way to tutor the student. This selection of teaching methods can include either the material to be taught based on the student's progress or the way to teach the student. For example, a grader unit 1730 can dynamically evaluate a student's performance based on observations related to multiple aspects, such as whether the student's answers are correct, whether the student's pronunciation has any accent, and whether the student's visemes match the required visemes. If the evaluation by the grader unit 1730 reveals that the student's visemes do not match the visemes required for the particular language being taught, then the robotic agent can dynamically decide how the correct visemes, or even visual animations of visemes, or sound tracks should be shown to the student in order to correct the student's pronunciation.
[0226] Figure 17C Exemplary aspects of user performance that a grader unit 1730 according to embodiments of the present teachings can evaluate during a tutoring session are provided. As shown, the grader unit 1730 can be designed to evaluate a student according to different aspects of language learning, such as language features (such as syntactic performance (whether the student uses correct syntax), semantics (whether the student understands the semantics of words, sentences)), pronunciation - related features (such as how the student pronounces, what visemes are observed, etc.), the student's reading fluency, the student's overall comprehension, the student's language use, and various other observations of the student that can be related to determining the appropriate way to teach the student (such as the student's gender, age group, or whether the language taught to the student is his / her first language or second language). A teaching plan adjustment unit 1750 can make and use these observations to facilitate its decision on how to adjust the teaching plan, including the material to be taught and the way to teach the student (auditory, visual, text). Figure 17D Examples related to teaching the English language are shown in FIGS. 17A - 17F.
[0227] Figure 17D An exemplary projection of the spoken - language profile distribution 1330 of American English speakers and exemplary acoustic waveforms of different phonemes for the centroid point of the distribution 1330 and the corresponding visual visemes are provided. As Figure 17DAs shown, on the left is the distribution 1330 of a set of accent profiles of Americans projected in a coordinate system, and on the right are the corresponding acoustic waveforms of different phonemes in American English and their corresponding visemes derived from the centroid (e.g., average value) of the distribution 1330. When teaching students American English, the goal is for the students to achieve an accent profile preferably within the range of 1330. Accordingly, to meet this goal, the tutoring program is to help the students achieve acoustic characteristics similar to the illustrated acoustic waveforms and the corresponding mouth shapes / movements shown for different phonemes when speaking English. This help can be provided via sound, i.e., the robotic agent speaks a phoneme or a word to the user and asks the student to imitate the sound. Additionally or alternatively, when pronouncing a phoneme or a word, the robotic agent can visually show the student the shape and movement of the mouth. Both the acoustic and visual means deployed to teach the students to achieve correct pronunciation can be based on the standard accent profile of the underlying language, e.g., the spoken profile of the centroid (average pronunciation) of the distribution 1330 for American English.
[0228] Figure 17E An example of the deviation between the accent profile of an ethnic group A4 and the accent profile of the language (A3) to be taught according to an embodiment of the present teachings is shown. In Figure 17E On the left, there are two distributions. One distribution corresponds to the distribution 1330 related to the standard accent profile distribution A3 1330 of American English, and the other distribution is the accent profile distribution A4 1340 of French people speaking American English. On the right, exemplary visemes from two ethnic groups are shown. For example, two exemplary visemes 1780-1 and 1790-1 from the standard accent profile corresponding to the centroid of group A3 and two exemplary corresponding visemes 1780-2 and 1790-2 observed from the users of the ethnic group A4 1340. As can be seen, there are observable differences between 1780-1 and 1780-2 and between 1790-1 and 1790-2. When there is a difference between the viseme from the user and the viseme from the standard accent profile of the language the user is learning, this can indicate incorrect pronunciation. Therefore, a teaching plan can be designed to incorporate accent correction measures / steps by teaching the students how to control the mouth shape for correct pronunciation. In some embodiments, the differences in the speech signals of various phonemes can also be used to determine whether oral correction is needed and, if so, what to incorporate in the teaching plan to make it happen. Thus, such observed differences in phonemes or visemes provide a basis for developing an adaptive teaching plan.
[0229] [[ID=Shows an example of tutoring content incorporated in a teaching plan according to an embodiment of the present teachings, the teaching plan being adaptively developed based on the visual feature of a user observed with respect to a standard viseme of an underlying spoken language. In this example, the teaching content 1790-3 is developed based on a first deviation of 1795 between the standard viseme 1790-1 and 1790-2 observed from the user, and the teaching content 1790-3 has an interface where the user can first see his / her own viseme and is prompted to, for example, click a "Create corrected shape" button to display the correct viseme (mouth shape) of the phoneme. In some embodiments, the robotic agent can also provide some accompanying verbal instructions to guide the student to correctly pronounce the phoneme while viewing the standard viseme. Such instructions can also be created adaptively based on, for example, the difference between 1790-1 and 1790-2. For example, if the user's viseme looks too wide instead of being rounder as in the standard viseme, the instruction can be designed to tell the user that his / her mouth needs to form a circle.
[0230] Is a flowchart of an exemplary process for adaptively creating a personalized tutoring plan via dynamic information tracking and performance feedback according to an embodiment of the present teachings. First, the multi-modal input of the user is received at 1705. To evaluate the user's performance, the current teaching plan, which is developed based on the curriculum, is accessed at 1715. Based on the user's input and the teaching plan, the user's performance is evaluated with respect to the expected goals of the curriculum at 1725, and the difference between the user's performance and the expected goals of the curriculum is identified at 1735. This difference can be identified in different modalities (e.g., identified in acoustic features and visual features), and then can be used to modify the current teaching plan at 1745 to derive a modified adaptive teaching plan. In some embodiments, the observations made from the user and the evaluation of the user's performance can be used to update the user profile at 1755. To continue the conversation with the user, the robotic agent continues the conversation at 1765 according to the modified teaching plan.
[0231] It is an illustrative diagram of an exemplary mobile device architecture that can be used to implement a dedicated system for implementing the present teachings according to various embodiments. In this example, the user equipment on which the present teachings are implemented corresponds to the mobile device 1800, including but not limited to smart phones, tablet computers, music players, handheld game consoles, global positioning system (GPS) receivers, and wearable computing devices (e.g., glasses, wristwatches, etc.), or any other form of device. The mobile device 1800 may include one or more central processing units (“CPUs”) 1840, one or more graphics processing units (“GPUs”) 1830, a display 1820, a memory 1860, a communication platform 1810 (such as a wireless communication module), a storage device 1890, and one or more input / output (I / O) devices 1840. Any other suitable components, including but not limited to a system bus or controller (not shown), may also be included in the mobile device 1800. As shown, the mobile operating system 1870 (e.g., iOS, Android, Windows Phone, etc.) and one or more applications 1880 may be loaded from the storage device 1890 into the memory 1860 for execution by the CPU 1840. The application 1880 may include a browser or any other suitable mobile application for managing the conversation system on the mobile device 1300. User interaction may be implemented via the I / O device 1840 and provided to the automated conversation partner via the network.
[0232] To implement the various modules, units, and their functions described in the present disclosure, a computer hardware platform may be used as the (one or more) hardware platform for one or more of the elements described herein. The hardware elements, operating systems, and programming languages of such computers are conventional in nature, and it is assumed that those skilled in the art are sufficiently familiar with these technologies to adapt those technologies to the appropriate settings as described herein. A computer with user interface elements may be used to implement a personal computer (PC) or other type of workstation or terminal device, but if appropriately programmed, the computer may also act as a server. It is believed that those skilled in the art are familiar with the structure, programming, and general operation of such computer equipment, and thus the drawings should be self-explanatory.
[0233] It is an illustrative diagram of an exemplary computing device architecture that can be used to implement a dedicated system for practicing the present teachings. Such a dedicated system incorporating the present teachings has a functional block diagram of a hardware platform that includes user interface elements. The computer can be a general-purpose computer or a dedicated computer. Both can be used to implement the dedicated system of the present teachings. This computer 1900 can be used to implement any component of a conversation or dialogue management system, as described herein. For example, the conversation management system can be implemented on a computer such as computer 1900 via its hardware, software program, firmware, or a combination thereof. Although only one such computer is shown for convenience, computer functions related to the conversation management system described herein can be implemented in a distributed manner on multiple similar platforms to distribute the processing load.
[0234] For example, computer 1900 includes a COM port 1950 that is connected to a network to which it is attached to facilitate data communication. Computer 1900 also includes one or more central processing units (CPUs) 1920 in the form of processors for executing program instructions. Exemplary computer platforms include an internal communication bus 1910, various forms of program storage devices and data storage devices (e.g., disk 1970, read-only memory (ROM) 1930, or random access memory (RAM) 1940) for various data files processed and / or transmitted by computer 1900, and program instructions that may be executed by CPU 1920. Computer 1400 also includes I / O components 1960 to support the input / output stream between the computer and other components therein, such as user interface elements 1980. Computer 1900 can also receive programming and data via network communication.
[0235] Thus, as outlined above, aspects of the method of dialogue management and / or other processes can be implemented in programming. The program aspects of the technology can be considered a "product" or "article of manufacture", typically in the form of executable code and / or associated data carried or embodied in a machine-readable medium. Tangible non-transitory "storage device" type media include any or all memory or other storage devices for a computer, processor, etc., or associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which can provide storage for software programming at any time.
[0236] All or part of software can sometimes be transmitted through a network such as the Internet or various other telecommunications networks. For example, such communication can enable software to be loaded from one computer or processor into another, for example, related to conversation management. Thus, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical landline networks, and through various air links. Physical elements that carry such waves (such as wired or wireless links, optical links, etc.) can also be considered as media that carry software. As used herein, unless limited to tangible "storage" media, terms such as computer or machine "readable media" refer to any medium that participates in providing instructions to a processor for execution.
[0237] Accordingly, machine-readable media can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media includes, for example, optical discs or magnetic disks, such as any storage device in any computer, etc., which can be used to implement a system or any of its components, as shown in the figure. Volatile storage media includes dynamic memory, such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables; copper wires and optical fibers, including the wires that form a bus within a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example: floppy disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched cards, paper tapes, any other physical storage media with a pattern of holes, RAM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, carriers that transport data or instructions, cables or links that transport such carriers, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media can involve transporting one or more sequences of one or more instructions to a physical processor for execution.
[0238] Those skilled in the art will recognize that this teaching is subject to various modifications and / or enhancements. For example, although the implementations of the various components described above can be implemented in hardware devices, it can also be implemented as a pure software solution - for example, installed on an existing server. In addition, the fraud network detection technology disclosed herein can be implemented as firmware, a firmware / software combination, a firmware / hardware combination, or a hardware / firmware / software combination.
[0239] Although the foregoing has described what is considered to constitute this teaching and / or other examples, it should be understood that various modifications can be made thereto and the subject matter disclosed herein can be implemented in various forms and examples, and that the teaching can be applied to many applications, only some of which are described herein. The following claims are intended to claim any and all applications, modifications, and variations that fall within the true scope of this teaching.
Claims
1. A method implemented on at least one machine, the at least one machine including at least one processor, a memory, and a communication platform capable of connecting to a network for human-machine dialogue, the method comprising the following steps: Receiving from a user participating in the human-machine dialogue an utterance regarding a topic in the dialogue scenario; Obtaining multi-modal ambient information related to the human-machine dialogue; Analyzing the multi-modal ambient information to track the multi-modal context of the human-machine dialogue; And Based on the tracked multi-modal context, personalizing the spoken understanding of the utterance in a context-aware manner to determine the semantics of the utterance; Wherein the multi-modal context includes at least one of the following: Objects in the dialogue scenario and their spatial relationships; At least one event related to the user that occurred in the past and / or was observed during the dialogue; One or more acoustic / visual activities observed in the dialogue scenario; Information related to previously recorded known characteristics of the user and / or characteristics of the user observed in the dialogue; and Common sense knowledge; Further comprising: Determining a response to the utterance based on a dialogue strategy governing the human-machine dialogue regarding the topic; Generating a personalized text response based on the response according to the multi-modal context; and Generating a personalized acoustic response corresponding to the personalized text response in a context-aware manner based on the multi-modal context; wherein, relevant information is retrieved from a user profile by an adaptive TTS analyzer, and the relevant information indicates converting the text response into a personalized acoustic response; the relevant information includes: the age group and / or ethnic group of the user, the known accent of the user, and gender.
2. The method according to claim 1, wherein The multi-modal ambient information includes acoustic and visual information.
3. The method according to claim 1, wherein The characteristics of the user observed in the dialogue include at least one of the following: Observations of the user regarding at least one of behavior, expression, and action; and Inferences about the emotion and / or intention of the user based on the observations.
4. The method according to claim 1, wherein The step of personalizing the spoken understanding of the utterance includes: Identifying, by automatic speech recognition ASR, each word spoken via the utterance, wherein the ASR disambiguates the words spoken by the user based on the characteristics of the user represented in the multi-modal context; and Determining the semantics of the utterance by natural language understanding NLU, wherein the NLU determines the semantics based on the acoustic / visual activities observed in the dialogue scenario and represented in the multi-modal context.
5. The method according to claim 1, wherein: The personalized text response is selected from a parameterized content set associated with the response; and The personalized text response is selected by: Being selected based on the characteristics of the user represented in the multi-modal context, the characteristics of the user being previously known and / or currently observed in the dialogue scenario, and Being selected in a context-aware manner based on the context information represented in the multi-modal context.
6. The method according to claim 1, wherein, The personalized acoustic response is rendered in a manner that is personalized and context-aware with respect to the user based on context information represented in the multimodal context.
7. The method according to claim 1, further comprising, when the conversation corresponds to a tutoring session regarding the topic according to the current tutoring plan, evaluating the performance of the user with respect to one or more aspects of the current tutoring plan based on the semantics of the utterance; adjusting the current tutoring plan in a context-aware manner based on the performance according to the multimodal context to generate a personalized tutoring plan; and applying the personalized tutoring plan in the conversation to continue the tutoring session regarding the topic.
8. The method according to claim 5, further comprising: Updating the dialogue strategy based on the tracked multimodal context such that the dialogue strategy is personalized and context-aware.
9. A system for conducting a human-machine conversation, comprising: A surrounding knowledge tracker configured to: obtain multimodal surrounding information related to the human-machine conversation, and analyze the multimodal surrounding information to track the multimodal context of the human-machine conversation; and A spoken language understanding engine configured to: receive an utterance about a topic in a conversation scenario from a user participating in the human-machine conversation, and personalize the spoken language understanding SLU of the utterance in a context-aware manner based on the tracked multimodal context to determine the semantics of the utterance; wherein the multimodal context includes at least one of the following: objects in the conversation scenario and their spatial relationships; at least one event related to the user that occurred in the past and / or was observed during the conversation; one or more acoustic / visual activities observed in the conversation scenario; information related to previously recorded known characteristics of the user and / or characteristics of the user observed in the conversation; and common sense knowledge; further comprising: A dialogue manager configured to determine a response to the utterance based on a dialogue strategy governing the human-machine conversation regarding the topic; A natural language generation NLG engine configured to generate a personalized text response based on the response according to the multimodal context; and A text-to-speech TTS engine configured to generate a personalized acoustic response corresponding to the personalized text response in a context-aware manner based on the multimodal context; wherein relevant information is retrieved from the user profile by an adaptive TTS analyzer, and the relevant information indicates converting the text response into a personalized acoustic response; the relevant information includes: the age group and / or ethnic group of the user, the known accent of the user, gender.
10. The system according to claim 9, wherein, The multimodal surrounding information includes acoustic and visual information.
11. The system according to claim 10, wherein, The characteristics of the user observed in the conversation include at least one of the following: observations of the user regarding at least one of behavior, expression, and action; and inferences about the emotion and / or intention of the user based on the observations.
12. The system according to claim 9, wherein, The spoken language understanding engine includes: An automatic speech recognition (ASR) engine configured to recognize individual words spoken via the utterance, wherein the ASR engine disambiguates the words spoken by the user based on characteristics of the user represented in the multimodal context; and A natural language understanding (NLU) engine configured to determine the semantics of the utterance, wherein the NLU engine determines the semantics based on acoustic / visual activities observed in the conversation scenario and represented in the multimodal context.
13. The system of claim 9, wherein: The personalized text response is selected from a parameterized content set associated with the response; and The selection is made by: Based on characteristics of the user represented in the multimodal context, the characteristics of the user being previously known and / or currently observed in the conversation scenario, and Based on context information represented in the multimodal context in a context-aware manner.
14. The system according to claim 9, wherein, The personalized acoustic response is rendered in a user-personalized and context-aware manner according to context information represented in the multimodal context.
15. The system of claim 9, when the conversation corresponds to a tutoring session regarding the topic according to a current tutoring plan, the system further includes: A scorer unit configured to evaluate the performance of the user with respect to one or more aspects of the current tutoring plan based on the semantics of the utterance; A tutoring plan adjustment unit configured to adjust the current tutoring plan in a context-aware manner based on the performance according to the multimodal context to generate a personalized tutoring plan; And A tutoring plan execution unit configured to apply the personalized tutoring plan in the conversation to continue the tutoring session regarding the topic.
16. The system of claim 13, further including an agent mind update engine configured to update the conversation strategy based on the tracked multimodal context such that the conversation strategy is personalized and context-aware.
Citation Information
Patent Citations
Methods and computer-program products for teaching a topic to a user
US20130029308A1
Context-based natural language processing
US20160259775A1
Real-time human-machine collaboration using big data driven augmented reality technologies
US20160378861A1