Systems and methods for adaptive dialogue management across real and augmented reality
By integrating artificial intelligence into the proxy device, using multimodal data to monitor the dialogue environment, and adaptively adjusting the conversation strategy, the problem that traditional dialogue systems cannot be dynamically adjusted is solved, and more attractive and efficient human-computer dialogue is achieved.
Patent Information
- Application Number
- CN202080053887.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-03
- Filing Date
- 2020-07-03
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-07-03
AI Technical Summary
Traditional computer-assisted dialogue systems cannot dynamically adjust conversation strategies, cannot effectively attract human conversation users, and do not consider emotional factors and conversational context, resulting in inefficiency in dialogue.
By incorporating artificial intelligence into the proxy device, leveraging multimodal data to monitor the conversation environment, adaptively adjust conversation strategies, including conversation topics, hardware configurations, and emoticons/behaviors, to personalize and dynamically adjust conversation strategies.
It realizes more attractive, real and efficient human-computer dialogue, which can flexibly adjust according to the reactions of human conversationalists and improve the conversation experience.
Smart Images

Figure CN114287030B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application 62 / 870,162, filed on July 3, 2019; U.S. Provisional Patent Application 62 / 870,168, filed on July 3, 2019; U.S. Provisional Patent Application 62 / 870,201, filed on July 3, 2019; U.S. Provisional Patent Application 62 / 870,174, filed on July 3, 2019; U.S. Provisional Patent Application 62 / 870,211, filed on July 3, 2019; U.S. Provisional Patent Application 62 / 870,217, filed on July 3, 2019; U.S. Provisional Patent Application 62 / 870,224, filed on July 3, 2019, each of which is hereby incorporated by reference in its entirety. Technical Field
[0003] The present disclosure generally relates to computers. More specifically, the present disclosure relates to computerized intelligent agents. Background Art
[0004] With the advancement of artificial intelligence technology and the proliferation of Internet - based communications due to ubiquitous Internet connectivity, computer - assisted dialogue systems have become increasingly popular. For example, more and more call centers deploy automated dialogue robots to handle customer calls. Various hotels install various self - service kiosks that can answer questions from tourists or guests. Online reservations (whether for travel accommodation or theater tickets, etc.) are also more frequently completed by chatbots. In recent years, automated human - machine communication in other fields has also become increasingly popular.
[0005] Such traditional computer - assisted dialogue systems typically pre - write certain questions and answers based on common conversation patterns in different fields. Unfortunately, human conversationalists can be unpredictable and sometimes do not follow the pre - planned dialogue patterns. Additionally, in some cases, human conversationalists may digress during the process, and continuing with a fixed conversation pattern may lead to irritation or loss of interest. When this happens, such traditional machine dialogue systems will no longer be able to attract human conversationalists, making it necessary either to abort the human - machine dialogue and hand it over to a human operator or for the human conversationalist to simply leave the dialogue, which is undesirable.
[0006] In addition, traditional machine-based conversation systems are generally not designed to take into account human emotional factors, let alone consider how to handle these emotional factors during a conversation with a person. For example, traditional machine conversation systems typically do not initiate a conversation unless someone activates the system or asks some questions. Even when a traditional conversation system does initiate a conversation, it has a fixed way of starting the conversation and does not vary from person to person or adjust based on observations. Thus, although they are programmed to faithfully follow a pre-designed conversation pattern, they generally cannot respond to and adjust to the dynamics of the conversation in a way that keeps the conversation engaging. In many cases, when the person participating in the conversation is clearly irritated or frustrated, traditional machine conversation systems are typically completely insensitive to this and instead continue to push the conversation forward in the same way, causing annoyance and undermining the person's interest. This not only results in the conversation ending unpleasantly (without the machine ever realizing it), but also makes the person less likely to engage in conversations with any machine-based conversation system in the future.
[0007] In some applications, it is crucial to conduct human-machine conversations based on what is observed from a person in order to determine how to conduct the conversation effectively. An example is a conversation related to education. When a chatbot is used to teach a child to read, it is necessary to continuously monitor and consider whether the child can perceive the way he / she is being taught in order for it to be effective. However, traditional systems cannot address such issues. Another limitation of traditional conversation systems is that they do not understand the context of the conversation in different dimensions. For example, traditional conversation systems do not have the ability to observe the context of the conversation and improvise conversation strategies to engage the user and improve the user experience.
[0008] Therefore, methods and systems are needed to address these limitations. SUMMARY OF THE INVENTION
[0009] The teachings disclosed herein relate to methods, systems, and programming for advertising. More specifically, the present disclosure relates to methods, systems, and programming related to exploring advertising sources and their utilization.
[0010] In one example, a method for managing a user machine conversation implemented on a machine having at least one processor, a memory, and a communication platform capable of connecting to a network. Information related to a user machine conversation in a conversation scenario involving a user is received, and the user machine conversation is managed by a conversation manager according to an initial conversation strategy. Based on this information, the initial conversation strategy is adapted to generate an updated conversation strategy, and based on the updated conversation strategy, it is determined whether the user machine conversation is to continue in an enhanced conversation reality having a virtual scenario presented in the conversation scenario. If so, a virtual agent manager is activated to create the enhanced conversation reality and manage the user machine conversation therein.
[0011] In various examples, a system for managing user-machine conversations includes a dialogue manager and an information state updater. The dialogue manager is configured to receive information related to a user-machine conversation in a conversation scenario involving a user, where the user-machine conversation is managed according to an initial dialogue strategy. The information updater is configured to update the information state based on the information to facilitate adjustment of the initial dialogue strategy stored in the information state, thereby generating an updated dialogue strategy. Based on the information state, the dialogue manager is further configured to determine, based on the updated dialogue strategy, whether the user-machine conversation is to continue in an enhanced dialogue reality having a virtual scenario presented in the conversation scenario, and if so, activate a virtual agent manager to create the enhanced dialogue reality and manage the user-machine conversation therein.
[0012] Other concepts relate to software for implementing the present disclosure. A software product according to this concept includes at least one machine-readable non-transitory medium and information carried by the medium. The information carried by the medium can be executable program code data, parameters associated with the executable program code, and / or information related to a user, a request, content, or other additional information.
[0013] In one example, a machine-readable non-transitory tangible medium having recorded thereon data for user-machine conversations, where when read by a machine, the medium causes the machine to perform a series of steps to implement a method for managing user-machine conversations.
[0014] Additional advantages and novel features will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the following description and the accompanying drawings, or may be learned by the production or operation of examples. The advantages of the present disclosure may be realized and obtained by means of the various aspects of the methods, means, and combinations set forth in the detailed examples discussed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The methods, systems, and / or programming described herein will be further described by way of exemplary embodiments. These exemplary embodiments will be described in detail with reference to the accompanying drawings. These embodiments are non-limiting exemplary embodiments, where in several views of the drawings, like reference numerals denote similar structures, and in the drawings:
[0016] Figure 1 depicts a networked environment for facilitating a conversation between a user of a user device and an agent device in combination with a user interaction engine according to an embodiment of the present disclosure;
[0017] Figure 2A - 2B depicts the connections between a user device, an agent device, and a user interaction engine during a conversation according to an embodiment of the present disclosure;
[0018] Figure 3A Illustrate an exemplary structure of an agent device having an exemplary type of agent body according to an embodiment of the present disclosure;
[0019] Figure 3B Illustrate an exemplary agent device according to an embodiment of the present disclosure;
[0020] Figure 4A Depict an exemplary high-level system diagram of an entire system for an automated companion according to various embodiments of the present disclosure;
[0021] Figure 4B Illustrate a part of a dialogue tree of an ongoing dialogue according to an embodiment of the present disclosure, and the path taken based on the interaction between the automated companion and the user;
[0022] Figure 5 Illustrate an exemplary multi-layer processing of an automated dialogue companion and the communication between different processing layers according to an embodiment of the present disclosure;
[0023] Figure 6A Depict an exemplary constitution of a dialogue system centered on an information state for capturing dynamic information observed during a dialogue according to an embodiment of the present disclosure;
[0024] Figure 6B Is a flowchart of an exemplary process of a dialogue system using an information state for capturing dynamic information observed during a dialogue according to an embodiment of the present disclosure;
[0025] Figure 7A Depict an exemplary structure of an information state according to an embodiment of the present disclosure;
[0026] Figure 7B Illustrate how different minds are connected in a dialogue with a robot teacher teaching a user fraction addition according to an embodiment of the present disclosure;
[0027] Figure 7C Indicate an exemplary relationship between the mind of an agent, the shared mind, and the mind of a user represented in an information state according to an embodiment of the present disclosure;
[0028] Figure 8A Illustrate an exemplary dialogue scenario according to an embodiment of the present disclosure, where actual and virtual reality can be combined to implement an adaptive dialogue strategy;
[0029] Figure 8B Illustrate an exemplary aspect of operations for implementing an adaptive dialogue strategy according to an embodiment of the present disclosure;
[0030] Figure 9AExemplary high-level system diagram depicting a system for adaptive environment modeling and robot / sensor collaboration according to an embodiment of the present disclosure;
[0031] Figure 9B Flowchart of an exemplary process of a system for adaptive environment modeling and robot / sensor collaboration according to an embodiment of the present disclosure;
[0032] Figure 10A Describe exemplary hierarchical dialogue agent collaboration in enhanced dialogue reality according to an embodiment of the present disclosure;
[0033] Figure 10B Represents an exemplary composition of a virtual agent to be deployed in enhanced dialogue reality according to an embodiment of the present disclosure;
[0034] Figure 11A Represents an exemplary dialogue scenario with a robot agent and a user according to an embodiment of the present disclosure;
[0035] Figure 11B Represents an enhanced dialogue reality scenario according to an embodiment of the present disclosure, where a robot agent creates an avatar designed to have a specified conversation with a user;
[0036] Figure 11C Represents an enhanced dialogue reality scenario according to an embodiment of the present disclosure, where a robot agent creates an avatar for a specified conversation with a user and a virtual companion for providing company to the user;
[0037] Figure 11D Represents an enhanced dialogue reality scenario according to an embodiment of the present disclosure, where a robot agent creates an avatar for a specified conversation, and the avatar further creates a virtual companion for providing company to the user;
[0038] Figure 11E Represents a robot agent that sequentially creates different avatars responsible for specified tasks according to an embodiment of the present disclosure;
[0039] Figure 12A Exemplary high-level system diagram depicting a virtual agent manager that jointly manages conversations in enhanced dialogue reality with a dialogue manager according to an embodiment of the present disclosure;
[0040] Figure 12B Flowchart of an exemplary process of collaborative management of conversations between a dialogue manager and a virtual agent manager according to an embodiment of the present disclosure;
[0041] Figure 12C Flowchart of an exemplary process of a virtual agent manager according to an embodiment of the present disclosure;
[0042] Figure 13AIllustrative example of types of constraints observed by an augmented reality launcher according to an embodiment of the present disclosure;
[0043] Figure 13B - 13E Depicts an example of presenting a virtual scene;
[0044] Figure 13F Represents different constraints observed when presenting a virtual agent according to an embodiment of the present disclosure;
[0045] Figure 13G Represents different constraints observed when presenting a virtual object according to an embodiment of the present disclosure;
[0046] Figure 14 Depicts an exemplary high - level system diagram of an augmented reality launcher for presenting a virtual scene in an actual conversation scenario according to an embodiment of the present disclosure;
[0047] Figure 15 Is a flowchart of an exemplary process of an augmented reality launcher for combining a virtual scene and an actual conversation scenario according to an embodiment of the present disclosure;
[0048] Figure 16 Is an illustrative schematic diagram of an exemplary mobile device architecture that can be used to implement a dedicated system for implementing the present disclosure according to various embodiments; and
[0049] Figure 17 Is an illustrative schematic diagram of an exemplary computing device architecture that can be used to implement a dedicated system for implementing the present disclosure according to various embodiments. Detailed Description
[0050] In the following detailed description, numerous specific details are set forth by way of example in order to provide a thorough understanding of the relevant teachings. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without such details. In other instances, well - known methods, procedures, components, and / or circuits have not been described in detail at a relatively high level in order to avoid unnecessarily obscuring aspects of the present disclosure.
[0051] The purpose of the present disclosure is to address the deficiencies of traditional human - machine dialogue systems and provide methods and systems that can implement a more effective and realistic human - machine dialogue framework. The present disclosure integrates artificial intelligence into an automated companion with an agent device that is connected to a human interface and works with the backbone support from a user interaction engine, enabling the automated companion to conduct conversations based on multi - modal data that continuously monitors the surrounding environment indicating the conversation, adaptively estimate the mindset / mood / intention of the conversation participants, and adaptively adjust the conversation strategy based on dynamically changing information / estimates / context information.
[0052] An automated companion according to the present disclosure can personalize conversations by adapting in multiple aspects, including (but not limited to) the topic of the conversation, the hardware / software / components or the physical or virtual conversation environment for conducting the conversation, and the expressions / behaviors / poses for delivering responses to human conversationalists. The adaptive conversation strategy is to flexibly change the conversation strategy based on the observation of the acceptance of the conversation by the human conversationalist, so as to make the conversation more attractive, more authentic and more productive. The adaptive conversation system according to the present disclosure can be configured to be driven by a goal-driven strategy by intelligently and dynamically configuring the hardware / software components, conversation settings (virtual or physical) or conversation strategies that are considered most suitable for achieving the intended goal. This adaptation / optimization is based on learning, including learning from previous conversations and ongoing conversations, such as observing the behaviors / responses / emotions shown by the human conversationalist during the conversation. Even in some cases where the conversation may seem to deviate from the path initially designed for the intended goal, this goal-driven strategy can be used to keep the human conversationalist engaged in the conversation.
[0053] More specifically, the present disclosure provides a user interaction engine that provides backbone support to an agent device to facilitate a more authentic and engaging conversation with a human conversationalist. Figure 1 A networked environment 100 is depicted in connection with a user interaction engine according to an embodiment of the present disclosure to facilitate a conversation between a user of an operating user device and an agent device. In Figure 1 this example, the exemplary networked environment 100 includes one or more user devices 110, such as user devices 110-a, 110-b, 110-c, and 110-d, one or more agent devices 160, such as agent devices 160-a,... 160-b, a user interaction engine 140, and a user information database 130, each of which can communicate with each other via a network 120. In some embodiments, the network 120 can correspond to a single network or a combination of different networks. For example, the network 120 can be a local area network ("LAN"), a wide area network ("WAN"), a public network, a private network, a private network, a public switched telephone network ("PSTN"), the Internet, an intranet, a Bluetooth network, a wireless network, a virtual network, and / or any combination thereof. In one embodiment, the network 120 can also include various network access points. For example, the environment 100 can include wired or wireless access points, such as (but not limited to) base stations or Internet exchange points 120-a,... 120-b. The base stations 120-a and 120-b can facilitate communication with one or more other components in the networked framework 100, such as between the user device 110 and / or the agent device 160, via different types of networks.
[0054] User devices, such as 110-a, can be of different types to facilitate a user to operate the user device to connect to network 120 and send / receive signals. Such user devices 110 can correspond to any suitable type of electronic / computing device, including (but not limited to) desktop computers (110-d), mobile devices (110-a), devices incorporated in transportation vehicles (110-b), …, mobile computers (110-c) or fixed devices / computers (110-d). Mobile devices can include (but not limited to) mobile phones, smart phones, personal display devices, personal digital assistants (“PDAs”), gaming devices / equipment, wearable devices such as watches, smart bracelets (Fitbit), brooches / pins, earphones, etc. Transportation vehicles embedded with devices can include cars, trucks, motorcycles, boats, ships, trains or airplanes. Mobile computers can include laptop computers, ultrabook devices, handheld devices, etc. Fixed devices / computers can include televisions, set-top boxes, smart home devices (e.g., refrigerators, microwave ovens, washing machines or dryers, electronic assistants, etc.), and / or smart accessories (e.g., light bulbs, light switches, electronic photo frames, etc.).
[0055] Agent devices, such as any one of 160-a, …, 160-b, can correspond to one of different types of devices that can communicate with user devices and / or the user interaction engine 140. As described in more detail below, each agent device can be regarded as an automated companion device connected to a user interface with the backbone support, for example, from the user interaction engine 140. The agent devices described herein can correspond to robots, which can be gaming devices, toy devices, designated agent devices such as travel agents or weather agents, etc. The agent devices disclosed herein are capable of facilitating and / or assisting user interactions with the operation of user devices. In doing so, the agent device can be configured as a robot that can control some of its parts via backend support from an application server to, for example, perform certain physical movements (such as the head), exhibit certain facial expressions (such as bent eyes when smiling), or speak in a certain voice or tone (e.g., an exciting tone) to express a certain emotion.
[0056] When a user device (e.g., user device 110-a) is connected to a proxy device, such as 160-a (e.g., via a contact or non-contact connection), the client running on the user device, such as 110-a, can communicate with the automation companion (proxy device and / or user interaction engine) to enable an interactive dialogue between the user operating the user device and the proxy device. The client can act independently in some tasks or can be remotely controlled by the proxy device or the user interaction engine 140. For example, in response to a question from the user, the proxy device or the user interaction engine 140 can control the client running on the user device to present a response voice to the user. During the session, the proxy device can include one or more input mechanisms (e.g., camera, microphone, touch screen, buttons, etc.) that enable the proxy device to capture inputs related to the local environment associated with the user or the session. Such inputs can help the automation companion understand the atmosphere around the session (e.g., the user's movement, the sound of the environment) and the mindset of the human conversant (e.g., the user picking up a ball may indicate that the user is bored), so that the automation companion can react accordingly and conduct the dialogue in a way that keeps the user interested and engaged.
[0057] In the illustrated embodiment, the user interaction engine 140 can be a backend server, which can be centralized or distributed. It is connected to the proxy device and / or the user device. It can be configured to provide backbone support to the proxy device 160 and guide the proxy device to conduct the session in a personalized and customized manner. In some embodiments, the user interaction engine 140 can receive information from the connected device (proxy device or user device), analyze such information, and control the flow of the session by sending instructions to the proxy device and / or the user device. In some embodiments, the user interaction engine 140 can also communicate directly with the user device, such as providing dynamic data, such as control signals, to the client running on the user device to present certain responses.
[0058] Generally, the user interaction engine 140 can control the state and flow of a session between a user and an agent device. The flow of each session can be controlled based on different types of information associated with the session, such as information about the user participating in the session (e.g., from the user information database 130), the session history, the ambient information of the session, and / or real-time user feedback. In some embodiments, the user interaction engine 140 can be configured to obtain various sensory inputs, such as (but not limited to) audio input, image input, tactile input, and / or context input, process these inputs, form an understanding of the human conversant, generate a response based on such understanding, and control the agent device and / or the user device to conduct the session based on the response. As an illustrative example, the user interaction engine 140 can receive audio data representing speech from a user operating the user device and generate a response (e.g., text), which can then be delivered to the user as a computer-generated speech as a response to the user. As another example, the user interaction engine 140 can also generate one or more instructions for controlling the agent device to perform a specific action or a set of actions in response to speech.
[0059] As shown in the figure, during a human-machine conversation, a user, who is a human conversant in the conversation, can communicate with the agent device or the user interaction engine 140 via the network 120. Such communication can involve various modalities of data, such as audio, video, text, etc. Via the user device, the user can send data (e.g., requests, audio signals representing the user's speech, or videos of the scene around the user) and / or receive data (e.g., text or audio responses from the agent device). In some embodiments, after being received by the agent device or the user interaction engine 140, various modalities of user data can be analyzed to understand the speech or gestures of the human user, so that the user's emotion or intention can be estimated and used to determine the response to the user.
[0060] Figure 2ADepict the specific connections among the user device 110-a, the proxy device 160-a, and the user interaction engine 140 during a conversation according to an embodiment of the present disclosure. It can be seen that, as described herein, the connection between any two parties can be bidirectional. The proxy device 160-a can be connected to the user interface via the user device 110-a to conduct a conversation in two-way communication. On the one hand, the proxy device 160-a can be controlled by the user interaction engine 140 to speak a response to the user operating the user device 110-a. On the other hand, the input from the user's location, such as including the user's words / actions and information about the user's surrounding environment, is provided to the proxy device via the connection. The proxy device 160-a can be configured to process such input and dynamically adjust its response to the user. For example, the user interaction engine 140 can instruct the proxy device to present a tree on the user device. Knowing that the user's surrounding environment shows green trees and lawns (based on the visual information from the user device), the proxy device can customize the tree to present as lush green trees. If the scene display from the user's location shows winter weather, the proxy device can control to present the tree on the user device using the parameters for a tree without leaves. As another example, if the proxy device is instructed to present a duck on the user device, the proxy device can retrieve information about the color preference from the user information database 130 and generate parameters for customizing the duck in the user's preferred color, and then send a presentation instruction to the user device.
[0061] In some embodiments, such input from the user's location and its processing results can also be sent to the user interaction engine 140 to facilitate the user interaction engine 140 to better understand the specific situation associated with the conversation, so that the user interaction engine 140 can determine the state of the conversation, the user's mood / mentality, and generate a response based on the specific situation of the conversation and the intended purpose of the conversation (e.g., for teaching children English vocabulary). For example, if the information received from the user device indicates that the user seems bored and becomes impatient, the user interaction engine 140 can determine to change the conversation state to a topic that the user is interested in (e.g., based on the information from the user information database 130) to continue to engage the user in the conversation.
[0062] In some embodiments, a client running on a user device may be configured to be capable of processing different modalities of raw input obtained from a user location and sending the processed information (e.g., relevant features of the raw input) to a proxy device or a user interaction engine for further processing. This will reduce the amount of data sent over the network and improve communication efficiency. Similarly, in some embodiments, the proxy device may also be configured to be capable of processing information from the user device and extracting information useful for customization, for example. Although the user interaction engine 140 can control the state and flow of a conversation, making the user interaction engine 140 lightweight better improves the scalability of the user interaction engine 140.
[0063] Figure 2B depicts the same setup as shown in Figure 2A and additional details regarding the user device 110-a. As shown, during a conversation between the user and the proxy 160-a, the user device 110-a can continuously collect multimodal sensor data related to the user and his / her surrounding environment, which can be analyzed to detect any information relevant to the conversation and used to intelligently control the conversation in an adaptive manner. This can further enhance the user experience or engagement. Figure 2B Illustrates exemplary sensors such as a video sensor 230, an audio sensor 240, …, or a tactile sensor 250. The user device may also send text data as part of the multimodal sensor data. These sensors together provide contextual information around the conversation and can be used by the user interaction system 140 to understand the situation in order to manage the conversation. In some embodiments, the multimodal sensor data may be processed first on the user device, and important features of different modalities may be extracted and sent to the user interaction system 140 so that the conversation can be controlled in the context of understanding. In some embodiments, the raw multimodal sensor data may be sent directly to the user interaction system 140 for processing.
[0064] As Figure 2A - 2B shown, the proxy device may correspond to a robot with different parts, including its head 210 and its body 220. Although the proxy device shown in Figure 2A - 2B is depicted as a humanoid robot, the proxy device may also be constructed in other forms, such as a duck, a bear, a rabbit, etc. Figure 3AAccording to an embodiment of the present disclosure, an exemplary structure of an agent device having an exemplary type of agent body is illustrated. As shown, the agent device may include a head and a body, where the head is attached to the body. In some embodiments, the head of the agent device may have additional parts such as a face, a nose, and a mouth, some of which may be controlled to make movements or expressions, for example. In some embodiments, the face on the agent device may correspond to a display screen on which a face may be presented, and the face may be a human or animal face. Such a displayed face may also be controlled to express emotions.
[0065] The body part of the agent device may also correspond to different forms, such as a duck, a bear, a rabbit, etc. The body of the agent device may be fixed, movable, or semi - movable. An agent device with a fixed body may correspond to a device that can sit on a surface such as a table to have a face - to - face conversation with a human user sitting next to the table. An agent device with a movable body may correspond to a device that can move back and forth on a surface such as a desktop or a floor. Such a movable body may include parts that can be controlled kinematically to make physical movements. For example, the agent body may include feet that can be controlled to move in space when needed. In some embodiments, the body of the agent device may be semi - movable, that is, some parts are movable and some parts are immovable. For example, the tail on the body of an agent device with a duck appearance may be movable, but the duck cannot move in space. A bear - shaped agent device may also have movable arms, but the bear can only sit on a surface.
[0066] Figure 3B An exemplary agent device or automated companion 160 - a according to an embodiment of the present disclosure is illustrated. The automated companion 160 - a is a device that interacts with people using voice and / or facial expressions or body postures. For example, the automated companion 160 - a corresponds to an electronic animal peripheral device having different parts, the different parts including a head 310, an eye part (camera) 320, a mouth having a laser 325 and a microphone 330, a speaker 340, a neck having a servo mechanism 350, one or more magnets or other components 360 that can be used for non - contact detection of presence, and a body part 370 corresponding to, for example, a charging dock. In operation, the automated companion 160 - a may be connected via a network connection to a user device that may include a mobile multifunctional device (110 - a). Once connected, the automated companion 160 - a and the user device interact with each other via, for example, voice, movement, posture, and / or via an indication using a laser pointer.
[0067] Other exemplary functions of the automated companion 160-a can include, for example, reactive expressions via an interactive video cartoon character (e.g., an avatar) in response to user responses, where the interactive video cartoon character is displayed on a screen that is, for example, part of the face of the automated companion. The automated companion can use a camera (320) to observe the user's presence, facial expressions, gaze direction, surrounding environment, etc. An electronic animal embodiment can "look" by aligning its head (310) that contains the camera (320), "listen" using its microphone (330), and "aim" by directing its head (310) that can be moved via a servo mechanism (350). In some embodiments, the head of the agent device can also be remotely controlled via a laser (325), for example, by the user interaction system 140 or by a client in the user device (110-a). It can also be controlled to "speak" via a speaker (340) as shown in Figure 3B the exemplary automated companion 160-a.
[0068] Figure 4A An exemplary high-level system diagram of the overall system for an automated companion in accordance with various embodiments of the present disclosure is depicted. In the embodiment shown in this illustration, the overall system can include components / functional modules residing in the user device, the agent device, and the user interaction engine 140. The overall system depicted here includes multiple processing layers and hierarchies that interact in an intelligent manner between humans and machines. In the embodiment shown in the illustration, there are a total of 5 layers, including layer 1 for the front-end application and front-end multimodal data processing, layer 2 for the representation of the dialogue setting, layer 3 where the dialogue management module resides, layer 4 for the estimated mental states of different parties (humans, agents, devices, etc.), and layer 5 for the so-called utility. Different layers can correspond to different processing levels, from the acquisition and processing of raw data in layer 1 to the processing of changing the utility of the dialogue participants in layer 5.
[0069] The term "utility" is hereby defined as the preference of a party identified based on the detected state associated with the conversation history. Utility can be associated with the parties in a conversation, whether the parties are human, an automated companion, or other intelligent devices. The utility regarding a specific party can represent different states of the world, whether physical, virtual, or even mental. For example, a state can be represented as a specific path that the conversation has taken in a complex map of the world. In different situations, based on the interaction among multiple parties, the current state evolves into the next state. The state can also be party-dependent, that is, when different parties participate in the interaction, the state generated by such interaction may vary. The utility associated with a party can be organized into a hierarchy of preferences, and such a hierarchy of preferences can evolve over time based on the choices made and the preferences shown by the party during the conversation. Such a preference, which can be represented as an ordered sequence of choices made from different options, is called utility. This disclosure presents a method and system by which an intelligent automated companion can learn the user's utility through a conversation with a human conversant.
[0070] Within the overall system for supporting an automated companion, the front-end application and front-end multimodal data processing in Layer 1 can reside in the user device and / or the proxy device. For example, a camera, microphone, keyboard, display, presenter, speaker, chat bubble, and user interface elements can be components or functional modules of the user device. For example, there can be an application or client running on the user device, which can include functions before the external application programming interface (API) as shown in Figure 4A . In some embodiments, functions outside the external API can be considered as the back-end system or reside in the user interaction engine 140. The application running on the user device can obtain multimodal data (audio, image, video, text) from the sensors or circuits of the user device, process the multimodal data to generate a text or other type of signal (such as an object like a detected user face, speech understanding result) representing the features of the original multimodal data, and send it to Layer 2 of the system.
[0071] In Layer 1, multimodal data can be obtained via sensors such as a camera, microphone, keyboard, display, speaker, chat bubble, presenter, or other user interface elements. Such multimodal data can be analyzed to estimate or infer various features, and the various features can be used to infer high-level characteristics such as expressions, character, poses, emotions, actions, attention, intentions, etc. Such high-level characteristics can be obtained by the processing unit at Layer 2 and via Figure 4AThe internal API shown is used by higher-level components to, for example, intelligently infer or estimate additional information related to a conversation at a higher conceptual level. For example, the estimated mood, attention, or other characteristics of the participants in the conversation obtained at layer 2 can be used to estimate the mindset of the participants. In some embodiments, such a mindset can also be estimated at layer 4 based on additional information (e.g., the recorded surrounding environment or other auxiliary information in such a surrounding environment, such as sound).
[0072] The conversation management at layer 3 can rely on the estimated mindset of the parties involved (whether related to a human or an automated companion (machine)) to determine, for example, how to conduct a conversation with a human conversant. How each conversation proceeds typically represents the preferences of the human user. Such preferences can be dynamically captured during the conversation at the utility layer (layer 5). As Figure 4A shown, the utility representation at layer 5 indicates the evolving state of the evolving preferences of the parties involved, which can also be used by the conversation management at layer 3 to decide on an appropriate or intelligent way to conduct the interaction.
[0073] Information sharing between different layers can be achieved via the API. In some embodiments, as Figure 4A illustrated in the figure, information sharing between layer 1 and other layers is via an external API, while information sharing between layers 2 - 5 is via an internal API. It should be understood that this is merely a design choice, and other implementations can also achieve the teachings presented herein. In some embodiments, via the internal API, each layer (2 - 5) can access information created or stored by other layers to support processing. Such information can include common configurations to be applied to the conversation (e.g., the role of the agent device is an avatar, the preferred voice, or the virtual environment created for the conversation, etc.), the current state of the conversation, the current conversation history, known user preferences, estimated user intent / mood / mindset, etc. In some embodiments, some information that can be shared via the internal API can be accessed from an external database. For example, certain configurations related to the desired role of the agent device (duck) can be accessed from, for example, an open-source database that provides parameters (e.g., parameters for visually presenting the duck and / or parameters required to present the voice from the duck).
[0074] Figure 4BIllustrates a portion of a dialogue tree for an ongoing dialogue according to an embodiment of the present disclosure, and the path taken based on the interaction between the automated companion and the user. In the example shown in this illustration, the dialogue management at layer 3 (of the automated companion) can predict multiple paths by which the dialogue (or more generally the interaction with the user) can continue. In this example, each node can represent a point in the current state of the dialogue, and each branch originating from the node can represent a possible response from the user. As shown in this example, at node 1, the automated companion may face three independent paths that it may take based on the response detected from the user. If the user replies with an affirmative response, the dialogue tree 400 can proceed from node 1 to node 2. At node 2, in response to an affirmative response from the user, a response can be generated for the automated companion and then presented to the user, and the response can include audio, visual, text, tactile, or any combination thereof.
[0075] If, at node 1, the user's response is negative, the path for this stage is from node 1 to node 10. If the user replies with a "neither good nor bad" response (e.g., not negative, but also not affirmative) at node 1, the dialogue tree 400 can proceed to node 3, where a response from the automated companion can be presented, and there can be three independent possible responses from the user corresponding to nodes 5, 6, and 7 respectively, "no response", "affirmative response", and "negative response". Depending on the user's actual response to the response of the automated companion presented at node 3, the dialogue management at layer 3 can then follow the corresponding dialogue. For example, if the user replies with an affirmative response at node 3, the automated companion transfers to respond to the user at node 6. Similarly, depending on the user's reaction to the response of the automated companion at node 6, the user may further reply with the correct answer. In this case, the dialogue state transfers from node 6 to node 8, etc. In the example shown in this illustration, during this period, the dialogue state transfers from node 1 to node 3, to node 6, and then to node 8. Traversing nodes 1, 3, 6, and 8 forms a path consistent with the underlying conversation between the automated companion and the user. As Figure 4B shown, the path representing the dialogue is indicated by the solid line connecting nodes 1, 3, 6, and 8, while the paths skipped during the dialogue are indicated by the dashed line.
[0076] Figure 5 Illustrates exemplary communication between different processing layers of an automated dialogue companion centered around a dialogue manager 510 according to embodiments of the present disclosure. Figure 5The dialogue manager 510 therein corresponds to the functional component of the dialogue management at layer 3. The dialogue manager is an important part of the automated companion and manages the dialogue. Traditionally, the dialogue manager takes the user's utterance as input and determines how to respond to the user. This is done without considering the user's preferences, the user's mindset / mood / intentions, or the context of the dialogue, i.e., without giving any weight to the different available states of the relevant world. The lack of understanding of the surrounding world often limits the perceived authenticity or engagement of the conversation between the human user and the intelligent agent.
[0077] In some embodiments of the present disclosure, the utilities of the various parties involved in the ongoing conversation are utilized to facilitate a more personalized, flexible, and engaging conversation. It facilitates the intelligent agent acting in different roles to be more effective in different tasks (such as scheduling appointments, booking trips, ordering equipment and supplies, and researching various topics online). When the intelligent agent knows the user's dynamic mindset, mood, intentions, and / or utilities, it enables the agent to engage the human conversant in the conversation in a more targeted and effective manner. For example, when an educational agent teaches a child, the child's preferences (e.g., his favorite color), the observed mood (e.g., sometimes the child does not want to continue the class), the intention (e.g., the child is not focusing on the lesson but reaching for a ball on the floor) can all allow the educational agent to flexibly adjust the focus topic to the toy and perhaps adjust the way of continuing the conversation with the child so that the child can take a break to achieve the overall goal of educating the child.
[0078] As another example, the present disclosure can be used to enhance the service level of the customer service agent by asking more appropriate questions considering the situation observed from the user in real time, thereby achieving an improved user experience. This is achieved by forming means and methods for learning and adapting to the preferences or mindsets of the parties involved in the conversation so as to be able to conduct the conversation in a more engaging manner, which is rooted in the fundamental aspects of the present disclosure as disclosed herein.
[0079] The dialogue manager (DM) 510 is the core component of the automated companion. As Figure 5 shown, the DM 510 (layer 3) receives inputs from different layers, including inputs from layer 2, and inputs from higher abstraction layers such as layer 4 for estimating the mindsets of the parties involved in the conversation, and layer 5 for learning utilities / preferences based on the dialogue and its evaluation performance. As shown, at layer 1, multimodal information is obtained from sensors of different modalities, and the multimodal information is processed to obtain features such as characterization data. This can include signal processing of visual, auditory, and text modalities.
[0080] Such multimodal information can be obtained by sensors deployed on a user device (e.g., 110-a) during a conversation. The obtained multimodal information may be related to the user operating the user device 110-a and / or the surrounding environment of the conversation scenario. In some embodiments, the multimodal information may also be obtained by a proxy device (e.g., 160-a) during a conversation. In some embodiments, sensors on both the user device and the proxy device may obtain relevant information. In some embodiments, as Figure 5 shown, the obtained multimodal information is processed at layer 1, which may include both the user device and the proxy device. Depending on the situation and configuration, the layer 1 processing on each device may be different. For example, if the user device 110-a is used to obtain information about the surrounding environment of the conversation, including information about both the user and the environment around the user, the raw input data (e.g., text data, visual data, or audio data) can be processed on the user device, and then the processed features can be sent to layer 2 for further analysis (at a higher level of abstraction). If some of the multimodal information about the user and the conversation environment is obtained by the proxy device, the processing of such obtained raw data can also be processed by the proxy device ( Figure 5 not shown in), and then the features extracted from such raw data can be sent from the proxy device to layer 2 (layer 2 may be located in the user interaction engine 140).
[0081] Layer 1 is also responsible for the information presentation of the response to the user from the automated conversation partner. In some embodiments, the presentation is performed by a proxy device (e.g., 160-a), and examples of such presentation include voice, expressions, which may be facial expressions or body movements made. For example, the proxy device can present a text string received from the user interaction engine 140 (as a response to the user) as voice so that the proxy device can speak the response to the user. In some embodiments, the text string may be sent to the proxy device together with additional presentation instructions such as volume, intonation, pitch, and the like, which may be used to convert the text string into sound waves corresponding to the utterance of the content in a certain way. In some embodiments, the response to be delivered to the user may also include an animation, for example, speaking the response in a gesture that can be conveyed via, for example, facial expressions or body movements such as raising an arm. In some embodiments, the proxy can be implemented as an application on the user device. In this case, via the user device, e.g., 110-a ( Figure 5 not shown in) implements the presentation of the response from the automated conversation partner.
[0082] The processed features of the multimodal data can be further processed at layer 2 to achieve language understanding and / or multimodal data understanding, including vision, text, and any combination thereof. Some such understanding can be for a single modality, such as speech understanding, while some can be based on comprehensive information for understanding the surrounding environment of the users participating in the conversation. Such understanding can be physical (e.g., identifying certain objects in a scene), perceivable (e.g., identifying what the user said, or certain important sounds, etc.), or psychological (e.g., certain emotions, such as the user's stress based on, for example, the intonation of speech, facial expressions, or the user's pose estimation).
[0083] The multimodal data understanding generated at layer 2 can be used by DM510 to determine how to respond. To enhance engagement and the user experience, DM510 can also determine the response based on the estimated mindsets of the user and the agent from layer 4, and the utility of the user participating in the conversation from layer 5. The mindsets of the parties involved in the conversation can be estimated based on information from layer 2 (e.g., the estimated emotions of the user) and the progress of the conversation. In some embodiments, the mindsets of the user and the agent can be dynamically estimated during the conversation, and such estimated mindsets can then be used together with other data to learn the user's utility. The learned utility represents the user's preferences in different conversation scenarios and is estimated based on historical conversations and their outcomes.
[0084] In each conversation on a certain topic, the dialogue manager 510 bases its control of the conversation on a relevant conversation tree that may or may not be associated with the topic (e.g., may add small talk to enhance engagement). To generate a response to the user in the conversation, the dialogue manager 510 can also consider additional information, such as the user's state, the surrounding environment of the conversation scenario, the user's emotions, the estimated mindsets of the user and the agent, and the known preferences (utility) of the user.
[0085] The output of DM510 corresponds to the correspondingly determined response to the user. To convey the response to the user, DM510 can also formulate the way to convey the response. The form of conveying the response can be determined based on information from multiple sources, for example, the user's emotions (e.g., if the user is an unhappy child, the response can be presented in a soft voice), the user's utility (e.g., the user may prefer a voice with an intonation similar to that of his parents), or the surrounding environment where the user is located (e.g., a noisy place, so that the response needs to be conveyed at a higher volume). DM510 can output the determined response and such conveyance parameters.
[0086] In some embodiments, the delivery of such determined responses is achieved by generating a deliverable form of the response according to various parameters associated with each response. In general, responses are delivered in the form of speech in some natural language. Responses can also be delivered with speech coupled with specific non-verbal expressions (such as nodding, shaking the head, winking, or shrugging) that are part of the delivered response. There can be other deliverable forms of responses that are auditory but not verbal, such as whistling.
[0087] To deliver a response, as Figure 5 shown, the deliverable form of the response can be generated via, for example, speech response generation and / or behavioral response generation. Such responses for which a deliverable form has been determined can then be used by a presenter to actually present the response in its intended form. For the deliverable form of a natural language, the text of the response can be used to synthesize a speech signal via, for example, text-to-speech technology, according to delivery parameters (such as volume, intonation, style, etc.). For any response or part thereof to be delivered in a non-verbal form with a certain expression, the expected non-verbal expression can be transformed, for example, via animation into control signals that can be used to control certain parts of the agent device (the physical representation of the automated companion) to perform certain mechanical movements to deliver the non-verbal expression of the response, such as nodding, shrugging, or whistling. In some embodiments, to deliver a response, certain software components can be invoked to present different facial expressions of the agent device. This presentation of the response can also be done simultaneously by the agent (for example, the agent says the response in a joking voice with a big smile).
[0088] Figure 6A Depicted is an exemplary configuration of a dialogue system 600 centered on an information state 610 that captures dynamic information observed during a conversation, according to an embodiment of the present disclosure. The dialogue system 600 includes a multimodal information processor 620, an automatic speech recognition (ASR) engine 630, a natural language understanding (NLU) engine 640, a dialogue manager (DM) 650, a natural language generation (NLG) engine 660, and a text-to-speech (TTS) engine 670. The system 600 is interface-connected to a user 680 for conversation.
[0089] During a conversation, multimodal information is collected from the environment (including from user 680), which captures the ambient information of the conversation environment, the speech from user 680, the user's expressions (whether facial or bodily), etc. The multimodal information thus collected is analyzed by the multimodal information processor 620 to extract relevant representative features of different modalities in order to estimate different characteristics of the user, the environment, etc. For example, the speech signal can be analyzed to determine speech-related features such as speech rate, pitch, or even accent. Visual signals related to the user can also be analyzed to determine, for example, facial features or body postures in order to determine the user's expression. Combining the auditory features and the visual features, the multimodal information analyzer 620 is also able to estimate the user's emotional state. For example, a high pitch and a fast speech rate plus an angry facial expression can indicate that the user is upset. In some embodiments, the observed user activities can also be analyzed to indicate, for example, that the user is facing or walking towards a specific object. Such information can provide a context that can be used to understand the user's intention or what the user mentions in his / her speech. The multimodal information processor 620 can continuously analyze the multimodal information and store such analyzed information in the information state 610, which is then used by different components in the system 100 to facilitate decision-making.
[0090] In operation, the speech information of user 680 is sent to the ASR engine 630 for speech recognition. The speech recognition can include identifying the language spoken and the words uttered by user 680. To understand the semantics of what the user said, the result from the ASR engine 630 is further processed by the NLU engine 640. Such understanding can depend not only on the words spoken, but also on other information such as the posture of user 680 and / or other context information such as previously spoken words. Based on the understanding of the user's utterance, the dialogue manager 650 (the same as the dialogue manager 510 in Figure 5 determines how to respond to the user, and such determined response can then be generated by the NLG engine 660 and further converted from text form to a speech signal via the TTS engine 670. Then, the output of the TTS engine 670 can be passed to user 680 as a response to the user's utterance. The process continues with such back-and-forth responses to conduct a conversation with user 680.
[0091] As Figure 6AAs shown, the components in system 600 are connected to information state 610, which, as described herein, captures the dynamics surrounding a conversation and provides relevant and rich contextual information that can be used to facilitate automatic speech recognition (ASR), natural language understanding (NLU) to determine an appropriate response (DM), generate a response (NLG), and convert the generated text response into a speech form (TTS). As described herein, information state 610 can represent the conversation-related dynamics obtained based on multimodal information related to user 680 or the surrounding environment of the conversation.
[0092] When multimodal information (about the user or about the surrounding environment of the conversation) is received from the conversation scenario, multimodal information processor 670 analyzes the information and represents the surrounding environment of the conversation at different levels, e.g., acoustic characteristics (e.g., pitch, speed, the user's accent), visual characteristics (e.g., the user's facial expression, objects in the environment), physical characteristics (e.g., the user's hand waving or pointing at an object in the environment), the estimated emotional and / or mental state of the user, and / or the user's preferences or intentions. Such information can then be stored in information state 610.
[0093] The rich media contextual information stored in information state 610 can help facilitate the different components in playing their respective roles so that the conversation can be conducted in a more engaging and effective manner with respect to the intended goal, e.g., understanding the user's utterances 680 based on what is observed in the conversation scenario, evaluating the performance of user 680, and / or estimating the utility associated with the user based on the intended goal of the conversation, determining how to respond to the user's utterances 680 based on the evaluated performance and utility of the user, and delivering the response in the most appropriate way based on the understanding of the user, etc.
[0094] For example, with respect to the intonation information represented in both auditory form (e.g., a specific way of uttering certain phonemes) and visual form (e.g., specific visemes of the user) captured in the information state regarding the user, the ASR engine 630 can utilize this information to determine the words spoken by the user. Similarly, the NLU engine 640 can also utilize rich context information to determine the semantics intended by the user. For example, if the user points to a computer placed on the table (visual information) and says "I like this", the NLU engine 640 can combine the output of the ASR engine 630 (i.e., "I like this") and the visual information that the user is pointing to the computer in the room to understand that what the user means by "this" refers to the computer. As another example, if the user 680 repeatedly makes mistakes during the tutoring process and, at the same time, based on the evaluation of the voice tone and facial expression (determined based on multimodal information), the user looks very annoyed, the DM can determine to temporarily change the topic based on the known preferences of the user (e.g., likes to discuss Lego games) in order to continue to attract the user instead of continuing to emphasize the tutoring content. The decision to temporarily divert the user's attention can be determined, for example, based on the previous utility observed for the user regarding what works (e.g., temporarily divert the user's attention based on some favorite topics that work) and what does not work (e.g., continue to pressure the user to do better).
[0095] Figure 6B is a flowchart of an exemplary process of a dialogue system 600 using the information state 610 that captures dynamic information observed during a dialogue according to an embodiment of the present disclosure. As Figure 6B shown, the process is an iterative process. At 605, multimodal information is received, and then the multimodal information is analyzed by the multi-information processor 670 at 625. As described herein, the multimodal information includes information related to the user 680 and / or information related to the dialogue surrounding environment. The multimodal information related to the user can include the user's utterances and / or visual observations of the user, such as body postures and / or facial expressions. The information related to the dialogue surrounding environment can include information related to the environment, such as the objects present, the spatial / temporal relationship between the user and such observed objects (e.g., the user stands in front of the table), and / or the dynamics between the user's activities and the observed objects (e.g., the user walks towards the table and points to the computer on the table). Then, the understanding of the multimodal information captured from the dialogue scenario can be used to facilitate other tasks in the dialogue system 600.
[0096] Based on the information stored in the information state 610 (representing the past state) and the analysis results from the multimodal information processor 670 (regarding the current state), the ASR engine 620 and the NLU engine 630 respectively perform speech recognition at 625 to identify the words spoken by the user and perform language understanding based on the identified words. ASR and NLU can be performed based on the current information state 610 and the analysis results from the multimodal information processor 670.
[0097] Based on the multimodal information analysis and the language understanding result, that is, what the user said or meant, at 635, the change in the dialogue state is tracked, and at 645, such a change is used to update the information state 610 accordingly for subsequent processing. To conduct a dialogue, the DM 650 determines a response at 655 based on the dialogue tree designed for the basic dialogue, the output of the NLU engine 630 (the understanding of the utterance), and the information stored in the information state 610. Once the response is determined, the NLG engine 650 generates a response based on the information state 610, for example, in text form. When the response is determined, there can be different ways to express the response. At 665, the NLG engine 650 can generate a response in a way based on the user's preferences or anything known to be more suitable for the specific user in the current dialogue. For example, if the user answers a question incorrectly, there are different ways to indicate that the answer is incorrect. For a specific user in the current dialogue, if the user is known to be sensitive and easily frustrated, a gentler way to tell the user that his / her answer is incorrect can be used to generate the response. For example, the NLG engine 650 can generate a text response of "The answer is not entirely correct" instead of saying "The answer is wrong".
[0098] The text response generated by the NLG engine 650 can then be presented in voice form as an audio signal by the TTS engine 660 at 675. Although standard or common TTS techniques can be used for TTS, it is disclosed herein that the response generated by the NLG engine 650 can be further personalized based on the information stored in the information state 610. For example, if a slower speech rate or a gentler speaking manner is known to be more effective for the user (for example, it is known that students with ADHD, for example, have a slower processing speed for speech), the generated response can be presented in voice form by the TTS engine 660 at 675 at a lower speed and pitch accordingly. Another example is to present the response in a tone consistent with the known accent of the student according to the personalized information about the user in the information state 610. Then at 685, the presented response can be passed to the user as a response to the user's utterance. When responding to the user, the dialogue system 600 then tracks additional changes in the dialogue and updates the information state 610 accordingly at 695.
[0099] Figure 7ADepicts an exemplary structure of an information state representation 610 according to an embodiment of the present disclosure. The information state 610 includes (but is not limited to) an estimated mindset or mental state. As shown, the estimated mindset includes the agent's mindset 700, the user's mindset 720, and a shared mindset 710 related to other information recorded therein. The agent's mindset 700 may refer to one or more expected goals that a dialogue agent (machine) aims to achieve in a specific dialogue. The shared mindset 710 may refer to a representation of the current dialogue situation, which is a combination of the agent's execution of the expected agenda according to the agent's mindset 700 and the user's performance. The user's mindset 720 may refer to a representation of the agent's estimation of which stage the student is at with respect to the expected purpose of the dialogue based on the shared mindset or the user's performance. For example, if the agent's current task is to teach a student user the concept of fractions in mathematics (which may include sub - concepts to build an understanding of fractions), the user's mindset may include an estimated user's mastery of various sub - concepts. Such an estimation can be derived based on an assessment of the student's performance at different stages of tutoring related sub - concepts.
[0100] Figure 7B Illustrates how these different mindsets are connected in an example where a robot teacher 705 teaches a student user 680 a concept 715 related to fraction addition according to an embodiment of the present disclosure. As shown, the robot agent 705 interacts with the student user 680 via multimodal interaction. The robot agent 705 can start tutoring based on an initial agent mindset 700 (e.g., a fraction addition course that can be represented as an AOG). During the tutoring, the student user 180 can answer questions from the robot teacher 705, and these answers to the questions form a certain path, thus generating a shared mindset 710. Based on the user's answers, the user's performance is evaluated, and the user's mindset 720 is estimated with respect to different aspects, such as whether the student has mastered the taught concept.
[0101] As Figure 7A shown, the estimated mindset is also related to or includes various representations of various types of other information, including (but not limited to) a spatio - temporal causal AND - OR graph STC - AOG 730, an STC parse graph (STC - PG) 740, a dialogue history 750, a dialogue context 760, event - centered knowledge 770, a commonsense model 780, …, and a user profile 790. These different types of information can be multimodal and constitute different aspects of the dynamics of each dialogue with respect to each user. Thus, the information state 610 captures both the general information of each dialogue and the individualized information with respect to each user and each dialogue.
[0102] These different mindsets are interconnected, and together they facilitate different components in the dialogue system 600 to perform corresponding tasks in a more adaptive, personalized, and engaging manner.Figure 7C Illustrates an exemplary relationship among the agent's thought 700, the shared thought 710, and the user's thought 720 represented in the information state 610 according to an embodiment of the present disclosure. As described herein, the shared thought 710 is a representation of the current conversation setting obtained based on what the agent says to the user and the user's response to the agent, and it is a combination of what the agent (according to the agent's thought) expects and how the user behaves when following the agent's expected agenda. Based on the shared thought 710, it can be tracked what the agent has been able to achieve so far and what the user has been able to accomplish.
[0103] Tracking such dynamic knowledge enables the system to estimate what the user has achieved so far, or in a tutoring setting, which concepts or sub - concepts the student user has mastered so far, i.e., to estimate the user's thought 720. The estimated thought of the student facilitates the agent to adjust or update the dialogue strategy in order to achieve the expected goal, or to adjust the agent's thought 700 by learning how to adapt to the user to derive an updated agent's thought 700. Based on the dialogue history, the dialogue system 600 learns the user's preferences or what is more effective (utility) for the user, and such information will be incorporated into the information state, which can be used by the agent to adjust the dialogue strategy based on utility - driven dialogue planning, which further leads to a further update of the shared thought based on the user's response. Repeating this process, the agent continues to adjust the dialogue strategy based on the information state.
[0104] Figure 8A Illustrates an exemplary dialogue scenario 800 according to an embodiment of the present disclosure, in which augmented and virtual reality can be combined to implement an adaptive dialogue strategy. In this example, it can be seen that the dialogue scenario is a room with multiple objects, including a table, chairs, a computer on the table, walls, windows, and various paintings and / or boards on the walls. In addition to such fixtures, there is a robotic agent 810 on the table, a user 840, some sensors such as cameras 820 - 1 and 820 - 2, a speaker 830, and a drone 850. Some of the sensors can be fixed (such as cameras 820 and speaker 830), while others can be deployed on - the - fly based on need. An example of such a deployable sensor based on need is a drone, such as 850. Such a deployable device can include both sensing capabilities and other functions (such as the ability to produce sound to convey, for example, the words of the robotic agent). Depending on the purpose of deploying such a deployable sensor, the location and orientation of the deployment can be determined based on such a purpose. For example, as Figure 8AAs shown, user 840 may enter the conversation scene 800 with their head turned towards the window, such that neither the robot agent 810 nor the cameras 820-1 and 820-2 can capture the user's face. When the robot agent wishes to identify the user based on the face, it can utilize the resources under its control to obtain the user's face. For example, by deploying the drone 850 and adjusting the target of the drone's camera to obtain the required data. This is as Figure 8A illustrated in the figure.
[0105] Meanwhile, in some cases, identifying a person based on the face can also be based on a side view of the face. In this case, the robot agent can also adjust the parameters of certain cameras in the conversation scene to capture the side view of the person for easy identification. Although the cameras deployed in the conversation scene can be installed as fixed objects, their poses can be adjusted by applying different tilts, angles, etc. to capture visual information of different regions of the scene. The robot agent 810 can be configured to be able to adjust the parameters of different sensors to achieve this. As Figure 8A shown, to obtain a side view of user 840, the robot agent 810 can control the parameters of camera 820-2 to achieve this. In addition to visual information, the robot agent can also control sensors of other modalities to obtain the required information. For example, it can control Figure 8A the audio sensor 830 in to collect auditory information in the conversation scene 800. In some embodiments, there may be multiple robots (not shown) with designated tasks in the conversation scene. For example, the robot agent 810 can act as the main robot, and there can be other deployable auxiliary robots in the scene that can be deployed by the main robot for dynamically designated tasks. For example, the main robot 810 can deploy an auxiliary robot to, for example, walk towards the user and have a conversation with the user to confirm, for example, the user's identity or certain information involved.
[0106] Through deployable and / or adjustable multi-modal sensors / robots, flexible and adaptive conversation strategies can be achieved. In some embodiments, in addition to deploying physical sensors / robots, virtual agents can also be generated and presented in the conversation scene. Figure 8BThe figure illustrates exemplary aspects of the operation of an adaptive dialogue strategy in accordance with an embodiment of the present disclosure. The adaptive dialogue strategy may include a spontaneous dialogue, which may rely on, for example, environmental modeling flexibly deployed via a requirements-based sensor / robot, and a dialogue based on collaborative multi-agents. As described below, in a spontaneous dialogue between a user and a robot agent, the robot agent may dynamically deploy a virtual agent when needed, such that the dialogue will take place in an augmented reality dialogue scenario that combines a virtual scenario and a physical scenario. Each virtual agent may have a specified task (e.g., testing a student on a mathematical concept in a virtual scenario), and may be configured to conduct a conversation with the user within the scope of the specified task.
[0107] During a conversation between a virtual agent and a user, the virtual agent may be configured to conduct the conversation based on virtual content that is part of a virtual scenario. For example, the robot agent may generate a virtual agent and a plurality of virtual objects and present them as a virtual scenario in the dialogue scenario. The virtual agent may conduct a specified conversation in accordance with the virtual objects presented in the virtual scenario. Such dynamic virtual content may be determined dynamically based on a main conversation between the robot agent and the user. For example, if the robot agent is teaching a student user the mathematical concept of addition and knows that the student likes fruits, the robot agent may generate a virtual scenario with a virtual agent picking fruits in an orchard and a plurality of apples and oranges thrown in space. The virtual scenario is generated to facilitate a conversation between the virtual agent and the student about adding the fruits picked from the orchard.
[0108] Each virtual agent may be configured with a sub-dialogue strategy determined adaptively based on the progress of the conversation. The sub-dialogue strategy is generated to achieve the specified purpose of the virtual agent and is used to control the activities of the virtual agent during the sub-dialogue. The virtual agent may also deploy additional virtual agents, each of which has its own specified purpose and activities for achieving the specified purpose. Once the specified purpose is achieved, each virtual agent will cease to exist, and control may then return to the agent (physical or virtual) that created it. In this way, the physical robot agent and one or more virtual agents may collaborate to achieve the overall goal of the conversation.
[0109] For environmental modeling, Figure 9A An exemplary high-level system diagram of a system 900 for adaptive environmental modeling via robot / sensor collaboration in accordance with an embodiment of the present disclosure is depicted. The depicted system 900 may be implemented in a main sensor such as a robot agent 810 and is configured to enable the robot agent 810 to coordinate the deployment / adjustment of robots / sensors for environmental modeling. The robot agent 810 may also be configured to be able to communicate with and control different deployable auxiliary robots / sensors (820, 830, and 850) in the dialogue scenario.
[0110] As Figure 9A shown, the system for environmental modeling includes a visual data analyzer 910, an object recognition unit 920, an audio data analyzer 930, a scene modeling unit 940, a data acquisition adjuster 950, a robot / sensor deployer 960, and a robot / sensor parameter adjuster 970. In operation, there may be a main robot (such as robot agent 810) that coordinates the deployment or adjustment of auxiliary robots / sensors based on what is observed and required. The purpose of deploying auxiliary robots / sensors may be to obtain the required information for environmental modeling. For example, if from the perspective of robot agent 810, an object may be occluded by an object in the scene, the robot agent may deploy sensors installed at appropriate positions in the scene so that the occluded area can be clearly captured for analysis. As previously mentioned, if robot agent 810 wishes to identify the user's identity through face recognition but does not have sensor data capturing the user's face, robot agent 810 may adjust the pose of the camera at an appropriate position to obtain the required data (e.g., adjust Figure 8A the parameters of camera 820-2 in Figure 8A to capture a side view of user 840) and / or deploy a drone with a camera facing the user (e.g.,
[0111] 850 in Figure 9B to obtain the required face information. Figure 8A The decision on which robot / sensor to deploy or adjust can be made based on observations made via the currently deployed existing robots / sensors. Data from the existing robots / sensors can be acquired and analyzed to understand the environment, which may be incomplete. Based on this understanding of the surrounding environment, the robot agent can determine whether additional information is needed and, if so, where to obtain it.
[0112] Then at 912, the visual / audio data analyzers 910 and 920 can analyze the received multimodal information. At 925, the object recognition unit 920 can detect objects present in the scene observable from the currently deployed cameras, based on, for example, the object detection model 915. Then at 935, the scene modeling unit 940 can perform scene modeling using such information according to, for example, the scene interpretation model 945. For example, the camera 820-1 can capture a scene in which the user 840 walks into a door in the scene, and the audio sensor 830 can capture the sound of opening the door. Based on the scene interpretation model 945, the analysis results of such multimodal data (e.g., detection of the door, opening the door, and the person walking into the door) can be interpreted. The robot agent 810 can obtain data from these deployed sensors through its connection with these deployed sensors, and control any connected robots / sensors by deploying and / or adjusting the parameters associated with any connected robots / sensors.
[0113] Based on the current modeling of the dialogue environment based on multimodal data, at 937, the robot agent can determine whether additional information is needed to understand the environment. For example, if some information is incomplete, e.g., a corner of the room is not observed or the face of the user 810 is not visible, then the robot agent 810 can proceed to 952 to analyze the known information to identify the missing information, and then at 965, the data acquisition adjuster 950 can determine based on the information stored in the robot / sensor deployment configuration 955 (see Figure 9A ) to understand what is available for deployment and which configuration parameters to use. With this knowledge, at 967, the data acquisition adjuster 950 can determine whether it is necessary to deploy auxiliary robots / sensors to obtain information from a certain space in the dialogue scene. For example, referring to Figure 8A , when the robot agent needs to see the face of the user walking into the dialogue scene and identifies that the user is facing a window that no camera can capture, the robot agent can then deploy a drone with an on-board camera to a position where the drone can use its camera to capture the user's face.
[0114] If the currently deployed robot / sensor can be adjusted to cover the desired space (at 967, the answer to the query of whether to deploy a robot / sensor is "no"), then at 985, the data acquisition adjuster 950 invokes the robot / sensor parameter adjuster 970 to calculate the required adjustments to the parameters of a specific robot / sensor. If at 967, the data acquisition adjuster 950 determines that an additional robot / sensor is to be deployed, then at 975, the data acquisition adjuster 950 invokes the robot / sensor deployer 960 to determine the robot / sensor to be deployed to obtain the required information. The determination is based on the information stored in the robot / sensor deployment configuration 955. Then at 985, the robot / sensor parameter adjuster 970 determines the parameters for deploying the new robot / sensor based on the information related to the configuration of such a robot / sensor. When an adjustment is made (a robot / sensor is deployed with a specific parameter configuration, or the parameters of the currently deployed robot / sensor are adjusted), the system 900 loops back to 905 to receive the multimodal information obtained from the adjusted robot / sensor. The process is repeated until the scene modeling is completed. In some embodiments, the master robot may also cancel the deployment of a robot / sensor. For example, if the purpose of the robot / sensor has been achieved (e.g., the required data has been provided), the deployment of the robot / sensor may be cancelled.
[0115] As described herein, in addition to adjusting robot / sensor parameters and / or dynamically deploying robot / sensors based on requirements, the master robot may also be configured to generate a virtual environment including virtual agents and / or virtual objects in a real conversation scenario to create an augmented reality conversation scenario. Figure 10A An exemplary hierarchical dialogue agent collaboration scheme in a virtual scene for enhancing dialogue reality according to an embodiment of the present disclosure is described. The virtual scene may include a virtual master agent and may have some virtual objects that can be presented together with the virtual master agent in a certain layout to form a virtual scene. The virtual master agent may be deployed with certain destinations to perform a specified dialogue task according to a dialogue strategy generated for the specified dialogue task to meet the purpose.
[0116] The virtual scene is presented within the dialogue scene such that it creates an augmented dialogue reality. In this augmented dialogue reality, an initial dialogue manager can hand over dialogue management to a virtual agent within the virtual scene or can cooperate in the coordinated assignment of tasks with the virtual agent. In some embodiments, the virtual agent can operate within a virtual scene having one or more virtual objects to perform a specified dialogue task, while the dialogue manager that creates the virtual scene can operate in a separate scene (i.e., the physical space) and time. In some cases, when the virtual agent operates within the virtual scene, the dialogue manager in the original dialogue scene can be paused. When the virtual agent achieves its purpose, the dialogue manager can be resumed to continue the original dialogue. In some embodiments, the dialogue manager can create multiple virtual scenes having virtual agents that operate within each virtual scene.
[0117] The visual master agent can further create additional secondary virtual scenes having secondary virtual agents, each of which can also have additional purposes regarding the specified task and corresponding specified dialogue strategies designed to achieve the additional purposes. In some embodiments, such secondary virtual scenes can replace the initial virtual scene having the primary virtual agent such that the initial virtual scene ceases to exist. Accordingly, the initial augmented dialogue reality dialogue is changed to create a modified augmented dialogue reality. In some embodiments, the secondary virtual scenes can coexist with the initial virtual scene. In such a case, the initial augmented dialogue reality is further enhanced such that both the initial virtual agent and the secondary virtual agents exist simultaneously (but operate in a coordinated manner, as described below).
[0118] Such secondary virtual agents can have some specified purposes that will be achieved via new specified operation strategies through new specified dialogue tasks. Once the specified purposes are achieved, each virtual scene (having the virtual agent and / or virtual objects) ceases to exist. In the case where the secondary virtual scene replaces the initial (or parent) virtual scene, the secondary virtual agent takes over control of the dialogue management while it is operating. When it exits operation and ceases to exist, the initial (or parent) virtual scene (that created the secondary virtual scene) will then resume its operation to continue the dialogue. In the case where the secondary virtual scene coexists with the initial virtual scene, each of the parent virtual agent and the secondary virtual agents can cooperate in performing independent yet complementary tasks.
[0119] In some embodiments, an auxiliary virtual agent can cease to exist when achieving its purpose, without affecting the parent virtual agent. In some embodiments, the exit of any one of the virtual agents may cause another virtual agent to also cease to exist. In some embodiments, there may be a seniority or priority order. For example, the initial dialogue manager that initiates the enhanced dialogue reality may have the highest seniority / priority, and each created virtual scene has a lower seniority / priority than the parent scene that created it. In some embodiments, the exit of a virtual agent with a higher seniority / priority may cause all other virtual agents with a lower seniority / priority to cease to exist, but not vice versa. For example, if a virtual agent exits a virtual scene, all the auxiliary virtual scenes created by that virtual agent will also cease to exist.
[0120] Figure 10B Illustrated is an exemplary composition of a virtual agent to be deployed in an enhanced dialogue reality according to an embodiment of the present disclosure. Each virtual agent can be represented by a role (the role can be dynamically determined during a conversation, based on, for example, user preferences represented in the information state 610). Such an agent can operate with one or more objects that appear with the virtual agent in a virtual scene, which can form the basis of a conversation. The virtual agent can also include a planner. The planner can include a subordinate dialogue manager for performing its designated dialogue tasks according to a designated dialogue strategy associated therewith, and a scheduler that can control the timing of the virtual agent entering or exiting a dialogue scene together with the subordinate dialogue manager. For example, an exit condition can specify the conditions under which it is considered that the virtual agent has achieved its purpose and thus it can cease to exist.
[0121] Figure 11A - 11E Examples of enhanced dialogue realities are provided in which an actual dialogue scene and a virtual dialogue scene are combined via the cooperation between some real and some virtual agents. Figure 11A Illustrated is an exemplary dialogue scene with a robotic agent 810 and a user 1110 according to an embodiment of the present disclosure. The robotic agent 810 here is capable of creating a virtual scene when needed to generate an enhanced dialogue reality. Figure 11B Illustrated is an enhanced dialogue reality 1100 according to an embodiment of the present disclosure, in which a robotic agent 810 interacting with a user 1110 in a dialogue scene creates a virtual scene with a virtual agent (avatar) 1120, and the virtual agent 1120 is intended to perform a designated dialogue task involving the user 1110. In this enhanced dialogue reality 1100, the dialogue scene and the virtual scene are mixed to create the enhanced dialogue reality 1100, where the virtual scene includes a virtual agent / avatar 1120, a set of virtual objects 1130, and a corresponding designated dialogue strategy that the virtual agent can use to perform the designated dialogue task to teach, for example, the user to count (e.g., how many children of a certain color there are).
[0122] A virtual agent / avatar 1120 can be created with the aim of attracting users to learn how to count in an interesting way and, for example, be invoked by a dialogue manager when it observes that the user (a child) loses attention or interest. When the user can count correctly according to a certain criterion (e.g., a certain number of times correctly), the virtual scenario with the virtual agent 1120 can cease to exist. In this example, there are four colors involved in the object, and each time the avatar 1120 can ask the user to count according to one color. An exit condition associated with the avatar can be specified (e.g., 80% of the time the counting is correct), and once the exit condition is met, the virtual avatar 1120 can be considered to have achieved its purpose, such that it can cease to exist together with the associated object 1130.
[0123] Figure 11C An enhanced dialogue reality 1150 with more than one virtual agent according to an embodiment of the present disclosure is shown, each virtual agent having an independent specified purpose. As Figure 11C shown, the robotic agent 810 creates an enhanced dialogue reality 1150 with two virtual agents 1120 and 1140. The virtual avatar 1120 corresponds to, for example, a first virtual agent created for a specified dialogue. According to an embodiment of the present disclosure, the virtual agent 1140, which is a clown in this example, is used to provide companionship to the user. In this example, in some embodiments, the robotic agent 810 can create the avatar 1120. The robotic agent 810 can create the virtual companion 1140 when it observes that the user 1110 is unhappy or looks depressed. The creation of the virtual companion 1140 may be to cheer up the user. In this case, the avatar 1120 and the companion 1140 can coexist in the enhanced dialogue reality 1150, but each can have a different purpose and corresponding dialogue tasks and strategies. For example, the avatar 1120 can be responsible for presenting virtual objects and asking questions, while the companion 1140 can be designed to talk to the user only when the avatar has finished asking questions, in order to, for example, encourage the user to answer or give hints, or do other things to help the user answer the questions.
[0124] In some embodiments, the virtual companion 1140 can be thrown or generated by the virtual avatar 1120. In this case, the virtual companion 1140 is created by the avatar 1120, which can create the virtual companion 1140 based on a similar observation that the user seems to need some cheering up when interacting with the avatar 1120. Figure 11D An embodiment is shown in which, according to an embodiment of the present disclosure, the virtual avatar 1120 generates the virtual companion 1140. Figure 11EShows another enhanced dialogue reality 1170 according to an embodiment of the present disclosure, where the robot agent 810 sequentially creates different virtual scenarios respectively responsible for specified tasks. In this example, the robot agent 810 can create multiple enhanced dialogue realities at different times. For example, after the virtual avatar 1120 achieves its purpose regarding counting and stops existing, the robot agent 810 creates another avatar 1180, and the purpose of the avatar 1180 may be to teach the user 1110 the concept of addition, which is further advanced than the counting performed by the avatar 1120. The subsequent avatar 1180 may be animated to pick different fruits in an orchard, present the picked fruits to the user, and ask the user to add different fruits to obtain the total. In the above example, it shows that the robot agent 810 can continuously create different virtual agents at different times, and each virtual agent can be responsible for achieving certain purposes. Additionally, these examples also show that the virtual agent can also create one or more auxiliary virtual scenarios with virtual agents based on the needs detected from the conversations it conducts in the enhanced dialogue reality to achieve the purposes associated with these needs.
[0125] Figure 12A Depicts an exemplary high-level system diagram of a virtual agent manager 1200 that operates in conjunction with a dialogue manager 650 to manage conversations in an enhanced dialogue reality. In a traditional dialogue system, the conversation is managed by the dialogue manager based on a pre-determined dialogue tree. According to the present disclosure, the dialogue manager 650 also manages the conversation with the user 680 based on the dialogue tree / policy 1212. Different from the traditional technology, the dialogue manager 650 also relies on the information from the information state 610 that is dynamically updated by the information state updater 1210 based on the conversation with the user 680, and coordinates or collaborates with the virtual agent manager 1200 to adaptively create an enhanced dialogue reality in order to improve the effectiveness of the conversation and enhance user participation.
[0126] As described herein with respect to Figure 7AAs described above, the information state 610 includes different types of information characterizing the dynamics, history, events, user preferences, and the estimated thinking of the different parties participating in the conversation of the conversation environment. Such information is continuously updated by the information state updater 1210 during the conversation and can be used to adaptively conduct the conversation based on what may be more effective for the users involved. Utilizing such rich information, the dialogue manager 650 can integrate information from different sources and different abstraction levels to make dialogue control decisions, such as including whether and when to include a virtual agent in the conversation scenario, the purpose to be achieved by the virtual agent, etc. For example, if the dialogue manager 650 can estimate that a young user is distracted during the conversation and the user likes avatars, such that deploying an avatar to continue the conversation with the user on a certain topic can help focus the user's attention, then the dialogue manager 650 can request the virtual agent manager 1200 to generate a virtual scenario with an avatar to conduct a conversation with the user on that topic.
[0127] Figure 12B FIG. 4 is a flowchart of an exemplary process of a dialogue manager 650 cooperating with a virtual agent manager to manage a conversation according to an embodiment of the present disclosure. When interacting with a user 680 in a conversation, each utterance from the user is processed to obtain a spoken language understanding (SLU) result (not shown). When at 1205, the dialogue manager 650 receives the SLU result obtained based on the user's utterance, at 1215, the dialogue manager 650 accesses relevant information about the user from the information state 610. Based on the SLU result and the information state, the dialogue manager 650 determines a dialogue strategy at 1225, and the dialogue strategy includes a determination of whether to generate a virtual scenario. If it is determined at 1235 that no virtual scenario will be generated, the dialogue manager 650 accesses the relevant dialogue tree in the dialogue tree / policy 1212 at 1265, determines a response to the user based on the dialogue tree at 1275, and then responds to the user based on the determined response at 1285. If a virtual scenario is to be generated, then at 1245, the dialogue manager 650 invokes the virtual agent manager 1200 to create a virtual scenario. The virtual scenario thus created will be presented or projected in the physical conversation scenario to form an augmented reality conversation scenario. As described herein, a virtual agent such as an avatar will appear in the virtual scenario to perform certain (virtual) activities to interact with the user, thereby conducting a sub-conversation on a specified topic.
[0128] In some embodiments, once the virtual agent manager is invoked to handle a sub - dialogue, the dialogue manager 650 can wait until the sub - dialogue ends when a release signal indicating that control of the dialogue returns to the dialogue manager 650 is received at 1255. The process then proceeds to step 1215 to continue the dialogue based on an assessment of the current information state. In some embodiments, the dialogue manager 650 can initiate a new dialogue, for example, determined based on how the virtual agent ends the sub - dialogue. For example, if the dialogue manager 650 invokes the virtual agent manager 1200 to teach the user addition, for example, via an interesting but virtual visual scenario, how the dialogue manager 650 continues the dialogue when the virtual agent releases itself may depend on the state of the sub - dialogue. If the sub - dialogue is successful, an exit condition based on the successful release of the virtual agent can be met, and in this case, the dialogue manager 650 can transition to the next topic in the dialogue. If the sub - dialogue is not successful, the virtual agent may exit without achieving the intended goal. In this case, the dialogue manager 650 can approach from some different strategies. In either case, once the virtual agent exits and the virtual scene is released, the dialogue manager 650 can return to 1215 to determine a strategy for continuing the session based on, for example, the information in the current information state.
[0129] To cooperate with the virtual agent manager 1200, the dialogue manager 650 can provide different operation parameters to the virtual agent manager 1200 based on, for example, the current state of the dialogue and the intended purpose of the dialogue. For example, the dialogue manager 650 can provide pointers to a portion of the currently used dialogue tree, the purpose to be achieved, the identity of the user, etc. In some embodiments, the virtual agent manager 1200 can also access information (e.g., information generated by the environment modeling system 900 as shown in Figure 9A - 9B to model the current dialogue environment stored in the environment modeling database 980) so that it can generate a virtual scene in a manner consistent with the current dialogue environment.
[0130] Utilizing information related to the request to generate a virtual scene from the dialogue manager 650, the virtual agent manager 1200 can access data in the information state 610, which, for example, characterizes the current dialogue and its intended goal, the user and preferences, the environment, the history, and the estimated mental state of the dialogue participants, etc. In this illustrated embodiment, the virtual agent manager 1200 includes a virtual scene determiner 1220, a virtual agent determiner 1230, a virtual object selector 1240, a dynamic policy generator 1250, an augmented reality launcher 1260, and a policy implementation controller 1270. Figure 12CFIG. 0 is a flowchart of an exemplary process of a virtual agent manager 1200 according to an embodiment of the present disclosure. When invoked, at 1207, a virtual scene determiner 1220 in the virtual agent manager 1200 accesses information from different sources to determine what the virtual scene is. Such information may include data stored in an information state 610, environmental modeling information of a conversation scene stored in a database 980, etc.
[0131] As described herein, the virtual scene may include virtual agents and / or some virtual objects to be thrown into the conversation scene. For example, as Figure 11B - 11E shown, the virtual agent may correspond to an avatar designed to perform some actions, such as throwing an object into space and asking the user to count or add. To create the virtual agent, at 1217, the virtual scene determiner 1220 may invoke a virtual agent determiner 1230 to determine an appropriate virtual agent based on information states 610 such as user preferences (liking avatars and colored objects), available models of virtual agents stored in a virtual agent database 1232. At 1227, it may also invoke a virtual object selector 1240 to select an object to be presented in the virtual scene based on available object models stored in a virtual object database 1222.
[0132] In some embodiments, the virtual agent may be determined based on, for example, user preferences stored in the information state 610. To control the virtual agent in the virtual scene, at 1237, a dynamic policy generator 1250 may dynamically generate a policy based on, for example, information passed from a conversation manager 650 (e.g., a part of a conversation tree) and various modeled sub-policies stored in a sub-policy database 1242. Such sub-policies for sub-conversations may be related to the overall conversation policy that controls the operation of the conversation manager 650. In some embodiments, such generated sub-policies may also be used to select virtual objects to be used in the virtual scene. Based on the virtual agent determined (by the virtual agent determiner 1230) and the virtual object selected (by the virtual object selector 1240), at 1247, the virtual scene determiner 1220 generates a virtual scene based on the virtual agent and the selected virtual object.
[0133] The sub - policies dynamically generated for the virtual agent can then be used by the policy operation controller 1270 to control the behavior of the virtual agent. To present a virtual scene in a dialogue scenario to create an augmented reality dialogue scenario, the virtual scene determiner 1220 sends information characterizing the virtual scene to the augmented reality launcher 1260. Then, at 1257, the augmented reality launcher 1260 launches the virtual scene in which there are virtual agents and virtual objects. Once launched, at 1267, an instance of the virtual scene is registered in the virtual scene instance registration table 1252. Any instance registered in the registration table 1252 can correspond to a live virtual scene with associated virtual agents and objects. The registration can also include the sub - policies associated with the virtual scene. At 1277, the policy enforcement controller 1270 can use the virtual scene registered in this way and the associated information to control the virtual characters / objects related to the virtual scene according to the sub - policies associated with the virtual scene. During the execution of the virtual agent based on the sub - policies, the policy enforcement controller 1270 controls the virtual agent to have a dialogue with the user 680. This interaction can continue to be monitored by the information state updater 1210. In this case, although the dialogue manager can standby without taking any action when the virtual agent intervenes in the dialogue, the information state 610 can be continuously monitored and updated.
[0134] When operating the virtual scene and controlling the performance of the virtual agent relative to the virtual object, the associated sub - policies can specify one or more exit conditions. At 1280, the one or more exit conditions can be checked to see if any of the exit conditions are met. If an exit condition is met (e.g., the user correctly answers a question from the virtual agent), then the policy enforcement controller 1270 proceeds to release the virtual agent and the scene by deleting the instance created for the virtual scene in the virtual scene instance registration table 1252, and then can send a release signal to the dialogue manager 650. At this time, the control of the dialogue can be transferred back from the virtual agent manager 1200 to the dialogue manager 650.
[0135] If none of the exit conditions are met, then at 1285, it can be further checked whether another virtual scene along with sub - policies having different virtual agents and associated objects needs to be created. In some cases, this might make sense. For example, if the user is simply uncooperative and seems bored, the current virtual agent can be programmed to activate a different virtual agent that can better enhance engagement. If this happens, then at 1287, the current virtual agent can generate a request with relevant information, at 1290, create a new instance of the virtual agent manager 1200, and then at 1295, call the new instance of the newly created virtual agent manager 1200, which will result in Figure 12CA process similar to the one depicted. Thus, the virtual scene and its management can be recursive. If none of the exit conditions associated with the virtual scene are met and there is no need to create another virtual scene, the virtual agent in the current virtual scene can be controlled to proceed to step 1207 to continue the process. In this way, the virtual agent can continue to execute the specified sub-strategy until it is released when any exit condition is met, thereby returning control to its creator, or the conversation is handed over to another virtual agent for some other specified purpose.
[0136] As described herein, creating a virtual environment can involve creating a virtual character such as an avatar and virtual objects, which are then presented in a recognized space. The determination of the virtual scene to be generated (performed by the virtual scene determiner 1220) can include the role of the virtual agent and the virtual objects to be used by the virtual agent. Once generated, the information related to the virtual scene is sent to the augmented reality launcher 1260 for presenting the virtual agent and objects in the conversation scene. In doing so, the augmented reality launcher 1260 may need to access various information to ensure proper presentation in accordance with different constraints. Figure 13A Illustrates an exemplary type of constraints that the augmented reality launcher 1260 may need to observe in accordance with an embodiment of the present disclosure. The constraints that can be used to control the presentation of the virtual scene in the real scene can include physical constraints, visual constraints, and / or semantic constraints.
[0137] Physical constraints can include the requirement to present an object on a tabletop or a virtual agent standing on the floor. Visual constraints can include the limitations implemented to make the visual scene meaningful. For example, the virtual character should be presented within the user's field of view and / or facing the user. Figure 13B An example shown in illustrates this, where the virtual scene 1310 is presented in the conversation scene in a way that is not within the field of view of the user 1320, so that the user 1320 cannot even see the presented scene. Similarly, the virtual objects to be presented to the user can be the basis of the conversation that the virtual character intends to have. For example, the objects can be presented in a way that serves the intended purpose, such as including that they also need to be presented within the user's field of view, and depending on the purpose of presenting such virtual objects, they may need to be arranged in a way that serves the said purpose. Figure 13C Another example is shown in, where the virtual avatar 1120 and multiple objects (flying birds) 1340 and 1350 are presented within the field of view of the user 1330.
[0138] Another exemplary type of constraint is a semantic constraint, which is a restriction on the rendering for a given known purpose of presenting a virtual scene. For example, if the purpose of throwing virtual objects into the scene is for the user to learn how to count, the virtual objects should not occlude each other. If the virtual objects are rendered to occlude each other, it makes it difficult for the user to count and for the robotic agent to evaluate whether the user actually knows how to count. Figure 13D An example is shown in Figure 13D , where several coins are rendered in a virtual scene 1360 such that they occlude each other, making it difficult for a person seeing the rendered coins to count. In contrast, given the known purpose (semantic constraint) of presenting the coins is to enable the user to count, another way of rendering the coins that complies with the semantic constraint is shown in Figure 13E Figure 13E , where there is no occlusion and counting is easy.
[0139] Thus, when creating a virtual scene in an augmented reality conversation scenario, the rendering of both virtual characters and objects may be subject to different constraints, and the constraints applied to rendering virtual characters may be different from those applied to rendering virtual objects. Figure 13F Figure 13F shows different constraints observed in rendering a virtual agent according to an embodiment of the present disclosure. As described herein, a virtual agent typically needs to be rendered within a certain distance (not too far and not too close) from the user in a certain pose, i.e., at a certain position, with a certain height / size and a certain orientation (e.g., facing the user) within the field of view of the user participating in the conversation, and should not occlude the virtual objects that the user needs to see.
[0140] Figure 13G Figure 13G shows different constraints observed in rendering a virtual object according to an embodiment of the present disclosure. Each virtual object may also need to be rendered within a certain distance (not too far and not too close) from the user in a certain pose, i.e., at a certain position, with a certain height / size within the field of view of the user participating in the conversation. Depending on the purpose of deploying the virtual objects, they may also need to comply with certain inter-object spatial relationship restrictions, such as no occlusion.
[0141] Certain dynamics during the conversation may cause the constraints to change over time. For example, when the user walks back and forth or changes pose. It may be necessary to continuously render the virtual scene relative to the changing constraints. For example, when the field of view changes, it may be necessary to re-render the virtual scene in a different space. In some embodiments, additional virtual characters may be introduced into the augmented reality scene, requiring the earlier deployed virtual characters (such as the virtual companion 1140 as shown in Figure 11C - 11E Figure 11C - 11E ) and objects to be rendered differently to accommodate the additional agent.
[0142] Figure 14FIG. 0 depicts an exemplary high-level system diagram of an augmented reality launcher 1260 for presenting a virtual scene in an actual conversation scenario, in accordance with an embodiment of the present disclosure. In the embodiment shown in the figure, the augmented reality launcher 1260 includes a user pose determiner 1410, a constraint generator 1420, an agent pose determiner 1430, an object pose determiner 1440, a visual scene generator 1460, a text-to-speech (TTS) unit 1450, an audio-video (A / V) synchronizer 1470, an augmented reality renderer 1480, and a virtual scene registration unit 1490. Figure 15 FIG. 2 is a flowchart of an exemplary process of an augmented reality launcher 1260 for combining a virtual scene with an actual conversation scenario, in accordance with an embodiment of the present disclosure.
[0143] In operation, at 1500, the agent pose determiner 1430 and the object pose determiner 1440 receive information related to a virtual agent and a virtual object from a virtual scene determiner 1220 (see Figure 12A ). Additionally, at 1510, the user pose determiner 1410 accesses information from the information state 610 and, at 1515, determines, for example, the pose of the user in order to estimate the field of view to be applied to the virtual scene. Further, in order to generate dynamic constraints for presenting the virtual scene, at 1520, the constraint generator 1420 receives information about the sub-strategy of the virtual agent and then, at 1525, generates different types of constraints (physical, visual, and semantic), taking into account the conversation scenario (from the information state 610), the intended purpose of the sub-conversation, and the field of view (estimated based on the current pose of the user).
[0144] The constraint generator 1420 obtains information from different sources in order to generate physical constraints 1422, visual constraints 1442, and semantic constraints 1432. For example, it may receive the estimated user pose information from 1410, information about the physical conversation scenario from the information state 610, and the current sub-conversation strategy associated with the virtual agent to be generated, in order to determine, accordingly, the constraints to be imposed on the physical position of the virtual agent / object, the field of view consistent with the pose / appearance of the user, and the limitations used when presenting the object, based on the purpose of the virtual agent and the physical conditions in the space in which the virtual object is to be presented.
[0145] Then at 1530, the agent pose determiner 1430 and the object pose determiner use the constraints thus generated, and also determine where and how to orient the virtual agent and the object based on the pose information of the user (from the user pose determiner 1410). With the separate and relative positions / orientations of the virtual agent / object determined, at 1535, the visual scene generator 1460 generates a virtual scene having such a virtual agent / object in a manner that satisfies the physical / visual / semantic constraints, to ensure that the virtual agent and the object will be presented in the field of view (e.g., estimated by the user pose determiner 1410 based on the estimated user pose) with appropriate size and distance and / or unoccluded. Meanwhile, when the virtual agent is to have a sub-conversation with the user according to the specified sub-strategy, i.e., during a conversation, at 1540, the TTS unit 1450 generates the voice of the virtual agent (e.g., determined by the sub-strategy). Then, at 1545, the A / V synchronizer 1470 can synchronize such voice with the generated visual scene. Then, the visual scene with synchronized vision / audio can be sent from the A / V synchronizer 1470 to the augmented reality renderer 1480, and at 1550, the augmented reality renderer 1480 presents the virtual scene in the physical conversation scene to form an augmented reality conversation scene. As described herein, then at 1555, such a virtual scene is registered in the virtual scene instance registration table 1252 by the virtual scene registration unit 1490.
[0146] Figure 16 is an illustrative schematic diagram of an exemplary mobile device architecture that can be used to implement a dedicated system for implementing the present disclosure according to various embodiments. In this example, the user equipment implementing the present disclosure corresponds to the mobile device 1600, including (but not limited to) a smart telephone, a tablet computer, a music player, a handheld game console, a global positioning system (GPS) receiver, and a wearable computing device (such as glasses, a watch, etc.), or any other form factor. The mobile device 1400 may include one or more central processing units (“CPUs”) 1640, one or more graphics processing units (“GPUs”) 1630, a display 1620, a memory 1660, a communication platform 1610 such as a wireless communication module, a storage device 1690, and one or more input / output (I / O) devices 1640. Any other suitable components, including (but not limited to) a system bus or a controller (not shown), may also be included in the mobile device 1600. As Figure 16As shown, a mobile operating system 1670 (such as iOS, Android, Windows Phone, etc.) and one or more applications 1680 can be loaded from a storage device 1690 into a memory 1660 for execution by a CPU 1640. The application 1680 can include a browser or any other suitable mobile application for managing a session system on the mobile device 1600. User interaction can be implemented via an I / O device 1640 and provided to an automated conversation partner via a network 120.
[0147] To implement the various modules, units, and their functions described in this disclosure, a computer hardware platform can be used as the hardware platform for one or more of the elements described herein. The hardware elements, operating systems, and programming languages of such a computer are conventional in nature, and it is assumed that those skilled in the art are familiar enough with them to adapt these techniques to the appropriate settings described herein. A computer with user interface elements can be used to implement a personal computer (PC) or other types of workstations or terminal devices. However, if appropriately programmed, the computer can also act as a server. It is believed that those skilled in the art are familiar with the structure, programming, and general operation of such computer devices, and thus the drawings should be self-explanatory.
[0148] Figure 17 is an illustrative schematic diagram of an exemplary computing device architecture that can be used to implement a dedicated system for realizing this disclosure in accordance with various embodiments. This dedicated system in combination with this disclosure has an example functional block diagram of a hardware platform that includes user interface elements. The computer can be a general-purpose computer or a special-purpose computer. Both can be used to implement the dedicated system of this disclosure. As described herein, the computer 1700 can be used to implement any component of a session or dialogue management system. For example, a session management system can be implemented on a computer such as the computer 1700 via its hardware, software program, firmware, or a combination thereof. Although for convenience only one such computer is shown, the computer functions related to the session management system described herein can be distributedly implemented on multiple similar platforms to distribute the processing load.
[0149] The computer 1700 includes, for example, a COM port 1750 connected to and from a network to which it is connected for facilitating data communication. The computer 1700 also includes a central processing unit (CPU) 1720 in the form of one or more processors for executing program instructions. Exemplary computer platforms include an internal communication bus 1710, various forms of program storage devices and data storage devices (e.g., a disk 1770, a read-only memory (ROM) 1730, or a random access memory (RAM) 1740) for various data files processed and / or transmitted by the computer 1700, and program instructions that may be executed by the CPU 1720. The computer 1700 also includes I / O components 1760 for supporting input / output streams between the computer and other components therein, such as user interface elements 1780. The computer 1700 may also receive programming and data via network communication.
[0150] Thus, as described above, aspects of the dialogue management method and / or other processes may be embodied in programming. The program aspects of the technology may be considered a "product" or "article of manufacture" generally in the form of executable code and / or associated data carried or embodied in some machine-readable medium. Tangible non-transitory "storage" type media include any or all memory or other storage devices that can readily provide storage for software programming for a computer, a processor, etc., or associated modules thereof, such as various semiconductor memories, tape drives, disk drives, etc.
[0151] All or part of the software can sometimes be transmitted via a network such as the Internet or various other telecommunications networks. For example, such communication can enable the loading of software from one computer or processor into, for example, another computer or processor related to session management. Thus, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices, over wired and fiber optic landline networks, and via various air links. Physical elements that convey such waves, such as wired or wireless links, fiber optic links, etc., can also be considered media that carry software. Unless restricted to tangible "storage" media, terms such as computer or machine "readable medium" as used herein refer to any medium that participates in providing instructions to a processor for execution.
[0152] Thus, a machine-readable medium can take many forms, including (but not limited to) tangible storage media, carrier media, or physical transmission media. Non-volatile storage media includes, for example, optical discs or magnetic disks, such as any storage device in any computer that can be used to implement the system or any of its components as shown in the attached drawings. Volatile storage media includes dynamic memory, such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables; copper wire and fiber optics, including the wires that form a bus within a computer system. Carrier transmission media can take the form of an electrical signal, an electromagnetic signal, or a sound wave or light wave, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example: floppy disks, hard disks, magnetic tape, any other magnetic media, CD-ROM, DVD or DVD-ROM, any other optical media, punched cards, paper tape, any other physical storage media with a pattern of holes, RAM, PROM, and EPROM, FLASH-EPROM, any other memory chip or memory card, a carrier wave that transports data or instructions, a cable or link that transports such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media may involve transporting one or more sequences of one or more instructions to a physical processor for execution.
[0153] Those skilled in the art will recognize that the present disclosure can be subject to various modifications and / or improvements. For example, although the implementation of the various components described above can be embodied in hardware devices, it can also be implemented as a pure software solution - for example, an installation on an existing server. Additionally, the fraud network detection techniques disclosed herein can be implemented as firmware, a firmware / software combination, a firmware / hardware combination, or a hardware / firmware / software combination.
[0154] Although the foregoing describes what is considered to constitute the present disclosure and / or other examples, it should be understood that various modifications can be made thereto, and the subject matter disclosed herein can be implemented in various forms and examples, and the present disclosure can be applied to numerous applications, only some of which are described herein. The following claims are intended to claim any and all applications, modifications, and variations within the true scope of the present disclosure.
Claims
1. A method for adaptive dialogue management across real and augmented reality, where the machine includes at least one processor, a memory, and a communication platform capable of connecting to a network, the method comprising: Receiving information related to a user machine dialogue in a dialogue scenario involving a user, the user machine dialogue being managed by a dialogue manager according to an initial dialogue strategy; Adapting the initial dialogue strategy based on the information to generate an updated dialogue strategy; wherein the adaptive dialogue strategy is to change the conversation strategy by observing the acceptance degree of a human conversationalist towards the dialogue; Based on the updated dialogue strategy, determining whether the user machine dialogue is to continue in an augmented dialogue reality with a virtual scenario presented in the dialogue scenario; and In response to the determination that the user machine dialogue is to continue in the augmented dialogue reality, activating a virtual agent manager to create an augmented dialogue reality and manage the user machine dialogue therein; Wherein the virtual scenario includes at least one of the following: A virtual agent created based on the information, for performing a specified dialogue task in the virtual scenario according to a specified dialogue strategy; and at least one virtual object to be presented in the virtual scenario; The constraints for controlling the presentation of the virtual scenario in the real scenario include semantic constraints, and the semantic constraints are the limitations for presentation under a known purpose of presenting the virtual scenario, to avoid mutual occlusion between virtual objects; During the session between the virtual agent and the user, the virtual agent is configured to conduct the session based on virtual content that is part of the virtual scenario; presenting the virtual agent and the virtual object as the virtual scenario in the dialogue scenario; and the virtual agent conducting a specified dialogue according to the virtual object presented in the virtual scenario.
2. The method according to claim 1, wherein the information related to the user includes at least one of the following: The result of spoken language understanding obtained based on the user's utterance; and Data from the information state, including at least one of the following: The user profile, A representation of the user's state; The dialogue history; The dialogue context representation, and A representation of the dialogue strategy associated with the user machine dialogue.
3. The method according to claim 1, wherein the information related to the dialogue scenario includes data on the surrounding environment of the dialogue scenario in one or more modalities.
4. The method according to claim 1, wherein, The specified dialogue task is determined based on the updated dialogue strategy, The specified dialogue strategy is generated according to the specified dialogue task, and The virtual agent is managed by the virtual agent manager when performing the specified dialogue task based on the specified dialogue strategy.
5. The method according to claim 4, wherein the at least one object is selected and / or presented based on the information; For the virtual agent to perform the specified dialogue task for the user in the augmented dialogue reality.
6. The method according to claim 1, further comprising pausing the dialogue manager from managing the user machine dialogue when the virtual agent manager is called.
7. The method according to claim 1, further comprising Receiving a release signal from the virtual agent manager indicating the completion of the user machine dialogue in the virtual scenario; Pausing the virtual agent manager when receiving the release signal; and When releasing the virtual agent manager, the dialogue manager is restored to manage the user-machine dialogue.
8. A system for adaptive dialogue management across real and augmented reality, comprising: A dialogue manager configured to receive information related to a user-machine dialogue in a dialogue scenario involving a user, where the user-machine dialogue is managed according to an initial dialogue strategy; An information state updater configured to update the information state based on the information to facilitate the adaptation of the initial dialogue strategy stored in the information state, thereby generating an updated dialogue strategy; wherein the adaptive dialogue strategy is to change the conversation strategy based on the observation of the acceptance degree of the human conversant towards the dialogue; And the dialogue manager is further configured to determine, based on the updated dialogue strategy, whether the user-machine dialogue is to continue in an augmented dialogue reality with a virtual scene presented in the dialogue scenario, and In response to the determination that the user-machine dialogue is to continue in the augmented dialogue reality, activate a virtual agent manager to create an augmented dialogue reality and manage the user-machine dialogue therein; Wherein the virtual agent manager includes at least one of the following: A virtual scene determiner configured to determine a virtual scene based on the information; A virtual agent determiner configured to determine a virtual agent that will perform a specified dialogue task in the virtual scene according to a specified dialogue strategy based on the information and the virtual scene; A virtual object selector configured to select at least one virtual object to be presented in the virtual scene; An augmented reality launcher configured to create an augmented dialogue reality by presenting the virtual scene in the dialogue scenario; The constraints for controlling the presentation of the virtual scene in the real scene include semantic constraints, which are the limitations for presentation under a given known purpose of presenting the virtual scene, to avoid mutual occlusion between virtual objects; During the conversation between the virtual agent and the user, the virtual agent is configured to conduct the conversation based on the virtual content that is part of the virtual scene; present the virtual agent and the virtual object as the virtual scene in the dialogue scenario; and the virtual agent conducts the specified dialogue according to the virtual object presented in the virtual scene.
9. The system according to claim 8, wherein the information related to the user includes at least one of the following: The speech understanding result obtained based on the user's utterance; and Data from the information state, including at least one of the following: A user profile, A representation of the user's state; A dialogue history; A dialogue context representation, and A representation of the dialogue strategy associated with the user-machine dialogue.
10. The system according to claim 8, wherein the information related to the dialogue scenario includes data on the surrounding environment of the dialogue scenario in one or more modalities.
11. The system according to claim 8, the virtual agent manager further includes at least one of the following: A dynamic policy generator configured to determine the specified dialogue strategy based on the specified dialogue task; and A policy enforcement controller configured to manage a virtual agent to perform the specified dialogue task based on the specified dialogue policy, where the specified dialogue task is determined based on an updated dialogue policy.
12. The system according to claim 11, wherein the at least one object is selected and / or presented based on the information; and is used by the virtual agent to perform the specified dialogue task for the user in an augmented reality dialogue scenario.
13. The system according to claim 8, wherein when the virtual agent manager is invoked, the dialogue manager is paused from managing the user machine dialogue.
14. The system according to claim 8, wherein the dialogue manager is further configured to: receive a release signal from the virtual agent manager indicating completion of the user machine dialogue in the virtual scene; pause the virtual agent manager upon receiving the release signal; and resume managing the user machine dialogue when the virtual agent manager is released.
Citation Information
Patent Citations
Augmented reality virtual personal assistant for external representation
US20140310595A1
Virtual personification for augmented reality system
US20160343168A1
Optimizing dialogue policy decisions for digital assistants using implicit feedback
US20180329998A1