Systems, devices and methods of providing an animated virtual agent
The virtual agent system addresses inefficiencies in customer service by using multimodal inputs to detect emotions and generate personalized responses, enhancing interaction efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CLOUDCONSTABLE INC
- Filing Date
- 2025-11-03
- Publication Date
- 2026-05-07
AI Technical Summary
Conventional customer service systems, including customer service agents and virtual agents, are resource-intensive and inefficient, and current animated virtual agents lack the ability to generate accurate and customized responses based on user inputs such as emotional cues.
A virtual agent system that utilizes multimodal inputs, including audio and visual data, to detect and classify user emotional states, retrieve relevant data from a custom dataset, and generate personalized responses through retrieval augmented generation architecture, combining audio and visual components.
The system provides efficient and accurate query responses by reducing the need for follow-up questions, saving computational resources and improving user interaction through personalized, contextually relevant answers.
Smart Images

Figure CA2025051464_07052026_PF_FP_ABST
Abstract
Description
Title: SYSTEMS, DEVICES AND METHODS OF PROVIDING AN ANIMATED VIRTUAL AGENTCross-Reference to Related Applications
[0001] The present application claims priority to United States provisional Patent Application No. 63 / 715,459 entitled “Systems, Devices and Methods of Providing an Animated Virtual Agent” filed on November 1 , 2024, the entire contents of which are hereby incorporated by reference herein.Technical Field
[0002] This disclosure relates generally to virtual agents, and, more specifically, to systems, devices and methods of providing animated virtual agents.Background
[0003] Customer service, including support and sales inquiries, are typically addressed by human agents. For example, if a customer needs an answer to a query, the customer may direct their inquiry to a customer service or support agent in-person within a business premises or through a web-based or in-app chat service.
[0004] Conventional customer service is however resource and labor intensive, in addition to being inefficient. Customer service agents need to be trained to provide clear and concise responses to queries and do so in an efficient manner.
[0005] Animated virtual agents have been developed as an alternative to in-person agents. However, current virtual agents are unable to generate accurate and specific responses to queries in an efficient manner. Also, current animated virtual agents lack the ability to customize their responses based on other inputs from users, such as but not limited to physical and / or emotional queues.
[0006] Accordingly, there is a need for new and improved systems, device and methods of providing an animated virtual agent.19359953Summary
[0007] In accordance with a broad aspect, a computer-implemented method of providing a query response with a virtual agent is described herein, the computer- implemented method comprising: receiving, by a virtual agent system, a multimodal input from at least one sensor, the multimodal input including one or more of an audio input through a microphone and a visual input through a camera; presenting a virtual agent on an output component of the virtual agent system, the virtual agent being configured to interact with the user through audio outputted from a speaker and / or video outputted through a display device; receiving, by the virtual agent system, a query from a user; generating a query response based at least in part on context of the query and content of the query, the generating including: utilizing retrieval augmented generation architecture of the virtual agent system upon receipt of the query to retrieve custom data from a custom dataset stored on a database of the virtual agent system relevant to the content of the query, the custom data forming content of the query response; detecting and classifying an emotional state of the user by the virtual agent system based on the multimodal input data; and formatting the query response based on the emotional state of the user and the content of the response; and providing the query response with the virtual agent presented on a graphical user interface on the display device.
[0008] In at least one embodiment, the query response includes an audio component spoken by the virtual agent and a visual component presented on the graphical user interface, the visual component being based on the custom data.
[0009] In at least one embodiment, providing the query response includes presenting relevant content used by the retrieval augmented generation architecture in generating the response on a graphical user interface as the visual component.
[0010] In at least one embodiment, generating the query response includes generating at least a portion of the query response in a text-based format and converting the portion of the query response from the text-based format to an audio response.
[0011] In at least one embodiment, prior to generating the query response, the method includes transcribing the query from an audio format to a text format.29359953
[0012] In at least one embodiment, providing the query response includes presenting text of the audio component on the graphical user interface on the display device as the virtual agent provides the audio component.
[0013] In at least one embodiment, detecting and classifying the emotional state of the user by the virtual agent system is based on one or more emotional indicators present in the input data.
[0014] In at least one embodiment, the emotional indicators present in the input data include: valence, attention and / or confidence in the vision data; arousal, stress and / or excitement in audio data; intent and / or positivity / negativity in linguistic content and / or comfort and / or engagement environment in contextual sensor data.
[0015] A system for providing a response to a query with an animated virtual agent is also described herein, the system including: at least one processor; at least one memory coupled to the at least one processor; a database storing a custom dataset; one or more input devices configured to provide input received from a user to the processor; one or more output devices; and one or more non-transitory computer-readable media having stored therein computer-executable instructions that, when executed by the computing system, cause the computing system to: receive a multimodal input from at least one sensor, the multimodal input including one or more of an audio input through a microphone and a visual input through a camera; present a virtual agent on the one or more output devices, the virtual agent being configured to interact with the user through audio outputted from a speaker and / or video outputted through a display device; receive a query from a user; generate a query response based at least in part on context of the query and content of the query, generating the query including: utilizing retrieval augmented generation architecture of the virtual agent system upon receipt of the query to retrieve custom data from a custom dataset stored on a database of the virtual agent system relevant to the content of the query, the custom data forming content of the query response; detecting and classifying an emotional state of the user by the virtual agent system based on the multimodal input data; and formatting the query response based on the emotional state of the user and the content of the response; and providing the query response with the virtual agent presented on a graphical user interface on the display.39359953
[0016] One or more non-transitory computer-readable media comprising computerexecutable instructions that, when executed by a computing system, cause the computing system to perform operations including the steps of the methods described herein are also described herein.
[0017] These and other features and advantages of the present application will become apparent from the following detailed description taken together with the accompanying drawings. It should be understood, however, that the detailed description and the specific examples, while indicating preferred embodiments of the application, are given by way of illustration only, since various changes and modifications within the spirit and scope of the application will become apparent to those skilled in the art from this detailed description.Brief Description of the Drawings
[0018] For a better understanding of the various embodiments described herein, and to show more clearly how these various embodiments may be carried into effect, reference will be made, by way of example, to the accompanying drawings which show at least one example embodiment, and which are now described. The drawings are not intended to limit the scope of the teachings described herein.
[0019] FIG. 1 is a block diagram of a system for providing an animated virtual agent, according to at least one embodiment described herein.
[0020] FIG. 2 is a pictorial diagram of an example of the system of FIG. 1 .
[0021] FIG. 3 is a flow chart of a method of providing a response to a query with an animated virtual agent system, according to at least one embodiment described herein.
[0022] FIG. 4 is a picture of an example of an animated virtual agent system, according to at least one embodiment described herein.
[0023] Further aspects and features of the example embodiments described herein will appear from the following description taken together with the accompanying drawings.49359953Detailed Description
[0024] Numerous embodiments are described in this application, and are presented for illustrative purposes only. The described embodiments are not intended to be limiting in any sense. The invention is widely applicable to numerous embodiments, as is readily apparent from the disclosure herein. Those skilled in the art will recognize that the present invention may be practiced with modification and alteration without departing from the teachings disclosed herein. Although particular features of the present invention may be described with reference to one or more particular embodiments or figures, it should be understood that such features are not limited to usage in the one or more particular embodiments or figures with reference to which they are described.
[0025] The terms "an embodiment," "embodiment," "embodiments," "the embodiment," "the embodiments," "one or more embodiments," "some embodiments," and "one embodiment" mean "one or more (but not all) embodiments of the present invention(s)," unless expressly specified otherwise.
[0026] The terms "including," "comprising" and variations thereof mean "including but not limited to," unless expressly specified otherwise. A listing of items does not imply that any or all of the items are mutually exclusive, unless expressly specified otherwise. The terms "a," "an" and "the" mean "one or more," unless expressly specified otherwise.
[0027] As used herein and in the claims, two or more parts are said to be “coupled”, “connected”, “attached”, “joined”, “affixed”, or “fastened” where the parts are joined or operate together either directly or indirectly (i.e. , through one or more intermediate parts), so long as a link occurs. As used herein and in the claims, two or more parts are said to be “directly coupled”, “directly connected”, “directly attached”, “directly joined”, “directly affixed”, or “directly fastened” where the parts are connected in physical contact with each other. As used herein, two or more parts are said to be “rigidly coupled”, “rigidly connected”, “rigidly attached”, “rigidly joined”, “rigidly affixed”, or “rigidly fastened” where the parts are coupled so as to move as one while maintaining a constant orientation relative to each other. None of the terms “coupled”, “connected”, “attached”, “joined”, “affixed”, and “fastened” distinguish the manner in which two or more parts are joined together.59359953
[0028] Further, although method steps may be described (in the disclosure and I or in the claims) in a sequential order, such methods may be configured to work in alternate orders. In other words, any sequence or order of steps that may be described does not necessarily indicate a requirement that the steps be performed in that order. The steps of methods described herein may be performed in any order that is practical. Further, some steps may be performed simultaneously.
[0029] As used herein and in the claims, a first element is said to be ‘communicatively coupled to’ or ‘communicatively connected to’ or ‘connected in communication with’ a second element where the first element is configured to send or receive electronic signals (e.g. data) to or from the second element, and the second element is configured to receive or send the electronic signals from or to the first element. The communication may be wired (e.g., the first and second elements are connected by one or more data cables), or wireless (e.g., at least one of the first and second elements has a wireless transmitter, and at least the other of the first and second elements has a wireless receiver). The electronic signals may be analog or digital. The communication may be one-way or two- way. In some cases, the communication may conform to one or more standard protocols (e.g., SPI, l2C, Bluetooth™, or IEEE™ 802.11 ).
[0030] As used herein and in the claims, a group of elements are said to ‘collectively’ perform an act where that act is performed by any one of the elements in the group, or performed cooperatively by two or more (or all) elements in the group.
[0031] Some elements herein may be identified by a part number, which is composed of a base number followed by an alphabetical or subscript-numerical suffix (e.g., 112a, or 112i). Multiple elements herein may be identified by part numbers that share a base number in common and that differ by their suffixes (e.g., 112i, 1122, and 112s). All elements with a common base number may be referred to collectively or generically using the base number without a suffix (e.g., 112).
[0032] Typically, customer service inquiries, including customer support inquiries and sales inquiries, are serviced by live, human agents (e.g., customer service representatives, agents and / or personal assistants). Traditional customer service is, however, resource-intensive and inefficient since it requires employing and training live-69359953agents. In addition, these live-agents are also often tasked with responding to the same or similar inquiries. Wait times associated with traditional customer service is also often long due to the limited number of live-agents.
[0033] Virtual digital assistants and / or chatbots have been developed and used in various customer service applications to support live-agents. These alternatives are typically used in combination with live-agents to manage and balance the workload of live-agents and to respond to simple and / or frequently asked questions.
[0034] Herein, a physical, in-person artificial-intelligence platform that performs customer service and sales interactions through autonomous terminals or devices is described. The platform combines real-time perception, cognitive analysis, and adaptive communication to detect emotion, intent, and context and respond with personalized empathy. Unlike conventional kiosks or chatbots, the platform unifies emotional intelligence and predictive sales analytics within a single operational flow, delivering consistent, human-like engagement.
[0035] The system’s architecture inlcudes perceptual sensors, cognitive processing, decision and response generation, and integration layers. Sensors capture visual, auditory, and contextual inputs; the cognitive engine interprets emotion and intent; the response layer adapts dialogue and offers accordingly; and the integration layer connects to CRM or PCS systems while ensuring privacy compliance. Together, these layers form a closed loop that senses, understands, responds, and learns from every interaction. This modular design supports scalability across industries while maintaining ethical data handling and brand consistency.
[0036] At its essence, the systems described herein transforms customer interaction from transactional to empathic by embedding scientific emotion understanding into physical service interfaces. The systems fuse perception and cognition into a unified, learning system capable of real-time emotional adaptation and predictive persuasion. This not only improves commercial performance but represents a significant advancement in human-machine interaction design. The outcome is a scalable, ethical, emotionally intelligent platform.79359953
[0037] The systems are networked, intelligent, in-person customer interaction solutions that combine real-time perception, contextual awareness, and adaptive communication to deliver human-like service, engagement, and sales conversion. The systems may operate through autonomous customer interaction terminals (ACMs) installed in physical venues such as but not limited to banks, hotels, clinics, retail stores, and showrooms and the like.
[0038] More specifically, the embodiments described herein describe an animated virtual agent engaging in an in-person conversation with a consumer and / or client. The animated virtual agent receives input from the user, including but not limited to audio input, visual input and / or input via a user interface presented to the user, and uses artificial intelligence (Al) to develop a response. The animated virtual agent can interact with human users through interactions, e.g., voice interactions, using speech transcription and voice synthesis, optionally with automatically generated subtitles, that can simulate a human conversation. The conversation can be optionally, and when appropriate, supplemented with visual aids to help illustrate or explain the answer being given by the animated virtual agent, and / or to provide insight into the source for the answer. The animated virtual agent is presented on a user interface by software that is trained using Al and machine-learning algorithms to respond to open-ended questions, that may use one or various data sources to provide a response to a consumer query and that can learn from (or has been trained from) interactions with consumers to improve its responses. The responses of the animated virtual agent can be built using large language models (LLMs), i.e. , computational models that use deep learning techniques and large data sets to predict and generate natural language text. The described embodiments involve the consumer or client initiating or responding to the conversation with the animated virtual agent.
[0039] The conversations described herein can be any type of conversation that may involve one or more question(s) from a consumer and that involves an agent. For example, the conversation can include but is not limited to a customer service interaction, including a customer support interaction or a personal assistance interaction. A consumer can be any individual who is seeking a response to a question. For example, the89359953consumer can be a customer of a business or other organized entity, or any individual consuming information.
[0040] The described embodiments may involve the animated virtual agent determining a response to a query received from the consumer and / or client. The response is based in at least in part on the content and / or context of the query.
[0041] At least some of the embodiments described herein involve determining the context of a query based on the conversation between the consumer and the animated virtual agent. For example, the conversation between the consumer and the animated virtual agent may involve a context that may be automatically determined by the animated virtual agent such as, based on the language of the query.
[0042] When compared to existing systems and methods, the described embodiments can determine responses to queries in a more efficient manner. When the context of a query can be determined by the animated virtual agent, the embodiments described herein can enable the animated virtual agent to more efficiently respond to queries from the consumer, particularly when the queries omit necessary context to respond to the question. By reducing the number of follow-up questions that need to be generated and transmitted to the consumer, the embodiments described herein can save computational resources associated with generating messages and save network resources associated with transmitting responses to the consumer, in addition to reducing the time needed for a consumer to receive a response to their query. The embodiments described herein can also reduce the computational load of the device presenting the animated virtual agent, since the device can transmit fewer messages in response to follow-up questions.
[0043] The term “module” used herein may refer to a hardware processor including a Central Processing Unit (CPU), an Application-Specific Integrated Circuit (ASIC), an Application-Specific Instruction-Set Processor (ASIP), a Graphics Processing Unit (GPU), a Physics Processing Unit (PPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), a Controller, a Microcontroller unit, a Processor, a Microprocessor, an ARM, or the like, or any combination thereof.99359953
[0044] The term “machine learning” may be used to refer to a computational or statistical or mathematical model that is trained on classical ML modelling techniques with or without classical image processing. The “machine learning model” is trained over a set of data and using an algorithm that it may use to learn from the dataset.
[0045] The term “neural networks” may be used to refer to a model built using simple or complex Neural Networks using deep learning techniques and computer vision algorithms. Artificial intelligence model learns from the data and applies that learning to achieve specific pre-defined objectives.
[0046] The term “artificial intelligence” may be used to include both of machine learning and neural networks as well as other classical options such as natural language processing and expert systems.
[0047] The term “virtual agent” may be used to refer to a virtual assistant that is computer program or Al system designed to simulate human-like conversations with users. They are typically powered by artificial intelligence and natural language processing technologies. The virtual agent can accept user inputs in natural (spoken, human) language, generate appropriate responses, and perform specific tasks or provide information. They are often used in customer support, information retrieval, and other applications to provide automated and efficient conversational experiences.
[0048] Referring first to FIG. 1 , therein is shown a block diagram of a virtual agent system 100, such as an animated virtual agent system, that includes a virtual agent subsystem 110 in communication with one or more remote devices 102via a network 104. It should also be understood that in optional embodiments, one or more user devices 108 may also be included as shown in FIG. 1. Virtual agent subsystem 110 may also be in communication with a greater number of user devices 108.
[0049] Virtual agent subsystem 110 can communicate with the one or more remote devices 102 and / or the consumer device 108 over a wide geographic area via the network 104.
[0050] It should be understood that the virtual agent system 100 includes multiple software modules (e.g., of devices 102 and / or subsystem 110) and the software modules109359953may be flexibly deployed in a ‘hybrid cloud’ model. That is, each animated virtual agent subsystem 110 includes a display device 120, as described in greater detail below, that presents a virtual agent character on a physical screen near the user. The virtual agent subsystem 110 can make eye contact with the user during an interaction and requires various local sensors (such as but not limited to a camera and / or a microphone) so that it can see and hear the user. However, potentially all of the computing required to provide for the virtual agent to interact with the user could run on any device within the system 100, or in some combination of local devices that are attached to the display device 120, or connected via network 104 and running on one or more remote devices (e.g., servers) 102, either, for example, positioned in a backroom (e.g., within a same premises as the virtual agent subsystem 110) or hosted anywhere in a nearby or faraway data centre or private or public cloud environment. Communication protocols can provide for the devices and subsystems of the system 100 to communicate with each other. It at least some embodiments, local processing may be preferred, especially for high bandwidth sensory input, e.g., video.
[0051] Animated virtual agent subsystem 110 includes at least a processor 112, storage 114 and a communication component 116. Animated virtual agent system 110 110 can be implemented with more than one computer server distributed over a wide geographic area and connected via the network 104. The processor 112, storage 114 and the communication component(s) 116 may be combined into a fewer number of components or may be separated into further components.
[0052] The processor 112 can be implemented with any suitable processor, controller, digital signal processor, graphics processing unit, application specific integrated circuits (ASICs), and / or field programmable gate arrays (FPGAs) that can provide sufficient processing power for the configuration, purposes and requirements of animated virtual agent subsystem 110. The processor 112 can include more than one processor with each processor being configured to perform different dedicated tasks.
[0053] The communication component 116 can include any interface that enables the subsystem 110 to communicate with various devices and other systems. For example, the communication component(s) 116 can receive inputs (e.g., free-form natural language119359953voice, user-initiated selections, video and / or audio) from the one or more input devices (e.g., microphone, camera, touchscreen, etc.) and store the inputs in storage 114 and / or remote device 102. The processor 112 can then process the inputs according to the methods described herein.
[0054] To further elaborate, here are some detailed examples of what these inputs may include:
[0055] Touch input: The user may provide touch input using a touch sensitive screen, and potentially including an on screen keyboard or other touch target such as a displayed button.
[0056] Speech Input: Users may interact with the virtual agent using spoken language.
[0057] Visual Input: Users may show their face expressions, emotions, or gestures, to show their feelings and the virtual agent may interpret and respond to these inputs accordingly. For instance: A user might smile to indicate happiness or agreement, or a user might nod in agreement, or even as a specific gesture to provide a positive response to a direct question.
[0058] Questions and Inquiries: Users often seek information or answers to questions.
[0059] Location-Based Queries: Input may involve location-based requests or queries. For example: “Give me directions to the nearest parking lot.”
[0060] In some embodiments, while communicating with the user, the virtual agent may combine the multimodal inputs, such as text, speech, and visuals of the user to understand the user's query / request and provide more personalized response in visual form.
[0061] By way of an example, while communicating with the user, the virtual agent may analyze the following:
[0062] Facial Expressions: Users may use their facial expressions to convey emotions or reactions. For instance: frowning or a furrowed brow may signify confusion or dissatisfaction, or raising an eyebrow may signal curiosity or skepticism.129359953
[0063] Emotional Cues: Visual input as well as audio-based input and other forms of input help understand the emotional cues of the user. For example: if the user appears sad or teary-eyed, the agent can respond with empathy and offer comforting words, or the user with an excited expression can prompt the agent to respond with enthusiasm.
[0064] Gestures: Users may use hand gestures or body language to communicate non-verbally. For example: pointing at an object in the environment, indicating interest or a question about that object, or waving a hand to get the agent's attention or to say goodbye. In another example, a user might wave or hold up a hand to indicate that the agent should stop talking, i.e. , to be able to interrupt the agent while it is talking.
[0065] Visual Cues: Users may provide visual cues by showing specific objects or scenes through their device's camera. For instance: displaying a broken appliance and asking for help in identifying the issue or sharing a photo of a product they want to purchase and asking for reviews or price information.
[0066] Facial Features: Detailed analysis of facial features, such as eye movements or the position of the mouth, may help gauge the user's emotional state or level of engagement. For example: dilated pupils may indicate excitement or interest, or avoiding eye contact might suggest shyness or discomfort.
[0067] The virtual agent system 110 is equipped with computer vision algorithms and advanced multimodal response generation techniques, and may combine, and analyze visual inputs to understand the user's intent, interest / attention, or emotions. This enables virtual agent subsystem 110 to generate a response in a visual, personalized and contextually relevant manner, enhancing the overall user experience during interactions. For example, the computer vision algorithms and advanced multimodal response generation techniques may include running artificial intelligence (Al) inference models on edges or in hybrid cloud to perform face detection, head pose estimation, emotion estimation, person recognition, etc.
[0068] For example, the system 110 may include a cognitive engine 122. The cognitive engine 122 may have various core components that interact dynamically. For example, the cognitive engine 122 may include a perception module configured to receive input data from each input component 118, which may be a multimodal sensor(s) (e.g.,139359953visual, audio and / or proximity sensor) as noted above. The perception module may, in some embodiments, perform preprocessing, normalization and / or feature extraction on the input data.
[0069] After receiving input data from each input component 118, the cognitive engine 122 is configured to: detect and classify emotional states of the user based on the input data. The cognitive engine 122 executes algorithms for emotion recognition, intent determination, and contextual correlation.
[0070] For example, the system 110 may include an emotion recognition module of the cognitive engine 122 that processes visual and auditory features to determine emotional state categories such as interest, excitement, confusion, or frustration. In this case, the output of the emotion recognition module may be an emotion vector comprising one or more probability scores.
[0071] For example, with regards to vision data, feature examples that may be processed include but are not limited to facial landmarks, gaze direction, microexpression intensity and / or body posture.
[0072] With regards to audio data, feature examples that may be processed include but are not limited to pitch, prosody, rhythm, speech rate and / ore timbre.
[0073] With regards to linguistic content, feature examples that may be processed include but are not limited to word choice, sentiment and / or semantic context.
[0074] With regards to contextual sensors, feature examples that may be processed include but are not limited to proximity, ambient noise and / or lighting.
[0075] Detection and classification of emotional states of the user may be based on one or more emotional indicators present in the input data, including but not limited to vision input data and / or audio input data and / or linguistic content data and / or contextual sensor data.
[0076] Emotional indicators that may be gleaned from the vision data may include but are not limited to valence, attention and / or confidence.
[0077] Emotional indicators that may be gleaned from the audio data may include but are not limited to arousal, stress and / or excitement.149359953
[0078] Emotional indicators that may be gleaned from the linguistic content data may include but are not limited to intent and / or positivity / negativity.
[0079] Emotional indicators that may be gleaned from the contextual sensor data may include but are not limited to comfort and / or engagement environment.
[0080] All features are synchronized along a common time axis. Frame timestamps and speech segments are aligned so that facial expressions and vocal tones occurring at the same moment can be analyzed jointly.
[0081] The system applies a weighted combination or neural-network integration of modality outputs. For example:Emotion_Score = w1 *Vision_Model + w2*Audio_Model + w3*Linguistic_Model + w4*Context_Model
[0082] where Wi-w4are adaptive weights determined by reliability or signal quality. Alternative embodiments may use Bayesian fusion, ensemble learning, or transformer-based cross-attention to merge features.
[0083] Fusion output includes a confidence metric derived from variance between modalities. High agreement between visual and auditory cues increases certainty; conflicting signals trigger clarification dialogue or model recalibration.
[0084] As another example, the system 110 may include an intent recognition module of the cognitive engine 122 that performs natural-language and menu-selection analysis to infer the user’s goal, such as but not limited to requesting information, making a purchase, or seeking assistance.
[0085] As another example, the system 110 may include a context memory module of the cognitive engine 122 that retrieves previous interaction records or customer preference data from local or remote storage. Context data may include but is not limited to prior purchases, feedback, or demographic information.
[0086] For example, the system 110 may include a predictive opportunity module of the cognitive engine 122 that combines emotion, intent, and context data to compute a conversion likelihood score or service priority value. This score guides subsequent response generation and escalation decisions.159359953
[0087] System 110 also includes a response generator of the cognitive engine 122 that dynamically selects tone, phrasing, and content based on detected emotion and intent. The response generator of the cognitive engine 122 may include a dialogue adaptation engine that uses the cognitive output to select a suitable communication mode and tone. For example, if the emotion vector indicates anxiety, the engine chooses a reassuring voice and simplified explanation; if it indicates enthusiasm, the engine may present promotional content or upgrade offers.
[0088] The response generator of the cognitive engine 122 may also include a recommendation generator that accesses product or service databases to compose personalized suggestions.
[0089] The response generator of the cognitive engine 122 may also include an escalation controller that evaluates whether live human intervention is warranted based on conversion likelihood, complexity, or emotion intensity. If escalation is triggered, a notification containing interaction context is transmitted to staff via the network interface.
[0090] The response generator of the cognitive engine 122 may also include a feedback composer that produces multimodal output — spoken dialogue, on-screen text, or animated graphics — and updates the analytics log with engagement data.
[0091] System 110 also includes a predictive decision module of the cognitive engine 122 that identifies purchase or engagement opportunities.
[0092] System 110 optionally includes a data interface connecting to external business systems, such as but not limited to a CRM, a PCS, or a scheduling system.
[0093] System 110 may optionally include a privacy controller for anonymization and consent management.
[0094] The ability to understand and respond to users across various modalities while considering individual user characteristics and emotions is a key differentiator and advantage of the present disclosure in the field of personalized multimodal response generation.
[0095] Storage 114 can include RAM, ROM, one or more hard drives, one or more flash drives or some other suitable non-volatile data storage elements (e.g., computer169359953readable mediums) such as disk drives. Storage 114 can store software modules including computer executable instructions to perform processing for the functions and methods described below. Storage 114 can also include one or more databases for storing inputs received by the animated virtual agent subsystem 110, query contexts determined by the animated virtual agent subsystem 110, transcripts of live conversations, summaries of live conversations, and any other information related to live conversations conducted by the animated virtual agent subsystem 110.
[0096] The communication component 116 can include an interface to component via one or more of an Internet, Local Area Network (LAN), Ethernet, Firewire, modem, fiber, or digital subscriber line connection. Various combinations of these elements may be incorporated within the communication component 116. The communication component 116 can enable the animated virtual agent subsystem 110 to communicate with the remote device(s) 102 and the device(s) 108 via the network 104 and / or can enable the animated virtual agent subsystem 110 to communicate with.
[0097] Input component, or sensor, 118 may be a multimodal sensor, for example any one or more of a microphone, camera and touchscreen for receiving both visual and auditory input(s) from a user. Additionally, the sensors may include proximity, biometric and / or gesture sensors.
[0098] Input component 118 can include any device for entering information into virtual agent subsystem 110. For example, input component 118 can be a keyboard, key pad, cursor-control device, touch-screen, camera, or microphone. Optionally, a gesturebased pointer input might be provided, using a camera and finger or hand tracking subsystem. Also, input might be optionally provided by attached card reader and / or document scanner. For example, a card reader might be used for reading magnetic strip cards or proximity cards, for example, health cards, credit / debit cards, access cards. Further, a document scanner might be used for example scanning identity documents such as a passport or a driver’s license. Also a barcode or QR code reader might be provided. In another example, one or more sensors could be provided that provide data regarding an environment of the subsystem 110 (e.g., temperature sensor) or measure a diagnostic of the user. For example, in some embodiments, the subsystem may be configured to179359953receive input from a thermometer, a heart rate monitor, a blood pressure cuff or the like to provide health information of the user to the system 100.
[0099] Input component 118 can also include input ports and wireless radios (e.g., Bluetooth®, or 802.11x) for making wired and wireless connections to external devices. It should be understood that although a single input component 118 is shown in FIG. 1 , the virtual agent subsystem 110 may include various input components 118 of several types.
[0100] Output component 120 can include any type of device for delivering output from the virtual agent subsystem 110, such as speakers and / or display devices, for example. In at least one embodiment, output component 120 includes one or more of output ports and wireless radios (e.g., Bluetooth®, or 802.11x) for making wired and wireless connections to external devices. Output component 120 may include a display device that can include any type of device for presenting visual information such as a live conversation. For example, the display device can be a computer monitor, a flat-screen display, a projector or a display panel. In at least one embodiment, the display device may be mounted on a cart or otherwise mobile component to provide for ease of movement of the virtual agent subsystem 110. Optionally, the display device might be mounted on an automatic, electrically activated lift, which would be controlled by the animated agent system so as to automatically match the height of the display device to the user’s head and reach, for example, in the case when the user is a child or uses an assistive device such as a wheelchair.
[0101] The remote device 102 can store data similar to that of the storage 114. The remote device 102 can, in some embodiments, be used to store data that is less frequently used and / or older data. In some embodiments, the remote device 102 can include a third-party data storage that stores past conversations. In some embodiments, the remote device 102 is a cloud storage server. The data stored in the remote device 102 can be retrieved by the virtual agent subsystem 110 via the network 104. In some embodiments, the virtual agent subsystem 110 only stores and retrieves data stored in the storage 114 and the virtual agent subsystem 110 is not in communication with a189359953remote device 102. In other embodiments, the one, two, three, four or more than four remote devices 102 may be included in system 100.
[0102] The user device 108, when present, can include a processor and memory, and may be an electronic tablet device, a personal computer, workstation, server, portable computer, mobile device, personal digital assistant, laptop, smart phone, an interactive television, video display terminals, gaming consoles, and portable electronic devices, any combination of these or any other device that can receive inputs from a user and that includes a display for displaying visual information to the user.
[0103] The network 104 can include any network capable of carrying data, including the Internet, Ethernet, plain old telephone service (POTS) line, public switch telephone network (PSTN), integrated services digital network (ISDN), digital subscriber line (DSL), coaxial cable, fiber optics, satellite, mobile, wireless (e.g. Wi-Fi, WiMAX), SS7 signaling network, fixed line, local area network, wide area network, and others, including any combination of these, capable of interfacing with, and enabling communication between, the subsystem 110, the remote device(s) 102, and the consumer device 108.
[0104] In at least one embodiment, the one or more remote devices 102 may be connected to the animated virtual agent subsystem 110 by a using secure web socket connection. The web socket connection may provide for modularity in the system 100.
[0105] Referring next to FIG. 2, shown therein is pictorial diagram of an example of the system 100 of FIG. 1 . Although FIG. 2 shows some specific branding as examples of services that can be used at the remote devices 102, it should be understood that within system 100, at any remote device 102 or at the subsystem 100, any Large Language Model or other conversational model, developed using generative Al techniques, could be used to power the conversational output of the virtual agent.
[0106]
[0107] Referring next to FIG. 3, FIG. 3 shows a flowchart illustrating a method 300 of conducting a conversation with a consumer by an animated virtual agent, in accordance with at least one embodiment. Method 300 can be implemented by the animated virtual agent subsystem 110. The flowchart shown in FIG. 3 illustrates the steps199359953of method 300 organized in a particular order. However, method 300 is not limited to the steps ordered as shown, and in some embodiments of method 300 some steps are practiced in a different order and some steps are practiced simultaneously.
[0108] At 310, one or more of an audio input through a microphone (e.g., input component 118) and a visual input through the camera (e.g., input component 118) is received by the animated virtual agent subsystem 110. In that manner, either an audio input may be received through the microphone, or visual input is received through the camera or both the audio input and the visual input are received in a multimodal manner, through the microphone and the camera, respectively. The one or more of the audio input and the visual input may be received at animated virtual agent subsystem 110 by the local processor 112.
[0109] According to various embodiments, the visual input may comprise any one of more of a gestural input from the user, a facial image of the user and an image of an object. For example, the visual input may be received when the user or the object comes within a predetermined distance of the camera. In this case, the system might be configured to automatically greet and welcome the user. For example, the subsystem 110 may include one or more proximity sensors that detect a presence of a user in front of or near display device 120. Additionally, a camera 118 may be used to detect a presence of a user in front of or near the display device 120.
[0110] Similarly, the audio input may be received when the user tries to communicate something to the microphone. In that manner it is envisaged that the user may start interacting with the animated virtual agent subsystem 110 by speaking or by playing an audio or by making various kinds of movements or just by changing facial expression or showing an object to the camera.
[0111] In another embodiment, input component 118 may be a full vision, audio and linguistic sensor fusion that provides input for a psychologically grounded understanding of the user. Fusion of these modalities yields a complete representation of emotion and reinforces the perception of empathy through synchronized verbal and non-verbal feedback. Optional variations include ARA / R and robotic implementations that preserve multimodal sensing. By aligning and integrating vision, audio, and linguistic209359953features, the platform may achieve near-human emotional recognition accuracy and robust operation under varied conditions.
[0112] At step 320, one or more animated virtual agents is shown on the output component 120 (e.g., display device). The one or more animated virtual agents are configured to interact with the user through an audio outputted from output component 120 (e.g., a speaker). The animated virtual agent can be configured to recognize the audio input of the user, the audio input being any combination of one or more sentences, phrases, words, music, song or any other verbal message or instructions from the user in one or more languages spoken by the user. The animated virtual agent can also be configured to interact with the user through a video outputted from the output component 120 (e.g., display device). Further, the one or more avatars can be adapted / configured to interact with the user in the one or more of the user's spoken languages. The audio and the video outputs may happen alternately, individually or in tandem.
[0113] The one or more virtual agents are rendered in real-time rendering as a central module of the subsystem 110. The virtual agent system 100 uses real-time feedback, including, for example, an estimation of a location of the human user's face relative to the display device 120 presenting the virtual agent, to track the user's eyes and make approximate eye contact. In this manner, the virtual agent may give the appearance of looking at the user during the interaction.
[0114] In another example, the virtual agent may be configured to greet the user with a smile upon determining that a user is proximate to the display device 120.
[0115] In at least one embodiment, the virtual agent may, at times, give other nonverbal feedback to the user during the conversation to give the appearance of engagement and encourage the user to continue the dialogue.
[0116] At 330, the animated virtual agent subsystem 110 receives, from the user, a query. The query can identify a query context for the query and / or a question for the animated virtual agent subsystem 110 (i.e., the query context and question can be determined from the audio and / or video input). The query can also elicit an Al generated response from the animated virtual agent subsystem 110.219359953
[0117] The query may also be a query that is received at the GUI of the display device 120 and inputted by the user, for example by typing the query on a keyboard (e.g., word by word).
[0118] The query context can include any relevant information that can assist the animated virtual agent subsystem 110 in responding to the question, including information used for establishing the meaning of the question or for establishing the circumstances in which the question is posed. For example, the query context can identify a field, a subject area or subject matter, or a scope of the question.
[0119] At 340, the animated virtual agent subsystem 110 determines a query response based in part on the query context and the question. The content (i.e., substance) of the query response is determined based at least in part on the query context.
[0120] To determine the query response, the virtual agent may retrieve information related to an input received by the virtual agent. This input may come in various forms, including text, speech, and visual cues. The virtual agent may employ an Artificial Intelligence (Al) model, which may, in particular embodiments, be a Generative Al (GenAI) model. The GenAI model represents a cutting-edge approach to artificial intelligence and is capable of multifaceted operations.
[0121] In specific implementations, the GenAI model may include, but not limited to, a Language model (LLM) for text, a Vision Language model (VLM) for vision-text, a speech model for speech, and other relevant modules. This comprehensive GenAI model is designed to process and respond to multimodal inputs effectively, making it exceptionally versatile in understanding and interacting with users across different modalities such as text, vision, and speech. In some embodiments, the GenAI model may take the form of an ensemble model allowing for even greater adaptability and proficiency in handling diverse inputs and user interactions.
[0122] To further elaborate, before generating the response, the virtual agent thoroughly analyzes the user's input. This analysis includes understanding the content, context, intent, and sentiment conveyed by the user across various modalities, such as text, speech, and visual cues.229359953
[0123] The virtual agent may refer to its knowledge base, which can be an internal database or an external data source, to gather additional information relevant to the user's query or input. This information retrieval process helps ensure that the response is factually accurate and contextually rich. Based on the analysis of the user's input and the retrieved information, the virtual agent generates an initial response. This response may take the form of text, speech, visual elements, or a combination of these modalities, depending on the nature of the user's input and the design of the virtual agent.
[0124] The virtual agent considers the user's characteristics, preferences, and historical interactions to tailor the response accordingly. For example, if the user has previously expressed a preference for a formal tone, the initial response may be composed in a formal style.
[0125] In scenarios where the user's input conveys emotional cues, such as sadness or frustration, the virtual agent may incorporate emotional intelligence. This means that the response may be designed to acknowledge the user's emotions and provide empathetic or supportive language.
[0126] More generally, in developing a response to the user’s query, to incorporate emotional intelligence, the system may utilize multimodal sensing and analysis that includes: sensing using one or more of a camera, microphone, proximity, and optional biometric or gesture sensors; real-time emotional analysis that interprets facial expression, tone, posture, and engagement level; and environmental context detection that identifies lighting, noise, crowd level, and time-of-day cues.
[0127] In determining the response, the virtual agent subsystem 110 may utilize Retrieval Augmented Generation (RAG) architecture responsible for retrieving relevant information from a database accessible by the virtual agent system. In at least one embodiment, the database is populated with content that is relevant to the business hosting the virtual agent. For example, if the business hosting the virtual agent is a museum, the database that is used by RAG is populated with content that is either specific to the museum or, in the least, is selected by the museum for use in developing responses to queries. RAG retrieves relevant information by referring to the references stored in the database that are needed to answer the user's query. In at least one embodiment, the239359953retrieved information may be customized based on a conversation history stored in the memory of the virtual agent. In this manner, it can be said that the response that is determined and generated is done so using a curated knowledge base, or is generated using custom knowledge corpus.
[0128] In at least one embodiment, the database is populated by ingesting content provided by the business that is hosting the virtual agent. In most cases, the ingest of user content, such as from a website and / or some documents such as digital text, spreadsheets, slideshows, etc., and potentially including pictures, voice recordings, videos, etc., is an important source of knowledge for the virtual agent. This content is indexed as part of the ingest process, for search and retrieval based on user queries. This process is known as “Retrieval Augmented Generation” or RAG. Once the highest ranked content that matches the user query is found, it is displayed to the user in the system user interface (III), while the answer is generated and spoken by the virtual agent.
[0129] In at least one embodiment, the system 100 may include a web portal or similar subsystem that is available to ingest and manage the knowledge base. Owners of the subsystem 110 may choose to ingest additional information via the portal.
[0130] In at least one embodiment, the database storing the content provided by the business (i.e., customer content) may be stored in a remote device 102, which as is noted above, or may be located in a same geographic place as the animated virtual agent subsystem 110, or stored within virtual agent subsystem 110. In some cases where the content provided by the business is stored in a remote device 102, which generally is not located in a same geographic place as the animated virtual agent subsystem 110, there may be potential for latency in generating a response to the user’s query. If the remote device 102 is located in a same geographic place as the animated virtual agent subsystem 110 and stores the database storing the content provided by the business, round trip delays and latency of the virtual agent may be minimized.
[0131] Further to this, it should be understood that the customer content that is used by the system 100 to develop and provide responses to user queries is available both in a visually pleasing way, suitable for redisplay to the user, as well as in an indexed database for use in comparing against user queries.249359953
[0132] Generally, the large-language model noted above is instructed to generate a text-based response to a user’s query. Once the text-based response has been generated, the text-based response needs to be converted to an audio based response. The audio based response is generally in the same language as the input received from the user. Optionally, the text of the response can be shown as subtitles while the response is spoken by the virtual agent.
[0133] Once on text in the voice of the character, real-time lip synchronization occurs to synchronize the audio response with movement of the lips of the virtual agent that is presented on the display.
[0134] In at least one embodiment, the response that is generated can be supplemented by providing the user with a visual indication of the source or sources that were used in developing the response. For example, the virtual agent subsystem 110 may present the source where the answer was derived. For example, the source may be a URL presented on the GUI of the display device 120 and be available for the user to select to present the content of the URL. In another example, the content may be a map, a brochure.
[0135] In at least one embodiment, before determining the response to the query, the virtual agent subsystem may determine that more information is needed to answer the query. In this case, the virtual agent may ask questions of or obtain information or documentation from the user. For example, the virtual agent might, as part of a customer service process, need to obtain answers to a few questions from the user, as in a healthcare screening or customer satisfaction survey. In such a scenario, the input provided by the user may trigger the virtual agent to pose additional questions to the user. The virtual agent would speak the additional questions, optionally with subtitles displayed on the display device 120. Answers would then be obtained by the virtual agent subsystem 110, optionally while displaying potential answers, for example in the form of multiple choice or scale of 1 to 10, etc., on the display device 120. Furthermore, the virtual agent may also determine that additional information from the user is required, for example in the form of identity information or may request to verify documentation from259359953the user, such as but not limited to by requesting them to show photo identification, a coupon or ticket, or a membership card.
[0136] At 350, the animated virtual agent subsystem 110 providing the response to the query with the virtual agent.
[0137] Turning now to FIG. 4, shown therein is an example of a portion of a virtual agent subsystem 110 of FIG. 1 . More specifically, FIG. 4 shows an example of a display device 120 having a camera 118 mounted thereon. Display device 120 is mounted on a mobile component 124 for ease of movement of the display device 120. Mobile component 124 may be configured to raise and lower the display device 120, for example in response to a height of a user of the device.
[0138] The display device 120 is configured to present a GU1 125 generated by the virtual agent subsystem 110, including a virtual agent 126 presented on a first portion of the display device 120. Relevant content to the user’s query is displayed on a second portion of the display device 120. The presentation of the virtual agent 126 and the relevant content 128 may be simultaneous to provide for the user to interact with both the virtual agent 126 and the relevant content 128 simultaneously. In an alternative embodiment, the GUI may be split and presented across various display devices 120. Additionally, as noted above, the textual response of the virtual agent may also be presented as “open captioning” at the bottom of the display device 120 as the virtual agent 126 is speaking.
[0139] Virtual agent 126 is a character presented in three dimensions on a two- dimensional display device 120.
[0140] In at least one embodiment, the system may include an application programming interface (API) or other interface into a host system. For example, the API the may be developed to support, for example, an Al Investment Advisor role, in which an Al Advisor might help the human client review his / her investment portfolio and make changes.
[0141] While the above description provides examples of the embodiments, it will be appreciated that some features and / or functions of the described embodiments are269359953susceptible to modification without departing from the spirit and principles of operation of the described embodiments. Accordingly, what has been described above has been intended to be illustrative of the invention and non-limiting and it will be understood by persons skilled in the art that other variants and modifications may be made without departing from the scope of the invention as defined in the claims appended hereto. The scope of the claims should not be limited by the preferred embodiments and examples, but should be given the broadest interpretation consistent with the description as a whole.279359953
Claims
ClaimsWhat is claimed is:
1. A computer-implemented method of providing a query response with a virtual agent, the computer-implemented method comprising: receiving, by a virtual agent system, a multimodal input from at least one sensor, the multimodal input including one or more of an audio input through a microphone and a visual input through a camera; presenting a virtual agent on an output component of the virtual agent system, the virtual agent being configured to interact with the user through audio outputted from a speaker and / or video outputted through a display device; receiving, by the virtual agent system, a query from a user; generating a query response based at least in part on context of the query and content of the query, the generating including: utilizing retrieval augmented generation architecture of the virtual agent system upon receipt of the query to retrieve custom data from a custom dataset stored on a database of the virtual agent system relevant to the content of the query, the custom data forming content of the query response; detecting and classifying an emotional state of the user by the virtual agent system based on the multimodal input data; and formatting the query response based on the emotional state of the user and the content of the response; and providing the query response with the virtual agent presented on a graphical user interface on the display device.
2. The computer-implemented method of claim 1 , wherein the query response includes an audio component spoken by the virtual agent and a visual component289359953presented on the graphical user interface, the visual component being based on the custom data.
3. The computer-implemented method of claim 2, wherein providing the query response includes presenting relevant content used by the retrieval augmented generation architecture in generating the response on a graphical user interface as the visual component.
4. The computer-implemented method of claim 1 , wherein generating the query response includes generating at least a portion of the query response in a text-based format and converting the portion of the query response from the text-based format to an audio response.
5. The computer-implemented method of claim 1 , wherein, prior to generating the query response, the method includes transcribing the query from an audio format to a text format.
6. The computer-implemented method of claim 2, wherein providing the query response includes presenting text of the audio component on the graphical user interface on the display device as the virtual agent provides the audio component.
7. The computer-implemented method of claim 1 , wherein detecting and classifying the emotional state of the user by the virtual agent system is based on one or more emotional indicators present in the input data.
8. The computer-implemented method of claim 7, wherein the emotional indicators present in the input data include: valence, attention and / or confidence in the vision data; arousal, stress and / or excitement in audio data; intent and / or positivity / negativity in linguistic content and / or comfort and / or engagement environment in contextual sensor data.
9. A system for providing a response to a query with an animated virtual agent, the system including: at least one processor;299359953at least one memory coupled to the at least one processor; a database storing a custom dataset; one or more input devices configured to provide input received from a user to the processor; one or more output devices; and one or more non-transitory computer-readable media having stored therein computer-executable instructions that, when executed by the computing system, cause the computing system to: receive a multimodal input from at least one sensor, the multimodal input including one or more of an audio input through a microphone and a visual input through a camera; present a virtual agent on the one or more output devices, the virtual agent being configured to interact with the user through audio outputted from a speaker and / or video outputted through a display device; receive a query from a user; generate a query response based at least in part on context of the query and content of the query, generating the query including: utilizing retrieval augmented generation architecture of the virtual agent system upon receipt of the query to retrieve custom data from a custom dataset stored on a database of the virtual agent system relevant to the content of the query, the custom data forming content of the query response; detecting and classifying an emotional state of the user by the virtual agent system based on the multimodal input data; and formatting the query response based on the emotional state of the user and the content of the response; and309359953providing the query response with the virtual agent presented on a graphical user interface on the display.
10. The system of claim 9, wherein the query response includes an audio component spoken by the virtual agent and a visual component presented on the graphical user interface, the visual component being based on the custom data.11 . The system of claim 10, wherein providing the query response includes presenting relevant content used by the retrieval augmented generation architecture in generating the response on a graphical user interface as the visual component.
12. The system of claim 9, wherein generating the query response includes generating at least a portion of the query response in a text-based format and converting the portion of the query response from the text-based format to an audio response.
13. The system of claim 9, wherein, prior to generating the query response, the method includes transcribing the query from an audio format to a text format.
14. The system of claim 9, wherein providing the query response includes presenting text of the audio component on the graphical user interface on the display device as the virtual agent provides the audio component.
15. The system of claim 9, wherein detecting and classifying the emotional state of the user by the virtual agent system is based on one or more emotional indicators present in the input data.
16. The system of claim 15, wherein the emotional indicators present in the input data include: valence, attention and / or confidence in the vision data; arousal, stress and / or excitement in audio data; intent and / or positivity / negativity in linguistic content and / or comfort and / or engagement environment in contextual sensor data.
17. One or more non-transitory computer-readable media comprising computerexecutable instructions that, when executed by a computing system, cause the computing system to perform operations comprising:319359953receiving, by a virtual agent system, a multimodal input from at least one sensor, the multimodal input including one or more of an audio input through a microphone and a visual input through a camera; presenting a virtual agent on an output component of the virtual agent system, the virtual agent being configured to interact with the user through audio outputted from a speaker and / or video outputted through a display device; receiving, by the virtual agent system, a query from a user; generating a query response based at least in part on context of the query and content of the query, the generating including: utilizing retrieval augmented generation architecture of the virtual agent system upon receipt of the query to retrieve custom data from a custom dataset stored on a database of the virtual agent system relevant to the content of the query, the custom data forming content of the query response; detecting and classifying an emotional state of the user by the virtual agent system based on the multimodal input data; and formatting the query response based on the emotional state of the user and the content of the response; and providing the query response with the virtual agent presented on a graphical user interface on the display device.
18. The one or more non-transitory computer-readable media of claim 17, wherein the query response includes an audio component spoken by the virtual agent and a visual component presented on the graphical user interface, the visual component being based on the custom data.
19. The one or more non-transitory computer-readable media of claim 18, wherein providing the query response includes presenting relevant content used by the retrieval329359953augmented generation architecture in generating the response on a graphical user interface as the visual component.
20. The one or more non-transitory computer-readable media of claim 17, wherein generating the query response includes generating at least a portion of the query response in a text-based format and converting the portion of the query response from the text-based format to an audio response.
21. The one or more non-transitory computer-readable media of claim 17, wherein, prior to generating the query response, the method includes transcribing the query from an audio format to a text format.
22. The one or more non-transitory computer-readable media of claim 18, wherein providing the query response includes presenting text of the audio component on the graphical user interface on the display device as the virtual agent provides the audio component.
23. The one or more non-transitory computer-readable media of claim 17, wherein detecting and classifying the emotional state of the user by the virtual agent system is based on one or more emotional indicators present in the input data.
24. The one or more non-transitory computer-readable media of claim 23, wherein the emotional indicators present in the input data include: valence, attention and / or confidence in the vision data; arousal, stress and / or excitement in audio data; intent and / or positivity / negativity in linguistic content and / or comfort and / or engagement environment in contextual sensor data339359953
Citation Information
Patent Citations
Electronic personal interactive device
US20170221484A1
Empathetic personal virtual digital assistant
US20190266999A1