Machine interaction

Embodied artificial agents facilitate improved human-computer interaction by simulating human-like behaviors and interactions within a shared environment, addressing the limitations of existing technologies and enhancing user engagement.

JP2025160239APending Publication Date: 2025-10-22SOUL MACHINES LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025117140
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-07-03
Filing Date
2025-07-11
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Existing human-computer interaction technologies lack the ability to create a seamless and engaging interaction experience using autonomous agents that mimic human-like communication and behavior, leading to perceived technology barriers and reduced user engagement.

Method used

The implementation of embodied artificial agents that interact with digital content within a simulated environment, allowing them to perceive and manipulate shared content through a continuous feedback loop between real and virtual environments, using neurobehavioral models and computer vision to simulate human-like interactions.

Benefits of technology

Enhances human-machine interaction by reducing technology barriers and increasing user engagement through realistic and dynamic agent behaviors, enabling flexible and scalable interaction with digital content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025160239000001_ABST
    Figure 2025160239000001_ABST
Patent Text Reader

Abstract

To improve human-machine interaction (human-computer interaction) or to at least provide the public or industry with a useful choice.SOLUTION: Interaction with a computer is provided via an autonomous virtual embodied agent. The computer outputs digital content, which includes any content that exists in the form of digital data and is representable to a user. A subset or all of digital content is configured as shared digital content which is representable to both the user and to the agent.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Computer science techniques are used to facilitate interaction between humans and machines, and more particularly, but not exclusively, use embodied agents to facilitate human-computer interaction. [Background technology]

[0002] Human-Computer Interaction (HCI) is a field of computer science that aims to improve the interaction between humans and computers using techniques from computer graphics, operating systems, software programming, and cognitive science. Autonomous agents can improve human-computer interaction by assisting human users in operating computers. Because autonomous agents are capable of flexible and / or autonomous operation in an environment (which may be virtual or real), autonomous interface agents can be viewed as robots whose sensors and effectors interact with the input and output capabilities of a computer interface. Autonomous interface agents can interact with a computer interface in parallel with the user, with or without user direction or collaboration.

[0003] EP1444687B1, entitled "Method for managing mixed-initiative human-machine dialogues based on interactive speech", discloses a method for managing mixed-initiative human-machine dialogues based on speech interaction, and EP2610724A1, entitled "System and method for online user assistance", discloses the display of virtual animated characters for online user assistance.

[0004] It may be desirable for autonomous interface agents to use artificial intelligence techniques to intelligently process information, interact, and present themselves in a more human-like manner. Human users may find it easier, quicker, and / or more engaging to interact with a computer using human-like communication methods, including body language and vocal tone. Second, increasing the realism of an agent's actions and responses may reduce perceived technology barriers, such as the uncanny valley effect. Summary of the Invention [Problem to be solved by the invention]

[0005] The purpose of this invention is to improve human-machine interaction (human-computer interaction), or at least to provide the public or industry with a useful option. [Means for solving the problem]

[0006] In one aspect, a method is provided for displaying an interaction between an embodied artificial agent and digital content on an end-user display device of an electronic computing device, the method comprising: creating an agent virtual environment having virtual environment coordinates; simulating digital content within the agent virtual environment; simulating an embodied artificial agent within the agent virtual environment; enabling the embodied artificial agent to interact with the simulated digital content; and displaying the interaction between the embodied artificial agent and the digital content on the end-user display device.

[0007] In another aspect, a method is provided for interacting with digital content on an electronic computing device through an embodied artificial agent, the method including the steps of displaying the digital content to a user on a user interface on the electronic computing device, creating an agent virtual environment having virtual environment coordinates, simulating the digital content in the agent virtual environment, simulating an embodied artificial agent in the agent virtual environment, enabling the embodied artificial agent to interact with the simulated digital content, translating the interaction into an actuation or manipulation of the digital content on the user interface, and displaying the interaction by overlaying the embodied artificial agent on the digital content and displaying the digital content and the overlaid embodied artificial agent on the user interface.

[0008] In another aspect, a system for facilitating interaction with an electronic computing device is provided, the system including at least one processor device; at least one memory device in communication with the at least one processor; an agent simulator module executable by the processor and configured to simulate the embodied agent; an interaction module executable by the processor arranged to transform the digital content into conceptual objects perceivable by the embodied agent and enable the embodied agent to interact with the digital content; and a rendering module executable by the processor configured to render the digital content, the embodied agent, and interaction of the embodied agent with the digital content.

[0009] In another aspect, an enforcement agent is provided that is located within a virtual environment created on an electronic computing device, the enforcement agent being programmed to receive input from the real-world environment, receive input from the virtual environment, and act in response to input from the real-world environment and the virtual environment, wherein the input received from the real-world environment and the virtual environment is received via a continuous feedback loop from both the real-world environment and the virtual environment. [Brief explanation of the drawings]

[0010] [Figure 1] Demonstrates interactions between user agents in a shared interaction environment. [Figure 2] Shows the user-agent interaction from the agent's perspective. [Figure 3] A schematic diagram of human-computer interaction is shown. [Figure 4] Shows the user agent's interaction with the digital content behind the agent [Figure 5] Showing a system diagram for human-computer interaction [Figure 6] Show process diagram for human-computer interaction [Figure 7] Shows a flow diagram of the feedback loop between the agent and the environment [Figure 8] 1 illustrates user interaction in a virtual reality context. [Figure 9] Shows a schematic diagram of a system for human-computer interaction in a virtual reality context [Figure 10] 1 illustrates user interaction in an augmented reality context. [Figure 11] 1 shows a schematic diagram of a system for human-computer interaction using mobile applications. [Figure 12] 1 shows screenshots of human-computer interaction via an agent. [Figure 13] 1 shows screenshots of human-computer interaction via an agent. [Figure 14] 1 shows screenshots of human-computer interaction via an agent. [Figure 15] 1 shows screenshots of human-computer interaction via an agent. [Figure 16] 1 shows screenshots of human-computer interaction via an agent. [Figure 17] 1 shows a schematic diagram of an agent's interaction with digital content using a computer vision system. [Figure 18] 10 illustrates digital content associated with embedded actions and / or agent-recognizable locators. [Figure 19] 1 shows an example of how a user expression captured by a camera is displayed on an end user display. [Figure 20] An example of turn-taking is shown below. [Figure 21] 10 shows a user interface for setting salience values. [Figure 22] Demonstrates an attention system for multiple visual sources. [Figure 23] 1 shows a class diagram of an example of user agent interaction. DETAILED DESCRIPTION OF THE INVENTION

[0011] [Embodiment 1] Interaction with the computer is provided through autonomous virtual embodied agents (hereafter referred to as "agents"). The computer outputs digital content, including any content that exists in the form of digital data and is presentable to the user. A subset, or all, of the digital content constitutes shared digital content presentable to both the user and the agent. Figure 1 shows a user 3 and agent 1 perceiving shared digital content 5. The shared digital content 5 can be manipulated and / or perceived (interacted with) by both user 3 and agent 1. A shared environment 6 contains all shared digital content 5 and is perceivable by both agent 1 and user 3. Agent 1 can "physically" (using its embodiment) interact with the shared digital content 5 within its agent virtual environment 8.

[0012] Agent 1 resides in an agent virtual environment 8, and shared digital content 5, which perceives the agent virtual environment 8, is displayed to Agent 1 within the agent virtual environment (AVE) 8. A plane 14 of the agent virtual environment 8 is mapped to an area on a computer screen (end-user display device 10). The agent virtual environment 8 may optionally contain AVE objects 11 with which Agent 1 can interact. Unlike the shared digital content 5, User 3 cannot directly manipulate AVE objects; however, AVE objects are represented to User 3 as part of Agent 1's environment, with a scaled-down view displayed to User 3 as shown by 11RER. Agent 1 can navigate the agent virtual environment 8 and manipulate digital and AVE objects within the context of the physics of the agent virtual environment 8. For example, Agent 1 can pull down shared digital content 5, walk up to AVE object 11 (a virtual stool), and sit on it. Agent 1 can also see representations of User 3 (3VER) and the real-world environment 7 overlaid and / or simultaneous with the shared digital content 5, as if Agent 1 were "inside" the screen and looking at the real world.

[0013] The user 3 is a human user 3 in a real-world environment 7 who sees and controls (interacts with) what is represented on an end-user display device 10 (a real-world display device such as a screen) from a computer. The end-user display device 10 displays a real-world environment representation (RER) of the shared digital content 5RER (e.g., from a web page via a browser) and a compressed, overlaid, and / or blended view of the agent's environment (AVE), which includes: Agent 1's representation of the real-world environment (1RER) Representation of real-world environments for AVE objects 11 (11RER)

[0014] When User 3 makes changes to the shared digital content 5 via the user interface, the changes are reflected in Agent 1's representation of the shared digital content 5 in Agent Virtual Environment 8. Similarly, when Agent 1 makes changes to objects (items / components of the shared digital content 5) in Agent Virtual Environment 8, the changes are reflected in User 3's representation of the shared digital content 5 on his / her screen 10.

[0015] The agent and its environment can be visualized in various ways on the end-user display device. For example, the figure shows user-agent interaction from a perspective rotated through the agent's virtual space. Agent 1 is simulated within a three-dimensional agent virtual environment 8. Shared digital content 5 resides within the agent virtual environment 8 and is relatively in front of agent 1 from the user's perspective. Agent 1 perceives a representation of the user's virtual environment (e.g., via a camera feed), which may optionally be simulated as a plane or otherwise within the agent virtual environment 8, or may be superimposed / blended with the plane of the shared digital content 5. Thus, agent 1 perceives the relative position of the digital content to the user and can ascertain whether the user is looking at the shared digital content 5. Agent 1 is shown with its hand outstretched as if to touch the shared digital content 5.

[0016] The plane 14 can be located anywhere within the agent virtual environment 8. Figure 4 shows the agent 1 facing the user 3 and the plane 14 containing the shared digital content 5 (objects) behind the agent 1. A representation of the agent 1 may be shown in front of a real-world environment representation (RER) of the shared digital content 5RER. The user 3 sees a representation of the agent 1 interacting with the shared digital content 5 as if the agent 1 were in front of a screen.

[0017] 3 shows a schematic diagram of human-computer interaction in which an agent 1 facilitates interaction between a user 3 and digital content 4. The user 3 can interact with the digital content 4 directly from the computer 2 by providing input to the computer 2 via any suitable human input device / primary input device and / or interface device, and receiving output from the computer 2 from any suitable human output device, such as a display or speaker. The agent 1 can similarly provide input to and receive output from the computer 2.

[0018] Providing an agent with the perception of objects outside its simulated environment is akin to providing the agent with "augmented reality," as it provides the agent with a real-time representation of the current state of real-world elements and the computer, blending representations of digital content items from the computer, real-world elements, and more into a unified view of the world. Because the virtual world is augmented with real-world objects, this may be more accurately defined as enhanced virtuality, which enhances the autonomous agent's vision as opposed to human vision.

[0019] (Agent-Autonomous Interaction with Digital and Real-World Inputs) Agents can include cognitive, location, embodiment, speech, and dynamic aspects. Agents are embodied, meaning they have a virtual body that can be articulated. The agent's body is graphically represented on a screen or other display. Agents may be simulated using a neurobehavioral model (a biologically modeled "brain" or nervous system), which includes multiple modules with coupled computational and graphical elements. Each module represents a biological process and includes computational elements that relate to and simulate the biological process, as well as graphical elements that visualize the biological process. Thus, agents are "self-animated" because they do not require external control and exhibit naturally occurring autonomous behaviors such as breathing, blinking, looking around, yawning, and lip movement. Biologically based autonomous behavior can be achieved by modeling multiple aspects of the nervous system, including, but not limited to, sensory and motor systems, reflexes, perception, emotional and regulatory systems, attention, learning and memory, reward, decision-making, and goals. An agent's face reflects both the agent's brain and body, revealing mental states (e.g., mental attention through eye direction) and physiological states (e.g., fatigue through eyelid position and skin color). The use of neurobehavioral models to animate virtual objects or digital entities is further disclosed in Sagar M, Seymour M, and Henderson A (2016), "Creating Coordination with Autonomous Facial Animation," ACM Communications 59(12), 82-91, and in WO2015016723A1, which is assigned to the assignee of the present invention and incorporated herein by reference. Time-stepping mechanisms such as those described in WO2015016723A1 can synchronize or coordinate an agent's internal processes.

[0020] Agent aspects may be dynamic, meaning that the agent's future behavior depends on its current internal state. The agent's nervous system, body, and environment are coupled dynamical systems. Complex behaviors can be generated with or without external input. Complex behaviors can be generated without relying on the agent's internal representation of its external environment. Elements around the agent (such as visual stimuli or content items from the user) can form part of a causal network leading to behavior generation. Agent behavior and behavior can emerge from the bottom up. Agent dynamics can be chaotic, making it difficult to predict the agent's behavior in a given situation. The agent's behavior depends on continuous feedback loops from both the real-world environment 7 and the agent's virtual environment 8, as shown in Figure 7. Thus, the agent 1 is situated in both environments, receiving input 1102 from the real-world environment 7, such as audio input from the user, the user's environment (real-world environment 7) via a microphone, and visual input from the user via a camera. Real-world input 1102 to the Agent1 simulation influences Agent1's behavior, which in turn is shown in Agent1's animation output 1106. For example, Agent1 may smile when seeing a human user through an indirect mechanism, in that recognizing the user releases a virtual neurotransmitter such as dopamine within a neurobehavioral model that may naturally induce a smile in Agent1. This output 1106 can then influence the real-world environment 7, for example, by eliciting a response or emotion in the user of the real-world environment 7, which is again input for the Agent1 simulation. Agent1's agent virtual environment 8 is also input 1102 to Agent1. For example, shared digital content items within the agent virtual environment may be perceived by Agent1 in real time as a stream, influencing Agent1's behavior, which also results in a particular Agent1 animation output 1106. This output 1106 may affect the agent virtual environment 8.For example, Agent 1 can change the position of a virtual object in the Agent virtual environment 8. The new object position then becomes input 1102 to Agent 1's simulation, driving a continuous feedback loop.

[0021] An agent may be simulated and / or represented in any suitable manner, in any suitable form, such as a human form, a fictional human form, a human form, a robot, an animal, etc. A user may be able to select or change the form that the agent takes, or the agent's form may change in response to the user and the real or virtual environment.

[0022] (Move to coordinate position) Placing an agent in the Agent Virtual Environment 8 enables natural-looking "on-the-fly" animations, allowing the agent to freely interact with surrounding digital content items (shared digital content 5) as if they were present in the agent's 1 environment. This differs from pre-generated or pre-recorded animation snippets, which are difficult to use when interacting with unpredictable or dynamic content. An example of a limitation of pre-recorded animations is using a simulated arm to reach a specific content item at location (X,Y). The animation is limited in functionality by taking into account exactly where the content item is placed. If the content item is moved or resized, the animation cannot be changed accordingly.

[0023] Agents can be animated using "on-the-fly" animation algorithms. In embodied agents, an effector (e.g., a hand) reaching an end-goal position (digital object coordinates) is achieved by calculating the vector of joint degrees of freedom that will move the end-effector to the goal state. Computational techniques such as inverse kinematics can be used to animate agents. Inverse kinematics can be approximated using known techniques such as Jacobian or circular coordinate descent (CCD). Neural networks can be trained to map body positions to represent reaching or pointing at objects in 2D or 3D coordinate space (i.e., learning hand-eye coordination). In one example, a movement success predictor uses a deep convolutional neural network (CNN) to determine how likely a given movement is to reach a specified coordinate and a continuous servomechanism that uses the CNN to continuously update the agent's motor commands.

[0024] (Virtual environment) The agent is in the agent virtual environment and is aware of objects and their location relative to the agent in the agent virtual environment. In most embodiments described herein, the agent is located in a three-dimensional agent virtual environment. In one embodiment, the agent virtual environment is a 2D virtual space represented by a 2D pixel array. The 2D agent is located in the 2D virtual space and can interact with shared digital content in the 2D space that can be displaced horizontally (left or right) and / or vertically (up or down) from the agent.

[0025] In Figure 1, Agent 1 is located in 3D space behind a plane 14 (screen) that reflects the shared environment. Agent 1 views the shared digital content on plane 14, which is at the outer boundary / surface of the agent virtual environment 8 (which is a rectangular prism). Figure 2 shows Agent 1 touching a content item directly in front of Agent 1. Figure 4 shows Agent 1 positioned in front of plane 14, where the on-screen shared environment digital content items are located. The content items / objects that are the shared digital content may be displaced in three dimensions from Agent 1 horizontally (left, right, front or back) and / or vertically (up or down) relative to Agent 1.

[0026] The coordinates of the plane 14 of the virtual 3D space may be mapped to browser (or other display) coordinates, so that when the movement of the agent 1 with respect to the shared digital content is superimposed on the screen, the movement aligns with the position of the digital content item.

[0027] The real-world coordinates of the display on which the agent is displayed can change. For example, an agent displayed in a browser can be moved or resized by the user, for example, by dragging / resizing the browser. The real-world physical dimensions of the display can be changed dynamically, and changing the physical dimensions of the display in the real-world environment will update the representation of the agent's virtual environment proportionally.

[0028] The agent virtual environment can be simulated by any suitable mechanism and can obey certain laws of nature. The physical properties of the agent virtual environment can be defined as collision detection between objects and / or agents, physical constraints (e.g., allowed movement between the joints of an object), simulation of atmospheric resistance, momentum, gravity, material properties (e.g., elastic items can return to their natural shape after being stretched by an agent), etc.

[0029] (Input received by the agent) The agent receives any suitable visual input from the real world (such as a video stream from a camera), depth-sensing information from a camera with range imaging capabilities such as a Microsoft Kinect camera, audio input from a microphone or the like, input from a biosensor, input from a thermal sensor or any other suitable input device, touchscreen input (when the user presses the screen). A single input may be provided to the agent, or a combination of inputs may be provided. The agent can spatially locate aspects of the real world in relation to itself, and the user can direct the agent's attention to objects, such as things or people, in the user's real-world space.

[0030] An agent can receive input / communication from a user via a computing device. An agent may perceive content items such as the user's cursor / mouse or touchscreen input (which is also input from the real world). An agent can follow the user's mouse movements with the user's eyes and / or hands. In some embodiments, an agent can perceive keyboard input. For example, a user can communicate with an agent via a keyboard rather than verbally.

[0031] Visual (e.g., video stream) input is provided to the agent's visual system. Routines from an appropriate interface may be used to interface with the hardware. This may be provided via an interface between the agent's program definition and the camera and / or pixel streamer. This interface is provided by a vision module, which provides a modular wrapper for capturing image data from the camera at every time step for the agent. The vision module may not perform any perceptual processing of the image data (similar to the human "retina").

[0032] An agent may provide output to the real world through output devices of a computing system, such as visual output on a display, audible output through a speaker, or any other suitable means. An agent may also output mouse movements, keyboard strokes, or other interface interactions such as clicks and presses.

[0033] An agent can have two or more vision systems, which allows the agent to receive visual input from different sources simultaneously without the need to overlay the different visual sources on top of each other. For example, an agent may have two vision modules, one for camera input and one for user interface input, and be able to view both simultaneously. A different saliency map can be applied to each input source. For example, a saliency map that prioritizes faces can operate on a vision system configured to receive real-world input, such as camera input, while a saliency map that prioritizes text recognition and / or UI element recognition can be applied to a vision system configured to receive UI input.

[0034] (Digital Content) Shared digital content associated with a website can include web forms, buttons, text, text boxes, text fields, video elements, or images. In mixed reality applications such as augmented reality, virtual reality, and holograms, digital content items can refer to 3D objects. In applications (such as VR / mobile phone / computer applications), digital content items can be objects defined by an object-oriented language.

[0035] Digital content may be dynamic (e.g., an item of digital content may move around the screen), and an agent may follow or guide the movement of such dynamic content items. Digital content items may be static or interactive. Static content items may be text or images. For example, an agent engaging in a user's academic tutorial may interact with the text by pointing out particular words displayed to the user and asking the user if they understand the meaning of the words. Interactive digital content items may respond in some way to user and / or agent actions. For example, a button is a digital content item that presents other content when clicked. Digital content items may be represented in two or three dimensions. For example, FIG. 15 shows a car object modeled in 3D and located within the agent's virtual environment.

[0036] Because the agent is embodied in the virtual space along with the content item, the agent can use gestures, view the item, approach the item, manipulate or handle the item, or invoke methods related to the item, and interact with the shared digital content in several different ways, in a flexible and scalable manner. · Leaning its body towards the item and / or tilting its head towards the item. · Look at the item with your eyes. · Make a gesture to point to the item, turn your head in the general direction of the item, or wave your hand in the direction of the item. Approach an item by walking towards it, teleporting close to it, or floating close to it.

[0037] (Enabling agent interaction with shared digital content) Embodied agents are programmatically defined independently of the digital content they interact with: there is no central controller that controls both the digital content and the embodied agents. This allows embodied agents to be used flexibly with digital content created by different third-party providers. The two mechanisms that enable agents to perceive and natively interact with new digital content comprise computer vision and interaction modules (supportive interfaces).

[0038] (Computer Vision) In one embodiment, the agent receives input in the form of visual data from a source representing a display to a user. For example, the agent views pixels on a screen or in a 3D virtual space via visual (virtual) images and / or object recognition. The agent can be equipped with visual object recognition via standard computer vision / image processing / image recognition / machine learning techniques, etc., to identify objects / subjects in an image or video, the contours / colors of objects in an image or video, and interact with them accordingly. The agent may be equipped with optical character recognition to recognize text and natural language processing to understand the text. In other words, digital content is visually represented to the agent in the same way as it is represented to a human user (visual representation of pixels). The agent system can include built-in image recognition and / or learning, or it can use third-party services for image recognition.

[0039] (Digital content input) A visual representation of the shared environment can be sent to the agent in a manner similar to how screen sharing software sends a visual representation of a screen to a remote location. The interface may be configured to send packets of information from the computing device to the agent that describe what is being output by the computing device at any given time. The data may arrive as image files (e.g., JPEG and GIF), or the data may arrive as individual pixels assigned to specific X and Y coordinates (and Z, in the case of mixed reality). To minimize the amount of bandwidth, the interface may be configured to send only information updates on sections of the screen that have changed and / or to compress the data being sent.

[0040] Figure 17 shows a schematic diagram of an agent's interaction with digital content using a computer vision system. The modules mentioned are not necessarily modular components of code, but may be functional networks of modules driven by a highly interconnected neurobehavioral model. In 2901, an end-user computing device user interface (e.g., a browser) renders updates to the display (e.g., a web page redraw). Through a shared memory buffer 2930, the agent's controller 2951 provides pixels from the display as input 2902 to the agent's retina 2903. This is the interface between the agent's program definition and the pixel streamer. The pixels may then be provided to a perception module 2952, where a visual inspection module 2904 performs image processing. The portion of the pixel stream being attended to / processed may be guided by an attention module 2905, which determines what the agent is paying attention to. Information resulting from the processing of pixel data may be passed to a reactive and / or decision-making module 2907, which drives the agent's behavior. For example, after recognizing a portion of the image as a button, the reaction and / or decision-making module 2907 can cause the agent's hand to reach out and touch the button 2910. The action or behavior taken by the agent is passed to the simulated physiology module 2908. The physiology module 2908 can include subcomponents for controlling various body parts, including an arm control device 2910. At 2911, the agent controller 2951 can operate to map interactions between the agent and digital content to user 3 interface interactions. In the virtual environment, "physical" actions by the agent in the AVE can be translated as actions on the user 3 interface. For example, when the coordinates of the agent's body intersect with a perceived shared environment plane, the coordinates of the plane intersection can be translated into a mouse click or touchpad touch event at a corresponding pixel location on the user interface.The event action / back channel 2932 is then sent to the computing system as a human input device event (e.g., a mouse click at a corresponding pixel location on a browser). In one embodiment, an open browser implementation of the Chromium Embedded Framework (CEF) is configured to allow agents to interface with web digital content. Off-screen rendering allows the contents of a browser window to be output to a bitmap and rendered elsewhere.

[0041] (embodied interactions in response to triggering events) Touching a digital content item is one type of embodied interaction that can result in interaction with a digital content item, although the invention is not limited in this respect. In other embodiments, a particular gesture of an agent directed at a digital content item can trigger an event on that item. For example, an agent looking at an item and blinking can trigger an event on that item. Another example is gesturing at a digital content item, such as a button.

[0042] (Direct control of the browser or computing system) In one embodiment, the agent directly controls a mouse, touchpad, or other primary input device as if the input device were an effector of the agent. In other words, the agent can control the input device in the same way that it controls its own body / muscle movements. For computing devices, direct control of the computing device by the agent can be enabled by any suitable technology, such as technology that enables remote access and remote collaboration on a person's desktop computer via a graphical terminal emulator.

[0043] (saliency map) The attention module 2905 can include a "saliency" map to guide the agent's attention. A salience map is an indication to the agent of importance on the screen. The salience map can define where the agent's attention and focus is. Examples of features that may be differentially treated as salient include: User: The agent may include a face detection module for detecting faces. Face detection may be related to the agent's emotional influence and user interaction loop. The face detection module uses a face tracking and resolution library to find faces within the agent's visual input stream. The presence of a face may be interpreted by the agent as a highly salient visual feature. The resolved facial expression from any detected face may be fed into an expression recognition network. Motion - Since the vision module does not attempt perceptual processing of the video input, a motion detection module may be provided. The motion detection module may be a component of the agent's visual perception system that compares temporally adjacent video frames to infer simple movements. The resulting "motion map" may be used as a driver of visual saliency. Recognizing specific objects or images Text recognition; salience attributed to specific keywords or text patterns ·color ·brightness Edge

[0044] Salience maps may be user-defined. In one embodiment, the user may interactively communicate to the agent which features the user would like the agent to treat as salient (focus on). If multiple salience maps are used, each may be weighted, with the agent's final focus of attention driven by a weighted combination of each active salience map. In other embodiments, salience may be defined externally, providing an artificial pointer for the agent to mark digital content items for focus.

[0045] (Switching UI control between user and agent) The user and agent may control the mouse, keyboard, or other primary input mechanism. In one embodiment, a mechanism for collaboration is provided in the form of a control mechanism that ensures that once either party moves the mouse, they move the mouse until the initiated action is completed before allowing the other party to move the mouse. The control mechanism may enforce turn-taking between the user and the agent. In other embodiments, the user and agent may use a dialog to determine who has control of the UI (e.g., the user may ask the agent if the agent can take over control, or vice versa).

[0046] (Perceptual Control) An agent can control its perceptual input; for example, it can choose to look at the user rather than the content, or vice versa. The benefit of enabling perception in an agent via visual pixel recognition is that it provides freedom / flexibility in what the agent can look / focus on and therefore perceive (using a "Fovea" subset / region of the pixel stream displayed to the agent at higher resolution). The agent can focus on any pixel displayed on the user interface, or on any superstructure / aspect displayed on the user interface, such as any generated pattern, color, or represented object.

[0047] (Integration with agent simulator) Figure 23 is a class diagram of one embodiment using a CEF browser. The CEF window defines what the user sees on the UI. Variables may be defined within the agent's neurobehavioral model to store interactions with the shared digital content (e.g., browser content). One set of variables may be for the user's interactions, and another set may store the agent's interactions. The agent's set of variables may be set by the runtime via the neurobehavioral modeling language. The runtime host 2304 may set up variable monitoring for both sets of variables. Upon receiving updates to these variables, the runtime host 2304 constructs UI events (e.g., mouse / keyboard events) and sends them to the shared environment VER (which may correspond to a plane in the agent's environment where the agent views the shared digital content). The shared environment VER is owned by the runtime host but rendered off-screen into a buffer, which the runtime host then sends to the neurobehavioral modeling framework (simulator) for the neurobehavioral modeling framework to render the content in 3D space.

[0048] When the user interacts with the browser, the UI sends set variable messages (user_mousedown, user_mouse_x, etc.) to the SDK. The coordinates received from the UI are relative to the Off-Screen rendered shared environment 2308 (the agent's window). The coordinates are converted to the x,y position of the browser. Neurobehavioral model methods convert the coordinates and determine if the browser object contains the mouse. The runtime host then constructs and forwards mouse and key events to the shared environment (shared browser).

[0049] When an agent interacts with the shared environment 2308, the runtime host 2304 receives callbacks for monitored variables that have changed. If there is no user event in the same callback, the agent's interaction is forwarded to the shared browser. Otherwise, the user's interaction overrides the agent's interaction. The neuroethological modeling language variables for shared interactions can be defined as follows: To track mousedown and mouseup event variables, use user_mousedown, user_mouseup, persona_mousedown, persona_mouseup To track key down and key up event variables, use user_keydown, user_keyup, persona_keydown, persona_keyup

[0050] To indicate that an event (such as a mouse event or keyboard event) is occurring, the neurobehavioral modeling framework variables (user_mousedown / user_mouseup / agent_mousedown / etc.) are different from the last time-step. Instead of using 1 / 0 switching to indicate that an event is occurring, a counter counts events, adding 1 to the previous value each time. When it reaches 1000, the counter is reset to 1. The reason for this is that down and up (on, off) events can be on the same time step to track all events, and the current and previous values ​​of the variable do not need to match. This ensures that no events are lost. Queues may be implemented to facilitate rapid (faster than the agent's time-stepping) and / or simultaneous input and output events. The agent internal 2314 can control the rate of the agent's time-stepping in the shared environment VER and update user interactions.

[0051] (Interaction Module) The interaction module facilitates agent recognition of shared digital content and can define and communicate to the agent the interaction affordances of content items represented in the agent's virtual environment. The interaction module may be a support library or an application programming interface (API). When the agent 1 decides to take a particular action, the interaction module 16 translates that action into instructions natively defined by the author of the third-party digital content. The agent can use the interaction module to directly and dynamically interact with digital content (e.g., web content, application content, other programmatically defined content, etc.). The interaction module translates between digital content defined by computer program-readable information and information understandable by the agent.

[0052] Shared digital content items may be represented to agents as concept objects, which are abstractions of native (or natively rendered) digital content items. Concept objects may be defined by certain characteristics such as virtual world environment coordinates, color, identifier, or anything else related to the interaction between the agent and the respective digital content item. Concept objects are objects that represent real-world digital content items in an abstracted way that translates across the AVE, and agents need only understand the "concept" (i.e., size, color, position / placement) and metadata associated with the concept.

[0053] The native digital content item exists and is presented to the user in its native format, but the digital content item has an additional ID that Agent 1 can use to reference the digital content item. In one embodiment, the agent sees the shared digital content items / objects through an interaction module 16 that translates HTML information about those objects and passes it to Agent 1. The interaction module 16 abstracts the native digital content item information, such as HTML information, and translates this for Agent 1 so that Agent 1 understands what the content item is, its properties, and what input is required.

[0054] Figure 5 shows a system diagram for human-computer interaction facilitating user 3's interaction with a web page. The system may include a client side 512, a digital content provider server side 520, and an agent side (agent system or simulation) 510, and optionally, communication with a third-party service 590. The digital content provider may define digital content on a web server 522 and / or a server-side database 524. The web server 522 may provide the digital content to a web client 511, such as a web browser, where the user can view it. Human-computer interaction is facilitated by the inclusion of an interaction module 16 on the client side 512. Agents are simulated on the agent system 510, which may be a cloud server. The interaction module 16 processes the digital content (which may be defined, for example, by HTML code) and transforms items relevant to the agent so that they can be perceived by the agent. This can provide the agent with a context map of the content items on the web page. The agent and its virtual environment reside on the agent system 510. The agent system 510 includes an agent modeling system 513 that simulates the agent within its virtual environment, and an animation rendering system 514 for rendering a representation of the agent. A knowledge base 515 can provide the agent with a base level of domain knowledge about the environment it resides in and the types of content items it can interact with. The agent may be supported by third-party services 590 (e.g., a natural language processing system provided by a third party).

[0055] FIG. 6 shows a swimlane process diagram for human-computer interaction. A digital content provider 620 defines a digital content item (621). The digital content provider 620 includes an interaction module 16 linked to the digital content item 621 so that the digital content item can support agent interaction. To augment the digital content with agent 1 interaction, the digital content provider may enable this by linking to or including the interaction module 16 when defining the website. In another embodiment, if the interaction module 16 is not linked, a proxy server may be provided to execute the digital content that links to or includes the interaction module 16 to enable interaction via agent 1.

[0056] The user device 612 natively displays digital content items to the user, who interacts with web pages, applications, or other computer programs defined by the digital content provider 620. The interaction module 16 converts a particular digital content item from its native definition into a conceptual object. The conceptual object is sent to the agent 1 cloud 610, allowing agent 1 to conceptualize the digital content (616). The conceptual object is input for agent 1's simulation (617). Thus, the conceptual object forms part of agent 1's environment. Thus, agent 1 can interact with and change its behavior based on the conceptual object. Agent 1 can directly manipulate the digital content (618), for example, by pushing or moving the conceptual object. When agent 1 directly manipulates the conceptual object, the interaction module 16 translates the agent's actions into changes to the digital content item. In other words, the interaction module 16 updates the digital content to reflect agent 1's changes to the content item (642). Indirect manipulation is also possible, for example, by the agent looking at the conceptual object or gesturing to the conceptual object's location. At 619, a representation of Agent 1 interacting directly or indirectly with the conceptual object is rendered.

[0057] The agent may not directly perceive the content item in the same way that a user does (e.g., through pixel recognition). The interaction module 16 can pass a "conceptual object" to the agent, which is an abstracted representation of the content item. The conceptual object may include basic properties associated with the item to the agent, such as a tag and location that define what the content item is. The interaction module 16 can provide the agent with a list of the affordances of the content item in the context of the agent's virtual environment. The conceptual object corresponding to the content item may include further information that defines the object as having "physical qualities" with which the agent can interact.

[0058] The interaction module 16 can provide support for "actions" that allow the user 3 to interact with the digital content item through the agent. This "interaction" functionality provides an abstraction layer for the agent 1 to perform "actions" on objects. Examples of actions are press, drag, push, look, grab, etc. These actions are translated by the interaction module 16 into operations that can be performed on the digital content item in a way that works for the interaction space at that time. Actions that are not translated can be ignored. For example, the agent 1 can "action" a press on a web button, which is translated by the interaction module 16 into a click on a button HTML element. A "push" action on the web element may be ignored, but if acted upon in the 3D interaction space on a ball, it will cause the ball to move.

[0059] For example, an agent wanting to scroll down a web page can send a scroll down command to the interaction module, which can translate the agent's actions into web-readable commands, such as JavaScript code. The JavaScript code activates the actions on the web page. Thus, the agent does not need to be able to communicate directly in web languages. This makes the system scalable, as the agent can be applied to a variety of contexts.

[0060] In another embodiment, an agent wishing to enter text into a content item can send an input command to the interaction module 16, which executes the necessary JavaScript instructions to position a cursor within that text field and enter the text the agent wishes to enter. Thus, in a web interaction context, the user's view of a digital content item may be a full-fidelity web page with correctly stylized web elements (e.g., HTML elements). The agent may have an abstracted visualization of the HTML page made up of conceptual objects. Conceptually, this is similar to the agent seeing a simplified view of the web page, with only those aspects of the web page that are relevant to the user's interaction with the web page.

[0061] In addition to converting information from web language into agent-perceivable information, the interaction module 16 can convert "physical" agent actions on conceptual objects into instructions for the movement of corresponding digital content items. Instead of directly manipulating the content items through native methods, the agent manipulates the content items as if they were physical items in the virtual environment. To this end, the interaction module 16 can further include physics transforms to move HTML elements to simulate physics. For example, such a "physics"-type action of pushing a content item can be converted into an HTML element by the interaction module 16 moving the object by a certain amount a certain number of frames per second, thereby "simulating" a physical push. Thus, the various "actions" included by the interaction module 16 can directly implement or simulate (by approximating) changes to elements such as their position, content, or other metadata.

[0062] In some embodiments, Agent 1 may query Interaction Module 16 to obtain additional information about the item. For example, Agent 1 wishing to "read" the text in an item (where the HTML is a text field) can query Interaction Module 16 to obtain the text in the text field.

[0063] (Dynamic Control) A digital content item may include web elements with a set of parameters established by a document designer to define the element's initial structure and content. These include both the element's physical characteristics, such as the element's absolute or relative spatial position within the document, and attributes applied to any user text content entered into the element, such as font type, font size, font color, and font attributes such as bold and italic. A document may also be designed to allow a user to reposition one or more elements using traditional click-and-drag techniques. When the digital content is in the context of a web page, an interaction module 16, such as a JavaScript interaction module 16, may be provided to allow an agent 1 to modify the physical characteristics and / or attributes of a web element. Elements on an HTML page may become controllable after the page is rendered through a combination and interaction of several web-related standards, such as Dynamic HTML (DHTML), HTML, Cascading Style Sheets (CSS), the Document Object Model (DOM), and scripting. When a web page is loaded, a browser can create a Document Object Model (DOM) to represent the HTML elements on the page. JavaScript can be used to interact with the DOM (an interface in a browser that allows programs to access and modify the content, structure, and style of a document). The JavaScript interaction module 16 may include methods that specifically enable certain types of interaction between the agent 1 and a web page via the DOM.

[0064] The QuerySelector can be used to query the DOM. The interaction module 16 allows the agent 1 to modify the web page: Modifying / deleting HTML elements in the DOM or on the page Modifying and / or adding CSS styles to elements Read and / or modify element attributes (anchor text href attribute, image text src attribute, alt attribute, or any custom attribute) Create new HTML elements and insert them into the DOM / page Attach event listeners to elements. For example, event listeners can listen for clicks, keypresses, and / or submits and react to these with JavaScript.

[0065] Although dynamic control of web pages has been described with reference to a JavaScript interaction module 16, the invention is not limited in this respect. For example, in another embodiment, JQuery may facilitate interaction between agent 1 and digital content. The interaction / assistance module may be implemented in any suitable web-related open technology standard.

[0066] (Other interaction contexts) FIG. 8 illustrates User 3's interaction in a virtual reality context, such as a virtual reality environment. The methods and systems for user interface interaction described above apply equally to virtual / mixed / augmented reality interactions. A conceptual shared environment may also be provided, including a set of objects accessible to both User 3 and Agent 1. An interaction module 16 may be used to translate between Agent 1's space and User 3's space. The interaction module 16 may be incorporated into a virtual reality application (VR application) having a virtual reality environment (VR environment). The interaction module 16 facilitates visual matching in interactions between Agent 1 and a digital content item. Alternatively and / or additionally, the agent may be provided with full-fidelity computer vision of the shared digital content defined by the VR environment (and similarly with the augmented reality embodiments described below).

[0067] A user 3 in a real-world environment 7 views a 3D VR environment 13, which may contain shared digital content 5, including 3D objects. An interaction module may convert digital content from the VR application into conceptual objects 9 perceivable by an agent 1. Thus, the agent 1 may directly or indirectly interact with or reference the conceptual objects 9. When the agent 1 directly interacts with the conceptual objects 9, for example, by pushing a cylinder along the virtual floor of the agent 1's environment, the interaction module 16 converts this into a change in a digital object natively defined in the VR application. FIG. 8 shows an agent virtual environment 8 that is smaller than the VR environment 13. However, the agent 1's agent virtual environment 8 may be co-augmented with or larger than the VR environment 13. The agent 1 may pass an item from the shared environment to the user 3. For example, the agent 1 may pass a soccer ball, which is shared digital content, to the user 3. In one embodiment, ray tracing is used to simulate the view for the agent 1 within the three-dimensional scene. The interface can cast rays into the three-dimensional scene from the perspective of Agent 1 and perform ray tracing using the rays to determine whether an object is within the field of view of Agent 1. Thus, the behavior of Agent 1 can be based on whether the shared digital content 5 is within its field of view.

[0068] FIG. 9 shows a system diagram for human-computer interaction in a virtual reality context. A VR digital content item 824 may be defined in a VR application 822 and displayed to a user 3 on a VR display 811. An interaction module 16 converts the VR digital content item (e.g., a VR object) into a conceptual object perceivable by Agent 1. Thus, Agent 1 can interact with the perceivable conceptual object. This interaction is transformed by the interaction module 16 to reflect corresponding changes on the VR digital content item. The VR application can render a scene to the user 3 that includes Agent 1, any aspect of Agent 1's environment, and the digital content item facilitated by the interaction module 16. The Agent 1 system 810 also includes a knowledge base 815 that provides Agent 1 with knowledge about the particular domain in which Agent 1 is interacting.

[0069] FIG. 10 illustrates user interface interactions in an augmented reality context. The interactions are similar to those described with reference to virtual reality, except that the user views virtual digital content through a viewport, depicted as, for example, a mobile phone screen, which may be overlaid on a view of the real world. FIG. 11 illustrates a system diagram for human-computer interaction using a mobile application. Content items 824 may be defined in mobile 1022 and displayed to a user 3 on a display 1011, such as a mobile device screen. The interaction module 16 converts the digital content items in the mobile application into conceptual objects that the agent can perceive. Thus, the agent 1 can interact with the recognizable conceptual objects. This interaction is then translated by the interaction module 16 to reflect corresponding changes in the mobile application.

[0070] In one embodiment, Agent 1 lives in WebGL, and everything in the scene can be an object that the persona can manipulate. WebGL (Web Graphics API) is a JavaScript API for rendering interactive 3D and 2D graphics within a compatible web browser without the use of plug-ins. A WebGL-compatible browser provides a virtual 3D space into which Agent 1 is projected. This allows the virtual space in which Agent 1 operates to be displayed in any compatible web browser or WebGL-compatible device, and Agent 1 can manipulate web and 3D objects within the same virtual space.

[0071] (Rendering agents to digital content) The animation renderer may render the agent's animation and the agent's environment for display to the user. The generated animation may then be streamed as a video stream to a UI device (such as a browser). In one embodiment, the agent may be rendered in a limited area of ​​the end user's display. In a web context, the agent may be restricted to an HTML DIV element. In another embodiment, the display of the agent on the end user's display may be unlimited.

[0072] Pixels may be blended so that either the agent or the digital content is transparent, allowing one to see what is behind the agent or digital content, respectively. If the AVE is a 3D environment and the display is a 2D screen, the AVE may be rendered as a 2D animation from the user's perspective. An AVE may be rendered as a moving background or foreground to digital content (e.g., containing natively rendered HTML web elements) for an interactive user experience.

[0073] The agent's virtual environment may be presented to the user through one or more viewpoints. The agent and / or user can change the user's viewport into the agent's virtual environment. For example, in a pinhole camera model of rendering the agent's environment, the agent can change the pinhole position to change the angle / orientation and / or zoom of the user's view of the agent's environment. Instead of rendering an animation of the agent's environment (which can be computationally intensive), in some embodiments, a 2D representation of a viewpoint corresponding to the agent's viewpoint (from the agent's point of view) can be rendered and presented to the user.

[0074] (camera image overlay) A user can use gestures to draw or indicate an agent's attention to an area on the computer screen. Figure 19 shows an example of how a user representation (3VER and / or 3RER) captured by camera 15 is displayed on screen 10. The representation may be overlaid on other representations on screen 10 and may be semi-transparent, allowing user 3 to see both user 3's body and other content on screen 10. Alternatively, user 3's background may be automatically cropped (using standard image processing techniques) so that only user 3's image or hand is displayed on the screen, and as a result, the representation need not be transparent. In a further embodiment, user 3's representation is visible only to agent 1. Two buttons A and B are displayed on screen 3120, and user 3's hand 3145 is hovering over button B. A representation of user 3's hand is displayed on the screen. Agent 1 can see the same representation user 3 sees and therefore can also see which button has user 3's attention. The saliency map can assign importance to human hands or movements (User 3's hand can move over Button B to draw attention to it). Thus, User 3 can interact with a non-touchscreen screen in a similar manner to a touchscreen, and the user's gestures (e.g., clicking with a finger) can be translated into input device events (e.g., clicks) using the interface module. In other embodiments, instead of User 3's representation being displayed on the screen, User 3 can receive some other visual indicator of where User 3 is pointing, such as by Agent 1 looking in that direction. Similarly, Agent 1 can perceive where User 3 is looking on the screen by tracking the user's gaze or from verbal instructions / commands from User 3.

[0075] (Variation) Multiple agents can interact with digital content independently. Multiple agents can interact with each other as well as with a user. Multiple agents can be simulated in the same virtual environment or in different virtual environments. Multiple agents can have the same sensory capabilities or different capabilities. One or more agents can interact with multiple users. Any one or more of the users can converse with one or more agents and direct one or more agents to manipulate a user interface with which the user is interacting as described herein.

[0076] (Combination of Computer Vision and Interaction Module 16) Embodiments of the computer vision and interaction modules may be combined. In one embodiment, an agent can perceptually recognize features of a content item, such as an image, by processing the pixels of the image. This allows the agent to discuss features such as the color of the item.

[0077] (Agent Knowledge) The agent may also access object metadata from a provided object database, for example. One example is a catalog of purchased items. Agent 1 can associate digital content items with the purchased items in the database catalog and use this information to interact with User 3. The agent may control the navigation or display aspects of the User 3 interface. For example, in the context of a website, Agent 1 may control which parts of the web page are displayed by scrolling up, down, left, right, or zooming in and out.

[0078] (persistent agent browser) In one embodiment, an agent can understand the nature of a particular digital content item, allowing it to be integrated with different digital content sources (e.g., different websites). Therefore, the agent can facilitate user content interaction over the Internet in a scalable manner. Such an agent can be provided via a bespoke browser. Using machine learning techniques, the agent can be trained to understand the nature of web language learning associations between content items and actions that can / should be taken in relation to such content items. For example, Agent 1 may be trained to identify text areas regardless of their exact configuration, read user-visible indicators of the text areas, and fill in those areas on behalf of the user.

[0079] In one embodiment, a user can teach an agent about a digital content item. For example, the user can hover the mouse over a digital content item and name the item. The agent can observe this and associate the item's name with a representation of the digital content item (either a pixel representation or a conceptual representation provided by the interaction module).

[0080] (embedded actions and agent-recognized locators) In one embodiment, as shown in FIG. 18, digital content may be associated with embedded actions and / or agent-perceivable locators. The agent-perceivable locators locate digital content items (and may be associated with spatial coordinates corresponding to the digital content items within the agent's virtual environment). The locators may be associated with metadata describing the digital content items. The locators may replace and / or support a saliency map. In one example of a locator replacing a microscope map, a button is tagged with metadata indicating that the locator corresponding to the button should be clicked by Agent 1. In an example of a locator supporting a saliency map, the locator is placed on the button (the locator is automatically generated by reading the HTML content of a webpage and assigning the locator to the item using an HTML button tag). The button saliency map may be provided together with any other saliency map, e.g., a color saliency map; for example, the saliency map may be configured to prompt the agent to click a red button. Embedded content may be provided on a website accessible by the agent, but not necessarily by the user. For example, embedded content that is visible to the agent allows the agent to click on links that are not visible to the user, navigate to another page, or read information that is not visible to the user.

[0081] (Conversational Interaction) Agent 1 can engage in human interaction using the same verbal and non-verbal means (gestures, facial expressions, etc.) that humans do. Responses can include computer-generated speech or other audio content played through one or more speakers of the end-user computing device. Responses generated by Agent 1 can be visible to User 3 in the form of text, images, or other user-visible content. Agent 1 can converse with the help of third party services such as IBM Watson or Google Dialogflow and / or conversation corpora.

[0082] Figure 20 shows User 3. Agent 1 can see User 3. Agent 1 receives the following information about User 3 and can use this information to inform its interactions in both the real and virtual worlds: The embodied agent can receive camera input and calculate where the user's gaze is. This may be mapped to a content item or object in user space / the real world that the user is looking at. The user's gaze may be tracked using the angle of the user's eyes and / or the angle of the user's head. The embodied agent can also track the user's eye and head movements and calculate the user's eye and head angles. The embodied agent further receives linguistic input, including instructions from a user, which may in some cases direct the actions of the embodied agent and / or the gaze of the embodied agent. Other input may include text, for example via a keyboard. Once an embodied agent has identified the user, it can track the user's position by following them with its eyes (looking at the user) and leaning towards them. The embodied agent can recognize where the user-controlled mouse is located on the screen relative to the digital content item. The agent 1 can also detect user touches, for example via a touchscreen monitor. The embodied agent can monitor the user's movements via a camera, specifically the movements of the user's arms, hands, and fingers. It can use facial expressions to detect the user's emotions and adapt accordingly. The user's voice tone may be used to detect user information so that the agent can adapt accordingly. · The agent can ensure that it receives the user's attention before proceeding with the conversation. The agent has memory of past interactions with the user and can use this information in the conversation.

[0083] The agent may use the context of the interaction, the digital content item, and user information to resolve ambiguity. Actions performed by the agent may be dynamic, tailored to the user, or contextually intent / goal-directed. The agent may have access to sources of information about the user. For example, the agent may recognize the user's location and / or time of day (e.g., via geolocation services) and use this to guide interactions accordingly. The agent combines conversation, emotion, cognition, and memory to create an interactive user experience. Embodiments provide a system for synthesizing the agent's emotional and gestural behavior with content interaction. The agent interacts with the user through conversation and user actions, including attention, eye direction, user movements, and other input received about the user, to establish the user's goals, beliefs, and desires and guide the user accordingly. The agent may include an emotional response module that reacts to the user. In one embodiment, the agent's interactions are guided or prescribed by learned responses (e.g., reinforcement learning). The prescribed behavior may be guided by a knowledge base of rules. In one embodiment, the way the agent interacts is guided by the user's psychometric profile.

[0084] In one embodiment, a simulated interaction between an embodied agent and a user may be implemented using an embodied dyadic turn-taking model that applies to the gaze of the user and the embodied agent. The embodied agent may indicate the end of a conversational turn during the interaction by attempting direct gaze with the user. In a similar manner, the embodied agent can know that the user has indicated the end of a turn when the embodied agent detects that the user has begun to look directly at the embodied agent.

[0085] Referring to FIG. 20, an example of taking turns is shown. At 3325, User 3 can prompt Agent 1 by looking at Agent 1, saying Agent 1's name, or pointing at Agent 1, for example. Once User 3 receives Agent 1's attention, Agent 1 can return User 3's eye contact, signaling to User 3 that Agent 1 has acknowledged User 3's attention (3330). User 3 can then respond with a smile, prompting Agent 1, for example, to provide information or otherwise proceed to interact with User 3 and take their turn. Once Agent 1 is finished, Agent 1 can signal User 3 by stopping and looking directly at User 3 (3345). User 3 can then smile at the agent, acknowledging the agent (3350), and take their turn (3355). Once User 3 is finished, User 3 can stop and turn their attention back to Agent 1, allowing Agent 1 to take their next turn (3365). The above description is merely an example, and the indication by the user 3 may take other forms that enable the agent 1 to recognize that it is the agent 1's turn. The user 3 may indicate the end of his / her turn during an interaction, for example, by a verbal or non-verbal cue. The non-verbal cue may include, for example, a smile, a blink, a head movement, or a body movement including the arms, hands, and fingers. Similarly, the indication by the agent 1 may take other forms that enable the user 3 to recognize that it is the user 3's turn. The agent 1 may indicate the end of his / her turn during an interaction, for example, by a verbal or non-verbal cue. The non-verbal cue may include, for example, a smile, a blink, a head movement, or a body movement including the arms, hands, and fingers.

[0086] (Attention Modeling) An agent attention model can be implemented as a saliency map of regions in the visual field, with visible locations competing for attention. More active locations are more salient. A saliency map is an image that indicates the inherent qualities of each location (pixel), with more active locations being more salient. Several types of saliency active in the human brain can be implemented in an embodied agent. These include a visual frame that updates with every eye or head movement of the user 3 or the embodied agent 1. Other saliency maps use a frame of reference that remains stable regardless of head and eye movement. Salience features that can be mapped include color, brightness, or intensity of stimuli present in the visual field. Still other saliency maps can be created that focus on expectations or desires and, from these expectations or desires, predict where salient locations on the saliency map are likely to be. As implemented in an embodied agent 1, these saliency maps are combined to derive an aggregate measure of saliency.

[0087] In one embodiment, the attention model implemented includes multiple saliency maps representing different types of salience targets, where the total saliency map is the weighted sum of the maps used. How the various saliency maps are weighted can be varied. In one example, weighting may be used as follows, so that an object (thing or person) becomes more salient when both the user and the embodied agent are focusing on the object: Gaze saliency = Weighting 1 × Embedded Agent Gaze Map + Weighting 2 × User Gaze Map + Weighting 3 × (Embedded Agent Gaze Map × User Gaze Map) Salience points = weight 1 × embedded agent's gaze map + weight 2 × user's points map + weight 3 × (embedded agent's points map × user's points map)

[0088] Information used to create the salience map includes the inputs discussed in conversational interaction above. Non-visual inputs, such as auditory and textual inputs, can also be applied to the salience map by mapping the input to the visual map space. For example, to map a user pointing at an object, the system can calculate the location where the user is pointing and map that location to the visual map space. If the input is auditory, the system calculates the location where the sound is coming from and maps that location to the visual map space. These pointing and auditory maps are combined with the visual map. In one embodiment, the attention model includes a sub-map (tracker) that allows the embodied agent to continue tracking previously attended objects, even if there is a subsequent shift in current attention.

[0089] Figure 21 shows a screenshot of the user interface for setting multiple saliency maps. The weighting of the maps can be changed using the slider 2110 shown in Figure 21. The slider changes the default weighting.

[0090] In certain embodiments, multiple visual feeds from different sources can both activate a visual saliency map and provide a representation of the multiple visual feeds to the embodied agent. Each field can be used to compute a saliency map that controls visuospatial attention. Each visual feed is associated with multiple saliency maps, which can highlight areas of the visual feed as more salient relative to other areas of the visual feed. For example, a camera visual feed (capturing user 3 interacting with embodied agent 1) and a browser window visual feed (capturing the computer browser with which user 3 and / or agent 1 interact) both activate a visual saliency map. Another visual feed may be provided by the agent's own 3D virtual environment. For example, a 2D plane corresponding to agent 1's view of the surrounding environment may be provided to agent 1 by ray casting from the agent's viewpoint.

[0091] A human-like attention model may be implemented to allow Agent 1 to focus on only one aspect of the visual feed at any given time. Thus, a single salient region across two or more maps is always selected for attention at any given moment, and upon an attention switch, the two or more regions can be thought of as a single region with two parts. Weighting may be applied across the visual feed so that certain visual feeds are determined to be more salient than others.

[0092] A salience map may be applied to verbal cues to assist the agent in locating the item to which the user is referring. For example, keywords such as "left," "right," "up," and "down" may be mapped to a "verbal cue" salience map that highlights corresponding areas of the visual feed as salient. A verbal cue salience map may be combined with other salience maps, such as those described above, to facilitate joint attention and interaction. For example, if User 3 says "button on the left," the verbal cue salience map may highlight the left half of the screen. This can then be combined with an object salience map that detects buttons and highlights the buttons on the left side as most salient, and therefore the buttons to which the agent will pay attention.

[0093] Referring to FIG. 22, a system implemented with multiple visual feeds is shown. The system extracts multiple low-level features from two visual streams 3770, 3775 (3780). A central circular filter 3740 is then applied to each map to derive feature-specific saliency maps. From these maps, feature maps are created (3725), and combinations of maps are created for specific features (3720), such as the user's face or the user's surroundings. A combined saliency map is then created (3710) from the individual saliency or feature maps. These feature-specific maps are combined in a weighted sum to generate a feature-independent saliency map 3710 (3710). The system then applies a "winner take all" (WTA) operation, selecting the most activated locations in the saliency maps as regions of interest (3710). In addition to camera feeds, audio or other feeds can be provided to the system to create the feature maps (3625, 3725). In one embodiment, a turn-taking feature map may be incorporated into the system such that the focus of the turn-taking feature map depends on who is taking the turn, creating a bias in the saliency map related to turn-taking.

[0094] (Example of interaction) Figure 12 shows a screenshot of an agent facilitating a user's interaction with a web page. The web page contains shared digital content with several menu items, a search bar, and navigation buttons. Relative to the user, Agent 1 is positioned in front of the shared digital content. Agent 1 has perceptual awareness of the shared digital content 5. Therefore, Agent 1 can refer to or point to different content items as part of its interaction with the user. Agent 1 can engage in an interaction with User 3 to ascertain what the user wants to navigate to next. Agent 1 can look around and see digital content items in which the user has expressed interest, triggering navigation to the URL to which the menu content item links. The agent can indicate and visually press a specific content item, as if clicking the item.

[0095] FIG. 13 shows a screenshot of Agent 1 positioned behind shared digital content 5. The example shown is a bank website showing several available credit cards for the user to choose from. Agent 1 is conversing with the user and can ask if the user would like more information about any of the credit cards shown in the shared digital content 5. The user can directly click on one of the shared digital content items (credit cards), which is a clickable HTML object that triggers a link to more information. Alternatively, the user can ask Agent 1 to provide more information about one of the credit cards. Agent 1 has a perception of the conceptual object representing the shared digital content 5 and can therefore determine which credit card the user is interested in from the information provided to Agent 1 via the interaction module. Instead of the user clicking on the credit card, Agent 1 can trigger an action on the conceptual object representing that credit card, which is then translated via the interaction module into a click on an item on the website. Because the credit card is a clickable image, Agent 1 can use pixel information from the image to determine the image's color and therefore understand whether User 3 is referring to an item of interest by color. FIG. 13B shows the web page when a digital content item is selected by the user. Agent 1 may access metadata about a digital content item, such as via a digital content provider database, and convey detailed information about that content item to the user. Furthermore, such information may be displayed to the user (not shown) within the AVE.

[0096] Figure 14 shows how Agent 1 engages the user through an interactive menu (shared digital content 5) to help them find the appropriate credit card. Again, the user can click on a menu item directly. Alternatively, the user can interact with Agent 1, who can click on a digital content item on their behalf. For example, Agent 1 can ask the user, "How often do you use your card?" The user can read one of three displayed menu options: "Always," "Sometimes," or "I don't know." Agent 1 matches User 3's utterance to one of the digital content items, with the text on the digital content item provided to Agent 1 as an attribute of the content item's corresponding concept object. Instead of "I don't know," the user might utter something slightly different from the given text, such as "I'm not sure," allowing Agent 1 to infer the option the user wants to select. Figure 14 shows the option "I don't know" already selected. Agent 1 is in the process of touching (corresponding to a mouse click) on a menu item that reads "Full Amount on My Statement." Agent 1 has a direct view of the digital content item that the agent is touching. When Agent 1 touches the context object, a click on the digital content item is triggered via interaction module 16. The user may choose not to go through the menu sequentially and may choose to skip the first set of options, for example to tell Agent 1 that they would like a lower rate. Agent 1 may select that option and then query the user for information on the previous step.

[0097] FIG. 15 shows a series of screenshots of Agent 1 facilitating user interaction with a website to help the user purchase a car. Thus, Agent 1 can walk up to the car and point out its features. If the car rotates or changes position, Agent 1 can continue to point to the same side of the car. FIG. 15A shows Agent 1 walking into the screen from the left. The website includes a top menu containing digital content 4 and a virtual showroom 1780 displaying a virtual car 1760. The virtual showroom is configured to display 3D models of items available for purchase. User 3 may not be able to directly manipulate objects in Agent 1's virtual environment because they are in Agent 1's virtual environment. Nevertheless, User 3 can indirectly interact with such objects by communicating with Agent 1 such that Agent 1 manipulates objects in Agent 1's virtual environment. For example, User 3 may ask Agent 1 to pick up an item, manipulate the item, change the item's color, or rotate the item so User 3 can view the item from a different angle. Figure 15B shows Agent 1 facing User 3 after entering the virtual showroom, conversing with User 3, and discovering User 3's interests. Figure 15C shows Agent 1 gesturing toward a virtual car 1750. Because Agent 1 perceptually knows where the car is relative to Agent 1 and, in turn, relative to the screen, Agent 1 can gesture toward the car's virtual space coordinates. Because Agent 1 is located in a virtual environment containing the virtual car, Agent 1 can walk toward the car and point to its various features (defined to Agent 1 by object metadata such as feature and tag coordinates). Because both Agent 1 and the car are in the same 3D virtual environment, Agent 1's "real world" on the screen shrinks in size as Agent 1 moves toward the car, adding realism to the interaction. Figure 15E shows Agent 1 presenting a menu 2150 of options for asking User 3 about User 3's interests.Figure 15F shows Agent 1 selecting a menu item on behalf of User 3. Agent 1's touch on the item triggers a click on the corresponding digital content item. Figure 15F shows navigation to an image showing the interior of a car, where the user's view is from Agent 1's perspective and Agent 1's hands are visible.

[0098] Figure 16 illustrates a sequence of nonlinear interactions that depend on user feedback. Figure 16A shows Agent 1 displaying two purchasing options, X and Y, to the user, while Figure 16B shows Agent 1 gently / tentatively setting option Y aside after receiving feedback from the user, implying that the user has a preference for other option Y but is not 100% sure. After receiving feedback from the user that option Y is categorically undesirable by the user, Figure 16C shows Agent 1 being able to eliminate the option by removing it entirely. Figure 16D shows Agent 1 with an expectant look on its face while waiting for the user to make up their mind about whether they want to bring item Y back into consideration. Figure 16E shows Agent 1 returning item Y to the screen.

[0099] An agent may assist a user in learning how to perform specific tasks using third-party software. For example, in an application like Photoshop, the agent can converse with the user and demonstrate how to navigate the application's interface by physically interacting with controllable digital content items in the interface. A general algorithm may be provided to the agent with steps on how to manipulate or navigate the user interface. The agent may be equipped with a knowledge library of tasks, each associated with a temporal sequence of actions and defined by finding a salient item (e.g., a symbol or text), performing an action on that item (e.g., click), and proceeding to the next step. The agent can take control of a human input device (e.g., mouse / touch input) to perform these steps when requested by the user to perform actions stored in the agent's knowledge library for UI interaction tasks.

[0100] An agent can assist a user in making purchase choices when shopping online. The agent can be embedded in an e-commerce platform and help guide the user through items. In one embodiment, the agent can receive information about the user's profile to guide the interaction. For example, a user can have a profile created from their engagement history with the e-commerce system (e.g., previous purchases). The user's profile is stored in a broader recommendation system, and the agent can use the recommendation system to appropriately recommend items to the user. The agent can purchase products on the user's behalf by navigating the e-commerce UI for the user. Accidental or unintentional purchases can be minimized by the agent observing the user's body language while verbally confirming that the purchase will proceed. For example, if the agent recognizes the user nodding, it can say, "Yes, buy it," and by looking at the relevant product, the agent can understand the user's desire and be confident in purchasing the product.

[0101] An agent may include a computer program configured to perform tasks or services for a user of an end-user computing device (personal assistant). Examples of tasks performed by an agent on behalf of a user include placing a phone call to a user-specified person, launching a user-specified application, sending a user-specified email or text message to a user-specified recipient, playing user-specified music, scheduling a meeting or other event on a user's calendar, getting directions to a user-specified location, getting scores associated with a user-specified sporting event, posting user-specified content to a social media website or microblogging service, recording a user-specified reminder or note, getting a weather forecast, getting the current time, setting an alarm for a user-specified time, getting stock prices for a user-specified company, finding nearby commercial establishments, performing an internet search, etc.

[0102] In one embodiment, the agent can play or collaborate with the user on an activity. For example, the agent and user can collaboratively paint a picture. The painting itself may occur in the shared environment. If the user draws an object such as an apple in the shared space, the agent may recognize the object as an apple and talk about apples. Alternatively, visual characteristics such as a drawn line or the color red may be added to the apple. In another embodiment, the agent and user can interact in the shared environment and collaborate on some other activity, such as playing music. The shared environment may include a virtual instrument, e.g., a saxophone. The user can strike the keys on the saxophone, and the agent can react accordingly. The agent can notice items moving in the shared space and recognize and react to objects in the shared space.

[0103] Navigating the Internet, an application, an operating system, or any other computer system may be performed collaboratively by an agent and a user. A user accessing the Web can search or browse (looking for new or interesting things). An agent can facilitate Web browsing, transforming Web browsing into a real-time activity rather than a query-based search. In one embodiment, an agent uses interactive dialogs to assist a user with a Web search. The agent can use a separate search tool "in the back end" that is not displayed to the user, such as traditional search engine (e.g., Google) results. Traditional recommender systems require users to perform a mental "context switch" from browsing the web page space to explicitly interacting with a search assistant. An embodied agent as described herein allows a user's train of thought during browsing activities to be uninterrupted by the need to switch to a separate query interface. Because the agent constantly monitors browsing activity, agent recommendations are made in real time as relevant pages arise in the user's browsing activities. The agent can show the user how to navigate or perform actions on Internet homepages. The agent can "preview" links or chains that are not visible to the user, and can therefore warn the user about "dead ends" (links that appear interesting from reading the link text or looking at the link image, but turn out not to be), or garden paths are sequences of links that provide enough incentive to follow the path, but ultimately lead to a dead end. In one embodiment, the agent 1 may be provided with a syntactic understanding of common web languages ​​and may assist the user in their online information search.The user can ask the agent to scroll up or down, click a specific search link, enter a URL into the browser, go back to the previous page, or navigate in other ways.

[0104] The agent may assist the user in filling out a form. If the agent has a profile of the user, it can automatically fill in details about the user that the agent knows. The agent may query the user about areas in which the agent is unsure. If the form was not filled out successfully (e.g., the user has lost a password constraint), the agent may identify the constraint that was violated from the error message and ask the user to correct it. The agent can automatically navigate and / or gesture towards fields that need to be re-entered.

[0105] The shared environment can be used as a canvas for Agent 1's tutorial of User 3. For example, Agent 1 can ask User 3 to solve a mathematical equation in the shared environment and demonstrate all the steps. Agent 1 can use character recognition to recognize the numbers and steps the user writes. Agent 1 can engage in dialogue with the user and interact with them about the work on the shared canvas by referencing different lines in the work, erasing or underlining mistakes, or adding more work to the shared canvas. In real-world subjects, agents can assist users in learning in virtual or augmented reality contexts. For example, to facilitate medical training, Agent 1 can use a tool in front of someone and hand it to them to try on their own.

[0106] (advantage) Early research in human-computer interaction emphasized the benefits of interacting with a "model world": a computing device interface in which objects and actions resemble / mirror objects and actions in the real world, making them more intuitive to human users. Traditional applications of artificial agents to human interfaces resulted in "magical worlds" in which the world is changed by the action of hidden hands. The embodiments described herein extend human-computer interaction and artificial intelligence by enabling interfaces in which the world is changed by visible, useful hands, visually displaying inputs received by the agent (via the agent's gaze direction / eye direction / body language) and information about the agent's mental state in an efficient manner. Instead of users seeing interface elements "move on their own," users can now visualize the agent's thought process and actions that lead to interface manipulation. The user directly observes the agent's autonomous actions, and the agent can observe actions taken autonomously by the interface user.

[0107] Embedded agents that operate directly in the user interface, rather than in the "background" or "backend," increase the degree to which users perceive software as acting like an assistant. When users perceive the agent's actions as actions they "could have done themselves," they are more willing to conceptualize the agent in the role of an assistant.

[0108] In a traditional command line or menu-driven interface, the user executes and types input, the system accepts the input, computes some action, displays the result, and waits for the next input. While the user is preparing the input, the system is not doing anything. While the system is running, the user is not doing anything in the interface. The embodiments described herein provide agents that can execute independently and simultaneously.

[0109] The methods described herein assist users in performing tasks that operate a computing device through continuous and / or guided human-machine interaction processes. Limiting the capabilities of an autonomous agent to actions that may be emulated through conventional input devices or methods may make it easier for the agent to guide the user, since the agent cannot take shortcuts.

[0110] Embodiments of the present invention may usefully convert between traditional application programs (now generally written to communicate with outdated input / output devices and user interfaces) and user interfaces that include autonomous agents embodied as described herein, so that the logic and data associated with the traditional programs can continue to be used in the new interaction context.

[0111] The embodiments described herein avoid the need for synchronization between pre-recorded or pre-defined animations and / or dialogue and runtime-generated desired behavior in the output. Thus, fixed-length verbal and / or non-verbal segments need not be mapped to the execution context and synchronized spatially or temporally to fit variable and / or dynamic digital content. There is no need to synchronize verbal and non-verbal behavior because the agent simulation, including the neuro-behavioral model, drives both.

[0112] The embodiments described with reference to agent perception via computer vision leverage techniques in DOM parsing, computer vision, and / or natural language processing to view digital content and automatically extract useful information, simulating the human process that occurs when interacting with digital content. Agents can send messages or requests to interaction modules that embody actions relevant and compatible with the objects being manipulated. These actions allow agents to control and manipulate content hosted in all types of browsers in an abstracted manner. User-machine interactions are therefore modeled in a way that allows for the separation of dialogue knowledge from application knowledge. This opportunity dramatically reduces the cost of moving interaction systems from one application domain to a new one. Real-time performance and timing control of UI interactions are enabled by the strict time model of the agent and architecture. The agent's real-time response to user input minimizes delays between parts of the system during on-time execution of actions.

[0113] The advantage of allowing the agent to interact with items in the simulated environment is that this creates more natural-looking interactions, making the agent 1 real to the user. For example, the content items with which the agent interacts may be simulated to have mass and other physical properties. Elements provided by digital content providers are rendered natively as defined by the digital content provider.

[0114] By providing conceptual objects with affordances, agents can perceive actions available to them without having to perform significant cognitive processing / image recognition / reasoning, thus reducing computational power and / or time. The benefit of using interaction modules is that they allow for scalable customization of agent interactions in different contexts or environments, and allow computer interfaces to keep up with increasing technology complexity.

[0115] (interpretation) The described methods and systems may be utilized on any suitable electronic computing system. According to the embodiments described below, the electronic computing system utilizes various modules and engines to utilize the methods of the present invention.

[0116] An electronic computing system may include at least one processor, one or more memory devices or interfaces for connecting to one or more memory devices, input / output interfaces for connecting to external devices to enable the system to receive and act on instructions from one or more users or external systems, a data bus for internal and external communication between the various components, and a suitable power source. Additionally, an electronic computing system may include one or more communication devices (wired or wireless) for communicating with external and internal devices, and one or more input / output devices such as a display, pointing device, keyboard, or printing device.

[0117] The processor is configured to execute program steps stored as program instructions in a memory device. The program instructions enable various methods of implementing the present invention as described herein to be performed. The program instructions can be developed or executed using any suitable software programming language and toolkit, such as, for example, a C-based language and compiler. Furthermore, the program instructions can be stored in any suitable manner, such as, for example, stored on a computer-readable medium, transferred to a memory device, and read by a processor. The computer-readable medium may be any suitable medium for tangibly storing program instructions, such as, for example, a solid-state memory, a magnetic tape, a compact disk (CD-ROM or CD-R / W), a memory card, a flash memory, an optical disk, a magnetic disk, or any other suitable computer-readable medium.

[0118] The electronic computing system is arranged to communicate with a data storage system or device (eg, an external data storage system or device) to retrieve relevant data.

[0119] It will be understood that the systems described herein include one or more elements configured to perform the various functions and methods described herein. The embodiments described herein are intended to provide the reader with examples of how the various modules and / or engines that make up the elements of the systems may be interconnected to enable functions to be implemented. Furthermore, the embodiments herein explain in system-related details how the steps of the methods described herein may be performed. Conceptual diagrams are provided to show the reader how various data elements are processed at different stages by the various modules and / or engines.

[0120] It will be understood that the arrangement and configuration of modules or engines may be adapted as appropriate depending on system and user requirements, such that various functions may be performed by modules or engines different from those described herein, and that certain modules or engines may be combined into a single module or engine.

[0121] It will be understood that the described modules and / or engines may be implemented and provided with instructions using any suitable form of technology. For example, a module or engine may be implemented or created using any suitable software code written in any suitable language, where the code is compiled to generate an executable program that can be run on any suitable computing system. Alternatively, or in conjunction with an executable program, a module or engine may be implemented using any suitable mix of hardware, firmware, and software. For example, portions of a module may be implemented using an application-specific integrated circuit (ASIC), a system-on-chip (SoC), a field-programmable gate array (FPGA), or any other suitable adaptable or programmable processing device.

[0122] The methods described herein can be implemented using a general-purpose computing system specially programmed to perform the described steps. Alternatively, the methods described herein may be performed using specific electronic computer systems such as data sorting and visualization computers, database query computers, graphical analysis computers, data analysis computers, manufacturing data analysis computers, business intelligence computers, artificial intelligence computer systems, etc., where the computer is specially adapted to perform the described steps on specific data captured from an environment related to a particular domain.

[0123] By providing conceptual objects with affordances, agents can perceive the actions available to them without having to perform significant cognitive processing / image recognition / reasoning.

[0124] The advantage of using the interaction module 16 is that it allows scalable customization of the agent's interaction in different contexts or environments, allowing the computer interface to keep up with increasing technological complexity.

[0125] (summary) In one aspect, a method is provided for displaying an interaction between an embodied artificial agent and digital content on an end-user display device of an electronic computing device, the method comprising: creating an agent virtual environment having virtual environment coordinates; simulating digital content within the agent virtual environment; simulating an embodied artificial agent within the agent virtual environment; enabling the embodied artificial agent to interact with the simulated digital content; and displaying the interaction between the embodied artificial agent and the digital content on the end-user display device.

[0126] In one embodiment, the virtual environment is a 3D virtual space, and the virtual environment coordinates are 3D coordinates. Optionally, the interaction is moving towards or looking at the digital content item. Optionally, the interaction is gesturing towards or touching the digital content item by moving a multi-joint effector. Optionally, the movement of the multi-joint effector is simulated using inverse kinematics. Optionally, the movement of the multi-joint effector is simulated using neural network-based mapping of joint positions to target positions.

[0127] In another aspect, a method is provided for interacting with digital content on an electronic computing device through an embodied artificial agent, the method including the steps of displaying the digital content to a user on a user interface on the electronic computing device, creating an agent virtual environment having virtual environment coordinates, simulating the digital content in the agent virtual environment, simulating an embodied artificial agent in the agent virtual environment, enabling the embodied artificial agent to interact with the simulated digital content, translating the interaction into an actuation or manipulation of the digital content on the user interface, and displaying the interaction by overlaying the embodied artificial agent on the digital content and displaying the digital content and the overlaid embodied artificial agent on the user interface.

[0128] Optionally, the virtual environment is a 3D virtual space, and the virtual environment coordinates are 3D coordinates. Optionally, the digital content is simulated in the agent virtual environment as pixels in the agent virtual environment, each pixel having a coordinate location in the agent virtual environment. Optionally, the simulated embodied interaction is an intersection between coordinates of pixels corresponding to the digital content and the body of the agent. Optionally, converting the simulated embodied interaction includes mapping the embodied interaction to an input device event. Optionally, the human input device event is a mouse event, a keyboard event, or a touchscreen event.

[0129] In another aspect, a system for facilitating interaction with an electronic computing device is provided, the system including at least one processor device; at least one memory device in communication with the at least one processor; an agent simulator module executable by the processor and configured to simulate the embodied agent; an interaction module executable by the processor arranged to transform the digital content into conceptual objects perceivable by the embodied agent and enable the embodied agent to interact with the digital content; and a rendering module executable by the processor configured to render the digital content, the embodied agent, and interaction of the embodied agent with the digital content.

[0130] Optionally, the interaction module is further configured to translate embodied agent actions on the conceptual objects into changes to the digital content. Optionally, the interaction module associates the conceptual objects with coordinates representing the position of the conceptual objects relative to the embodied agent. Optionally, the interaction module associates the conceptual objects with one or more affordances of the corresponding digital content. Optionally, the computing device is web content and the interaction module is JavaScript code integrated into the web content. In another aspect, an enforcement agent located within a virtual environment created on an electronic computing device is provided, the enforcement agent being programmed to receive input from the real-world environment, receive input from the virtual environment, and act in response to the input from the real-world environment and the virtual environment, wherein the input received from the real-world environment and the virtual environment is received from both the real-world environment and the virtual environment via a continuous feedback loop.

[0131] In another aspect, an embodied agent simulator for simulating interaction with a user is provided that is implemented on an electronic computing device and is programmed to receive user input, receive digital content input including information regarding digital content displayed to a user, and simulate a conversation with the user by generating responses to the user based on both the user's natural language declarations and the digital content input.

[0132] Optionally, the response and / or input to the user is verbal. Optionally, the verbal response and / or input is auditory or textual. Optionally, the response and / or input to the user is visual. Optionally, the verbal response and / or input is a gesture or a facial expression.

[0133] In another aspect, there is provided a method for facilitating user interaction with an electronic computing device having a display and input means, and at least one processor and memory for storing instructions, the processor being programmed to: define a virtual environment having virtual environment coordinates; determine the magnitude of the coordinates relative to real-world physical dimensions of the display; simulate an embodied agent in the virtual environment, the position of the agent relative to the virtual space defined by the virtual environment coordinates; simulate one or more digital objects in the virtual environment of the agent, the positions of the one or more digital objects relative to the virtual environment defined by the virtual environment coordinates; enable the embodied agent to interact with the one or more digital objects using information about the virtual environment coordinates of the agent and the virtual environment coordinates of the virtual objects; and enable the interaction of the agent with the one or more digital objects to be displayed to a user on the display.

[0134] In another aspect, a method is provided for providing an embodied agent having a simulated environment with a substantially real-time perception of continuous visual input from outside the simulated environment, the method comprising the steps of: providing an interface between a programmatic definition of the embodied agent and / or the simulated environment and the continuous visual input; capturing the visual input from outside the simulated environment at each time step of simulation of the agent and at each time step transferring the input data to the embodied agent and / or the simulated environment; and inputting the visual input into a vision system of the agent or simulating the visual input within the simulated environment of the agent.

[0135] In another aspect, a method is provided for enabling an artificial agent to interact with a user interface, comprising the steps of presenting digital content to the artificial agent that is displayable to a user via an end-user display, and translating cognitive judgments and / or physical movements of the artificial agent into actuations or manipulations of human input devices that control the user interface.

[0136] Optionally, converting the simulated embodied interaction includes mapping the embodied interaction between the artificial agent and the digital content to input device events. Optionally, the human input device events are mouse events, keyboard events, or touchscreen events.

[0137] In another aspect, a method for interacting with an artificial agent is provided, the method including the steps of simulating the artificial agent in an agent virtual space; representing digital content perceivable by the artificial agent by simulating digital content in the agent virtual space; displaying the artificial agent, the artificial agent virtual space, and the digital content to a user on a display; receiving an image of the user from a camera; tracking the user's gaze on the display based on the received image; and simulating an embodied interaction between the artificial agent, the user, and the digital content based on user input including at least the user's gaze.

[0138] Optionally, the user's gaze is tracked using the user's eye angle and / or the user's head angle. Optionally, further comprising tracking the user's eye movements, wherein the simulated embodied interaction is further based on the user's eye movements.

[0139] Optionally, the user's eye movements are tracked using the angle of the user's eyes and / or the angle of the user's head. Optionally, the user input comprises linguistic input. Optionally, the user input comprises auditory input or text input. Optionally, the user input comprises touchscreen or mouse movements. Optionally, the user input comprises visual input. Optionally, the visual input is a gesture or a facial expression. Optionally, the gesture comprises one or more movements of an arm, a hand, or a finger. Optionally, the simulated embodied interaction between the artificial agent, the user, and the digital content base comprises the user directing their attention to an object in the digital content of the agent virtual space.

[0140] Optionally, the method includes tracking the gaze of the artificial agent, and the simulated embodied interaction is further based on a dyadic turn-taking model applied to the gaze of the user and the gaze of the artificial agent. Optionally, the artificial agent indicates the end of its turn in the interaction by attempting direct gaze with the user. Optionally, the artificial agent perceives that the user has indicated the end of its turn when the user initiates direct gaze with the artificial agent. Optionally, the interaction takes place in a virtual reality environment. Optionally, the interaction takes place in an augmented reality environment.

[0141] In another aspect, there is provided an embodied agent simulator for simulating interaction with a user implemented on an electronic computing device, the embodied agent simulator being programmed to: simulate an artificial agent within an agent virtual space; represent digital content perceivable by the artificial agent by simulating the digital content within the agent virtual space; display the artificial agent, the artificial agent virtual space, and the digital content to a user on a display; receive an image of the user from a camera; track the user's gaze on the display based on the received image; and simulate the embodied interaction between the artificial agent, the user, and the digital content based on user input including at least the user's gaze.

[0142] Optionally, the user's gaze is tracked using the angle of the user's eyes and / or the angle of the user's head. Optionally, further comprising tracking the user's eye movements, wherein the simulated embodied interaction is further based on the user's eye movements. Optionally, the user's eye movements are tracked using the angle of the user's eye head angle. Optionally, the user input comprises linguistic input. Optionally, the user input comprises auditory input or text input. Optionally, the user input comprises touchscreen or mouse movement. Optionally, the user input comprises visual input. Optionally, the visual input is a gesture or a facial expression. Optionally, the gesture comprises one or more movements of an arm, a hand, or a finger. Optionally, the simulated embodied interaction between the artificial agent, the user, and the digital content base comprises the user directing their attention to an object in the digital content of the agent's virtual space. Optionally, the method includes tracking the gaze of the artificial agent, and the simulated embodied interaction is further based on a dyadic turn-taking model applied to the gaze of the user and the gaze of the artificial agent. Optionally, the artificial agent indicates the end of its turn in the interaction by attempting direct gaze with the user. Optionally, the artificial agent perceives that the user has indicated the end of its turn when the user initiates direct gaze with the artificial agent. Optionally, the interaction takes place in a virtual reality environment. Optionally, the interaction takes place in an augmented reality environment.

[0143] In another aspect, there is provided a method for interacting with an artificial agent, the method including the steps of simulating the artificial agent in an agent virtual space; representing digital content perceivable by the artificial agent by simulating digital content in the agent virtual space; displaying the artificial agent, the artificial agent virtual space, and the digital content to a user on a display; receiving images of the user and the user's environment from a camera; tracking the user's gaze based on the received images; and simulating an embodied interaction between the artificial agent, the user, the user's environment, and the digital content based on user input including at least the user's gaze.

[0144] Optionally, the user's gaze is tracked using the user's eye angle and / or the user's head angle. Optionally, further comprising tracking the user's eye movements, wherein the simulated embodied interaction is further based on the user's eye movements. Optionally, the user's eye movements are tracked using the user's eye angle and / or the user's head angle. Optionally, the user input comprises linguistic input. Optionally, the user input comprises auditory input or text input. Optionally, the user input comprises touchscreen or mouse movements. Optionally, the user input comprises visual input. Optionally, the visual input is a gesture or a facial expression. Optionally, the gesture comprises one or more movements of an arm, a hand, or a finger. Optionally, the simulated embodied interaction between the artificial agent, the user, the user's environment, and the digital content comprises the user directing attention to an object in the digital content of the agent virtual space or an object in the user's environment. Optionally, the method includes tracking the gaze of the artificial agent, and simulating the embodied interaction is further based on a dyadic turn-taking model applied to the user's gaze and the gaze of the artificial agent. Optionally, the artificial agent indicates the end of its turn in the interaction by attempting direct gaze with the user. Optionally, the artificial agent perceives when the user begins to gaze directly at the artificial agent, indicating the end of the user's turn. Optionally, the interaction takes place in a virtual reality environment. Optionally, the interaction takes place in an augmented reality environment.

[0145] 1. A method of an embodied agent simulator for interacting with an artificial agent, the method comprising: simulating the artificial agent in an agent virtual space; representing digital content perceivable by the artificial agent by simulating digital content in the agent virtual space; displaying the artificial agent, the artificial agent virtual space, and the digital content to a user on a display; receiving images of the user and the user's environment from a camera; tracking the user's gaze based on the received images; and simulating the embodied interaction between the artificial agent, the user, the user's environment, and the digital content based on user input including at least the user's gaze.

[0146] In another aspect, a method-implemented system for interacting with an artificial agent includes simulating the artificial agent in an agent artificial virtual space, displaying the artificial agent and the artificial agent virtual space to a user on a display, receiving an image of the user from a camera, tracking the user's attention, applying an attention model to the user's attention, providing output to the artificial agent, and simulating an embodied interaction between the artificial agent and the user based on the output of the attention model.

[0147] In another aspect, an embodied agent simulator for simulating an interaction with a user implemented on an electronic computing device, the embodied agent simulator programmed to: simulate an embodied agent of an agent virtual space; display the artificial agent and the embodied agent virtual space to a user on a display; receive images of the user from a camera; track the user's attention based on the received images; apply an attention model to the user's attention; provide output to the embodied agent; and simulate an embodied interaction between the embodied agent and the user based on the output of the attention model.

[0148] Optionally, the attention model is also applied to the attention of the artificial agent. Optionally, the method includes receiving an image of the user's space from a camera. Optionally, the method includes representing digital content perceivable by the artificial agent by simulating digital content in the agent's virtual space, the digital content being visible to the user. Optionally, tracking the user's attention includes tracking the user's gaze. Optionally, tracking the user's attention includes tracking the user's eye movements. Optionally, tracking the user's attention includes tracking the user's eye movements on a display. Optionally, the method includes tracking the gaze of the artificial agent, the attention model also being applied to the gaze of the artificial agent. Optionally, the attention model includes a saliency factor and an object to which the artificial agent and the user jointly pay attention with increased salience. Optionally, the object is a human and / or an object. Optionally, the salience factor includes a weighted saliency-gaze factor based on saliency-gaze maps of the user and the artificial agent. Optionally, the weighted salience gaze factor is calculated as a first weighting multiplied by the artificial agent's gaze map, a second weighting multiplied by the user's gaze map, and a third weighting multiplied by the artificial agent's gaze map multiplied by the user's gaze map. Optionally, the salience factor includes a weighted salience point factor based on the user and artificial agent's salience point maps. Optionally, the weighted salience point map factor is calculated as a first weighting multiplied by the artificial agent's point map, a second weighting multiplied by the user's point map, and a third weighting multiplied by the artificial agent's point map and the user's point map. Optionally, the attention model is also applied to user eye movements, user actions, objects in the user's environment, and actions in the user's environment. Optionally, the objects are humans and / or objects. Optionally, the attention model is also applied to user input. Optionally, the user input comprises an auditory input or a textual input. Optionally, the input comprises a linguistic input.Optionally, the user input comprises a touchscreen or a user mouse movement. Optionally, the user input comprises a visual input. Optionally, the visual input is a gesture or a facial expression. Optionally, the gesture comprises a movement of one of an arm, a hand, or a finger. Optionally, an attention model is also applied to the actions of the artificial agent and the background actions of the artificial agent. Optionally, the user's eye movements and gaze are tracked using the user's eye angle and / or the user's head angle. Optionally, the simulated embodied interaction between the artificial agent, the user, and the digital content includes the user directing attention to an object in the digital content in the agent's virtual space. Optionally, the method includes tracking the gaze of the artificial agent, and the simulated embodied interaction is further based on a dyadic turn-taking model applied to the user's gaze and the gaze of the artificial agent. Optionally, the artificial agent indicates the end of its turn in the interaction by attempting direct gaze with the user. Optionally, the artificial agent indicates the end of its turn during the interaction by verbal cues. Optionally, the artificial agent indicates the end of its turn during the interaction by non-verbal cues. Optionally, the non-verbal cues include smiling, blinking, head movements, and body movements including arms, hands, and fingers. Optionally, the artificial agent perceives that the user indicates the end of their turn during the interaction when the user begins direct gaze with the artificial agent. Optionally, the artificial agent perceives that the user indicates the end of their turn when the user uses verbal cues. Optionally, the artificial agent perceives that the user indicates the end of their turn when the user uses non-verbal cues. Optionally, the non-verbal cues include smiling, blinking, head movements, and body movements including arms, hands, and fingers. Optionally, the interaction occurs in a virtual reality environment and / or an augmented reality environment. Optionally, the artificial agent is not visible on the display during at least a portion of the interaction.

Claims

1. 1. A method for interacting with an artificial agent and digital content on an end-user display device, comprising: Simulating artificial agents in an agent virtual space; displaying said artificial agent and said agent virtual space to an end user in front of an end user display device; receiving an image of said end user from a camera; generating one or more cue saliency maps using one or more cues from said end user selected from the group of: verbal, auditory or textual input, touchscreen input, mouse movements, visual input corresponding to said end user's gaze, visual input corresponding to said end user's gestures, and visual input corresponding to said end user's facial expressions; generating an object saliency map representing the digital content; and using a weighted sum of the object saliency map and one or more of the cue saliency maps to guide the gaze of the artificial agent. method.

2. The method of claim 1 , wherein the embodied agent simulator for simulating interaction with the end user is further programmed to receive an image of the user's space from the camera.

3. 3. The method of claim 1 or 2, further comprising representing digital content perceivable by the artificial agent by simulating the digital content within the agent virtual space, wherein the digital content is visible to the end user.

4. 4. The method of any one of claims 1 to 3, wherein visual input corresponding to said end user's gestures is provided by tracking said end user's eye movements and by tracking said end user's eye movements on said display to track said end user's gaze.

5. The method of any one of claims 1 to 4, wherein objects to which both the artificial agent and the end user pay attention increase in salience in the object salience map.

6. The method of claim 5 , wherein the salience factor comprises a weighted salience attention factor based on salience attention maps of the end user and the artificial agent.

7. The method of any one of claims 1 to 6, wherein the attention model is also applied to the end user's eye movements, the end user's actions, objects in the end user's environment, and actions in the end user's environment.

8. The method of claim 1 , wherein the attention model is also applied to end-user input.

9. The method of claim 8 , wherein the gesture comprises a movement of one of an arm, a hand, or a finger.

10. 1. An embodied agent simulator for simulating interaction with an end user implemented on an electronic computing device, comprising: simulating the embodied artificial agent within an agent virtual space and displaying the artificial agent and the embodied agent virtual space to an end user in front of an end user display; receiving an image of the end user from a camera; generating one or more cue saliency maps using one or more cues from the end user selected from the group of: verbal, auditory or textual input, touchscreen input, mouse movements, visual input corresponding to the end user's gaze, visual input corresponding to the end user's gestures, and visual input corresponding to the end user's facial expressions; generating an object saliency map representing the digital content; and guiding the gaze of the artificial agent using a weighted sum of the object saliency map and one or more of the cue saliency maps; An embodied agent simulator.

11. The embodied agent simulator of claim 10 , wherein the embodied agent simulator is further programmed to receive an image of the end user's space from the camera.

12. 12. The embodied agent simulator of claim 10 or 11, wherein the embodied agent simulator is further programmed to represent digital content perceivable by the artificial agent by simulating the digital content within the agent virtual space, the digital content being visible to the end user.

13. 13. The embodied agent simulator of any one of claims 10 to 12, wherein visual inputs corresponding to the end user's gestures track the end user's gaze by tracking the end user's eye movements and by tracking the end user's eye movements on the display.

14. An embodied agent simulator according to any one of claims 10 to 13, wherein objects to which both the artificial agent and the end user pay attention increase in salience in the object salience map.

15. The embodied agent simulator of claim 14 , wherein the salience factor comprises a weighted salience attention factor based on salience attention maps of the end user and the artificial agent.

16. 16. The embodied agent simulator of any one of claims 10 to 15, wherein the attention model is also applied to eye movements of the end user, actions of the end user, objects in the end user's environment, and actions in the end user's environment.

17. An embodied agent simulator according to any one of claims 10 to 16, wherein the attention model is also applied to end-user input.

18. The embodied agent simulator of claim 17 , wherein the gesture comprises a movement of one of an arm, a hand, or a finger.

Citation Information

Patent Citations

  • Multimodal interface device and multimodal interface method

    JP1998301675A

  • Information processor

    JP1999073289A

  • Interface

    JP2000200125A

  • Presentation of results of visual attention modeling

    JP2016530595A

  • Glove interface object

    JP2017530452A