Interactive visual content for interactive systems and applications

By adopting technologies such as standardized interaction classification schemes and multimodal human-computer interaction, the challenges of multimodal interaction system development and deployment are solved, more natural and flexible human-computer interaction is achieved, and the authenticity and complexity of the user experience is improved.

CN120066326APending Publication Date: 2025-05-30NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411751043.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-09
Filing Date
2024-12-02
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively support the development and deployment of multimodal interaction systems, especially with challenges in providing flexible and robust user interaction.

Method used

The use of standardized interactive classification schemes, multimodal human-computer interaction, reverse channel mechanism, event-driven architecture, interaction flow management, language model deployment, sensing processing and action execution is adopted. Through the interactive modeling language and interaction modeling API, the use of interactive modeling languages ​​and interpreters is supported to achieve more complex and subtle human-computer interaction.

Benefits of technology

It realizes more natural and flexible human-computer interaction, improves the development and deployment efficiency of multimodal interaction systems, and enhances the authenticity and complexity of user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066326A_ABST
    Figure CN120066326A_ABST
Patent Text Reader

Abstract

The invention discloses interactive visual content for interactive systems and applications. In various examples, an interactive agent platform hosting development and / or deployment of an interactive agent may use a GUI service to perform interactive visual content actions and generate a corresponding GUI. The interaction modeling API may use an interaction classification scheme that defines a standardized format that specifies an event (e.g., a visual information scene, a visual selection, or a visual form action) that indicates visual content coverage that supplements a conversation with the interactive agent. A GUI service may translate a standardized representation of a GUI specified by an interactive modeling API event into a modular GUI configuration defining blocks of visual content specified by the event, and may use the blocks to populate a (e.g., template or shell) visual layout of a GUI overlay layout. Thus, a visual layout representing the GUI specified by the interactive modeling API event may be generated and presented (e.g., via a user interface server).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 604,721, filed on November 30, 2023, the content of which is incorporated herein by reference in its entirety. BACKGROUND OF THE INVENTION

[0003] Conversational artificial intelligence (AI) allows a computer to engage in natural language conversations with a user, thereby facilitating human-like interactions and understanding. Multimodal conversational AI integrates multiple communication modalities (e.g., text, speech, gestures, emotions, visual elements, etc.), thereby allowing for more comprehensive and natural interactions between a user and an AI system. Multimodal conversational AI is being integrated into an increasing number of applications, from simple chatbots to voicebots, to fully interactive avatars (also known as digital characters or avatars) and robots. However, designing these applications to interact flexibly and robustly with users is a challenging task.

[0004] First, designing compelling avatar interactions is very challenging. Interactions with avatars are increasingly perceived as interactions with another person, but interactions with other people are complex, subtle, multimodal, and discontinuous. Since we humans spend our entire lives communicating with other people, we can usually quickly sense whether there is a sense of disharmony, unease, or incongruity in a conversation, or when our interlocutor responds in an unusual or unnatural way. These types of nuances are not easy to model, and the deficiencies in modeling human interactions actually become more apparent the closer one gets to modeling reality. A similar effect is known as the "uncanny valley" effect in three-dimensional (3D) graphics, where 3D models of humans are very close to being lifelike but still exhibit minor flaws or differences from real humans, which can appear creepy and cause feelings of unease or discomfort.

[0005] In addition, to illustrate the complexity in many of these design challenges, consider what it takes to upgrade a chatbot that can only interact through turn-based text conversations into a multimodal interactive avatar that you can see and talk to. Changing from a single interaction modality (e.g., text conversation) to supporting multiple user input interaction modalities (e.g., text, touch, speech, gesture, emotion, etc.) and / or supporting multiple output interaction modalities in response to the user (e.g., through text / voice, graphical user interface (GUI), animation, sound effects, lighting, etc.) adds significant complexity. Additionally, changing from a turn-based interaction system to a system that supports non-sequential interaction (e.g., multiple simultaneous, potentially overriding inputs and / or outputs) adds even more complexity. In many cases, interaction systems that provide a single interaction modality or use turn-based interaction are simply not suitable for multimodal and / or non-sequential interaction systems.

[0006] For some interaction systems (e.g., systems that provide interactive avatars), it may be desirable to support speech input and output and also utilize the screen space by displaying dynamic information on the screen and having the user interact with that information. Therefore, it may be desirable to dynamically adapt the visual presentation on the screen to the content of the conversation to provide useful context information (e.g., by showing visual representations of some of the options that the avatar verbally presents to the user). Nowadays, conversational AI models are customized to handle verbal input and output (e.g., speech in text form), but lack the ability to directly generate the corresponding visual elements or graphical user interface. This is just one example illustrating the limited capabilities of conventional tools in supporting multimodal interaction.

[0007] In addition, AI systems that provide multimodal dialogue experiences come in a variety of different forms, and different systems rely on a variety of different technologies. This means that most interactive systems use custom application programming interfaces (APIs) and architectures customized for each specific interactive system to connect their constituent components (e.g., decision-making units, AI models such as deep neural networks (DNNs) and machine learning models, cameras, user interfaces, etc.) in an application-specific manner. Today, there are a large number of toolkits and frameworks for modeling dialogue interactions, and many different applications are built on top of these technologies. As a result, components cannot be easily swapped or updated based on the latest technologies, which leads to an increase in the time from research to product. In addition, heterogeneous systems that represent multimodal interactions in different ways make it more difficult to train AI models on historical multimodal interactions, thus limiting their ability to improve the user experience over time. Moreover, in many systems, the interaction data is closely related to the specific implementation of the interactive system. For example, the specific format used by any given interactive system to encode or represent interaction data (e.g., how humans talk to a bot) typically depends on the specific implementation. This makes it difficult to reason about multimodal interactions without understanding the technical complexity of any given interactive system, thus limiting the ability to leverage existing frameworks or extend existing technologies.

[0008] Accordingly, there is a need for improved systems to provide and support the development and / or deployment of multimodal interactive systems. Summary of the Invention

[0009] Embodiments of the present disclosure relate to the development and deployment of interactive systems, such as systems that implement an interactive agent (e.g., a bot, an avatar, a digital human, or a robot). For example, systems and methods for implementing or supporting an interactive modeling language and / or an interactive modeling application programming interface (API) are disclosed, the interactive modeling language and / or the interactive modeling API using a standardized interaction classification scheme, multimodal human-computer interaction, a backchanneling mechanism, an event-driven architecture, interaction flow management, the deployment of one or more large language models, sensory processing and action execution, interactive visual content, animation of an interactive agent (e.g., a bot), expectation actions and signaling, and / or other features.

[0010] For example, an interactive agent platform that hosts the development and / or deployment of an interactive agent (e.g., a bot or robot) can provide an interpreter or compiler that interprets or executes code written in an interactive modeling language, and a designer can provide custom code written in the interactive modeling language for the interpreter to execute. The interactive modeling language can be used to define an interaction flow that indicates which actions or events the interpreter (e.g., an event-driven state machine) generates in response to a detected and / or executed sequence of human-machine interactions. An interaction classification scheme can use standardized action keywords and classify interactions by standardized interaction modalities (e.g., BotUpperBodyMotion) and / or corresponding standardized action categories or types (e.g., BotPose, BotGesture), and the interactive modeling language can use keywords, commands, and / or syntax that incorporate or classify the standardized modalities, action types, and / or event syntax defined by the interaction classification scheme. Thus, the flow can be used to model bot intent or inferred user intent, and a designer can use the bot intent or inferred user intent to build more complex interaction patterns with the interactive agent.

[0011] In some embodiments, one or more flows can implement the logic of an interactive agent and can specify a sequence of multimodal interactions. For example, an interactive avatar (e.g., an animated digital character) or other bot can support any number of simultaneous interaction modalities and corresponding interaction channels for interacting with a user, such as channels for character or bot actions (e.g., speech, gesture, pose, movement, sound burst, etc.), scene actions (e.g., two-dimensional (2D) GUI overlay, 3D scene interaction, visual effects, music, etc.), and user actions (e.g., speech, gesture, pose, movement, etc.). Actions based on different modalities can occur sequentially or in parallel (e.g., waving and saying hello). Thus, the interactive agent can execute any number of flows that specify sequences of multimodal actions (e.g., different types of bot or user actions) using any number of the supported interaction modalities and corresponding interaction channels.

[0012] To make the conversation with an avatar or other interactive agent feel more natural, some embodiments employ a backchannel mechanism to provide feedback to the user when the user is speaking or doing something detectable. For example, the backchannel mechanism can be implemented by triggering an interactive agent pose (e.g., based on the user or avatar speaking, or based on the avatar waiting for the user's response), such as pose mirroring (e.g., where the interactive avatar substantially mirrors the user's pose), short vocal bursts like "yes", "aha", or "hmm" when the user is speaking (e.g., signaling to the user that the interactive agent is listening), gestures (e.g., shaking the head of the interactive bot or robot), and / or other means. Thus, designers can specify various backchannel mechanism techniques to make the conversation with the interactive agent feel more natural.

[0013] In some embodiments, the platform hosting the development and / or deployment of the interactive system can use a standardized interaction modeling API, plug-ins, and / or an event-driven architecture to represent and / or communicate human-computer interactions and related events. In an example implementation, the standardized interaction modeling API serves as a common protocol, where components of the interaction system use a standardized interaction classification scheme to represent all activities of the bot and the user as actions in a standardized form, represent the states of the multimodal actions of the bot and the user as events in a standardized form, implement a standardized mutual exclusion modality that defines how to resolve conflicts between actions in the standardized action categories (e.g., it is impossible to say two things at the same time, while it is possible to say something and make a gesture at the same time), and / or implement a standardized protocol for any number of standardized modalities and actions, regardless of the implementation.

[0014] In some embodiments, the interpreter of the interactive agent can be programmed to iterate one or more flows until reaching an event matcher. The top-level flow can specify one or more instructions to activate any number of flows including any number of event matchers. The interpreter can use any suitable data structure to track the active flows and the corresponding event matchers (e.g., using a tree or other representation of the nested flow relationship), and the interpreter can employ an event-driven state machine to listen for various events and trigger the corresponding actions specified in the matching flows (which have event matchers that match the incoming interaction modeling API events). Thus, the interpreter can execute a main processing loop that processes the incoming interaction modeling API events and generates the outgoing interaction modeling API events for the interactive agent.

[0015] In some embodiments, the interaction modeling language and corresponding interpreter can support the use of natural language descriptions and one or more language models (e.g., large language models (LLMs), vision language models (VLMs), multimodal language models, etc.) to reduce the cognitive burden on programmers and facilitate the development and deployment of more complex and nuanced human-computer interactions. For example, the interpreter can parse one or more specified flows that define the logic of an interactive agent (e.g., at design time), identify if any of the specified flows lack a corresponding flow description, and if so, prompt the language model to generate a flow description based on the name and / or instructions of the flow. Additionally or alternatively, the interpreter can identify if any of the specified flows lack an instruction sequence, and if so, prompt the language model to generate an instruction sequence. In some embodiments, the interpreter can use one or more target event parameters and / or one or more parameter values generated using the language model to determine if an event matches an active event matcher of any active flow, can prompt the language model to determine if an event matches a flow description, can prompt the language model to determine if a non-matching event matches the name and / or one or more instructions of an active flow, can prompt the language model to generate a flow in response to a non-matching event, and / or otherwise.

[0016] Typically, an interactive agent platform that hosts interactive agents (e.g., chatbots, voicebots, digital assistants, interactive avatars, non-player characters (NPCs), digital humans, robots, etc.) can support any number of input and output interaction channels. In some embodiments where sensing processing, interaction decision-making, and action execution are decoupled, the interactive agent platform can support a sensing server for each input interaction channel and an action server for each output interaction channel. The sensing server of the corresponding input interaction channel can convert input or non-standard technical events into a standardized format and generate corresponding interaction modeling API events, and the interaction manager can process these incoming interaction modeling API events and generate outgoing interaction modeling API events representing commands to take certain actions. The action server of the corresponding output interaction channel can interpret these outgoing interaction modeling API events and execute the corresponding commands. A combination of asynchronous event loops and processes can be used to implement the sensing server and / or the action server to ensure that multiple user sessions and system pipelines can be served in parallel. To handle all supported actions of at least one interaction modality, the action server can be equipped with action handlers for each standardized action category or type and / or action event supported by the interaction modeling language and / or defined by the interaction classification scheme of a given interaction modality. Each action server can manage the lifecycle of all actions within its jurisdiction and can synchronize action state changes with designated conditions (e.g., waiting for the previous action of the same modality to complete before starting an action, aligning the completion of two different actions of different modalities, aligning the start of one action with the end of some other action, etc.).

[0017] In some embodiments, an interactive agent platform that hosts the development and / or deployment of an interactive agent can use a graphical user interface (GUI) (or generally a UI) service to perform interactive visual content actions and generate a corresponding GUI. For example, an interaction modeling API can use a standardized interaction classification scheme that defines a standardized format (e.g., standardized and semantically meaningful keywords) that specifies events related to interactive visual content actions (e.g., actions indicating an overlay or other visual content arrangement to supplement a conversation with an interactive agent), such as visual information scene (e.g., displaying non-interactive content, such as images, text, and video, next to an interaction) actions, visual selection (e.g., presenting visual selections to a user in the form of multiple selectable buttons or a list of options) actions, and / or visual form (e.g., presenting a visual web form to a user to input user information) actions. A sensory server can translate detected interactions with GUI interaction elements into standardized interaction modeling API events that represent possible interactions with these elements in a standardized format. The standardized interaction modeling API events can be processed by an interpreter that implements the logic of the interactive agent to generate outgoing interaction modeling API events that specify commands for responding to GUI updates. An action server that implements the GUI service can convert a standardized representation of a particular GUI specified by a particular interaction modeling API event into a representation (e.g., JavaScript Object Notation (JSON)) of a modular GUI configuration that defines visual content blocks specified or otherwise represented by the interaction modeling API event, such as paragraphs, images, buttons, multiple selection fields, and / or other types. Thus, the GUI service can use these blocks to populate the (e.g., template or shell) visual layout of a GUI overlay (e.g., a HyperText Markup Language (HTML) page that can be rendered in a web browser) with the visual content specified by the interaction modeling API event. Thus, a visual layout representing the GUI specified by the interaction modeling API event can be generated and presented (e.g., via a user interface server) to the user.

[0018] In some embodiments, an interaction modeling API event that specifies a command to make a bot expression, gesture, or other interaction or movement (e.g., executed by an interpreter to run code written in an interaction modeling language) can be generated and translated into a corresponding bot animation. More specifically, an interpreter implementing the logic of an interactive agent can use a standardized interaction classification scheme to generate an interaction modeling API event representing the target bot expression, gesture, or other interaction or movement, and an action server implementing an animation service can use a standardized representation of the target bot movement to identify a corresponding supported animation or generate a matching animation on the fly. The animation service can implement an action state machine and an action stack for all events related to a specific interaction modality or action category (e.g., bot gesture), connect to an animation graph of a state machine that implements the transitions between animation states and animations, and instruct the animation graph to set corresponding state variables based on commands that change the state of the bot movement represented by the interaction modeling API event (e.g., initialize, stop, or resume).

[0019] In some embodiments, an interpreter associated with an interactive agent can generate an interaction modeling API event that conveys an expectation that certain events will occur and commands or otherwise triggers corresponding preparatory actions, such as turning down the speaker volume in anticipation of user speech, enabling computer vision and / or machine learning algorithms in anticipation of a visual event, and / or signaling to the user that the interactive agent is waiting for input (e.g., on a designated user interaction modality). The interaction modeling API event can include one or more fields that represent the expectation that a specified target event will occur using a standardized interaction classification scheme that identifies the expectation as a supported action type (e.g., ExpectationBotAction, ExpectationSignalingAction) and represents the corresponding expected event (e.g., indicating an expected state such as start, stop, and complete), the expected target event (e.g., UtteranceUserActionStarted), and / or the expected input interaction modality (e.g., UserSpeech) using standardized (e.g., natural language, semantically meaningful) keywords and / or commands.

[0020] Thus, the present technology can be used to develop and / or deploy interactive bots or robots (e.g., chatbots, voicebots, digital assistants, interactive avatars, non-player characters (NPCs), digital humans, etc.) that engage in more complex, nuanced, multimodal, non-sequential, and / or more realistic conversational AI and / or other types of human-machine interactions than the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The present system and method for developing and deploying an interaction system will be described in detail below with reference to the accompanying drawings, in which:

[0022] Figure 1 illustrates an example interaction system according to some embodiments of the present disclosure;

[0023] Figure 2 illustrates an example interaction modeling API according to some embodiments of the present disclosure;

[0024] Figure 3 illustrates an example interaction system that can be supported by the example interaction modeling API and / or the example interaction modeling language according to some embodiments of the present disclosure;

[0025] Figure 4 illustrates an example modality policy according to some embodiments of the present disclosure;

[0026] Figure 5 illustrates an example interaction classification scheme according to some embodiments of the present disclosure;

[0027] Figure 6 illustrates an example event-driven interaction system according to some embodiments of the present disclosure;

[0028] Figure 7 illustrates an example interaction manager according to some embodiments of the present disclosure;

[0029] Figure 8 is a flowchart showing an example event-driven state machine for an interaction manager according to some embodiments of the present disclosure;

[0030] Figure 9 illustrates an example action server according to some embodiments of the present disclosure;

[0031] Figure 10 illustrates an example event flow through the example action server according to some embodiments of the present disclosure;

[0032] Figure 11 illustrates an example action lifecycle according to some embodiments of the present disclosure;

[0033] Figures 12A - 12DShows some example action handlers for an example GUI service and an example animation service according to some embodiments of the present disclosure;

[0034] Figures 13A - 13F Shows some example interactions with visual selection according to some embodiments of the present disclosure;

[0035] Figure 14A Shows an example graphical user interface presenting an interactive avatar and interactive visual content according to some embodiments of the present disclosure, and Figures 14B - 14L Shows some example layouts of visual elements of interactive visual content according to some embodiments of the present disclosure;

[0036] Figure 15 Shows an example event flow of user speech actions in an implementation where a user is speaking to an interactive avatar according to some embodiments of the present disclosure;

[0037] Figure 16 Shows an example event flow of user speech actions in an implementation where a user is speaking to a chatbot according to some embodiments of the present disclosure;

[0038] Figure 17 Shows an example event flow of bot expected actions in an implementation where a user is speaking to an interactive avatar according to some embodiments of the present disclosure;

[0039] Figure 18 Is a flowchart showing a method for generating a representation of a responsive agent action classified using an interaction classification scheme according to some embodiments of the present disclosure;

[0040] Figure 19 Is a flowchart showing a method for generating a representation of a responsive agent action based at least on executing one or more interaction flows according to some embodiments of the present disclosure;

[0041] Figure 20 Is a flowchart showing a method for triggering an interactive avatar to provide backchannel mechanism feedback according to some embodiments of the present disclosure;

[0042] Figure 21 Is a flowchart showing a method for generating an interaction modeling event for commanding an interactive agent to execute a responsive agent or scenario action according to some embodiments of the present disclosure;

[0043] Figure 22 Is a flowchart showing a method for triggering one or more responsive agent or scenario actions specified by one or more matching interaction flows according to some embodiments of the present disclosure;

[0044] Figure 23is a flowchart showing a method for generating a response agent or scenario action based at least on prompting one or more large language models;

[0045] Figure 24 is a flowchart showing a method for generating one or more outgoing interaction modeling events indicating that one or more action servers execute one or more response agents or scenario actions;

[0046] Figure 25 is a flowchart showing a method for generating a visual layout representing an update specified by an event;

[0047] Figure 26 is a flowchart showing a method for triggering an animation state of an interactive agent;

[0048] Figure 27 is a flowchart showing a method for performing one or more preparatory actions;

[0049] Figure 28A is a block diagram of an example generative language model system suitable for implementing at least some embodiments of the present disclosure;

[0050] Figure 28B is a block diagram of an example generative language model including a transformer encoder-decoder suitable for implementing at least some embodiments of the present disclosure;

[0051] Figure 28C is a block diagram of an example generative language model including a decoder-only transformer architecture suitable for implementing at least some embodiments of the present disclosure;

[0052] Figure 29 is a block diagram of an example content streaming system suitable for implementing some embodiments of the present disclosure;

[0053] Figure 30 is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and

[0054] Figure 31 is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. Detailed Description

[0055] Systems and methods related to the development and deployment of interactive systems (e.g., interactive systems implementing interactive agents (e.g., bots, non-player characters, digital avatars, digital humans, robots, etc.)) are disclosed. For example, systems and methods for implementing or supporting an interactive modeling language and / or an interactive modeling API are disclosed, which use a standardized interaction classification scheme, multimodal human-computer interaction, backchanneling mechanisms, event-driven architectures, interaction flow management, deployment using one or more language models (e.g., LLM, VLM, multimodal language models, etc.), sensing processing and action execution, interactive visual content, interactive agent animation, anticipated actions and signaling, and / or other features.

[0056] Introduction . At a high level, an interactive agent platform can be used to author and / or execute interactive agents (e.g., chatbots, voicebots, digital assistants, interactive avatars, non-player characters, robots, etc.) that participate in conversational AI or other types of human-computer interaction. When designing such a platform and / or an interactive system for an interactive agent, it can be insightful to consider some possible features that can contribute to compelling human-computer interaction and interaction flow.

[0057] Multimodality is a factor that contributes to compelling human-computer interaction. For example, when designing an interactive avatar experience, a designer may want to support many different output interaction modalities or ways of interacting with the user. The designer may want their avatar to be able to speak, gesture, display something in a GUI, make a sound, or otherwise interact. Similarly, the designer may want to support different types of input interaction modalities or ways for the user to interact with the system. For example, the designer may want to support detecting and responding to when the user verbally answers a question by selecting an item on the screen or making a gesture such as a thumbs up to confirm a selection. One possible implication of multimodality is that the designer may want to have flexibility in how to temporarily align interactions. For example, the designer may want the avatar to say something while performing a gesture, or may want to initiate a gesture at a specific moment when the avatar has said something specific. Thus, supporting different types of independently controllable interaction modalities can be desirable.

[0058] Backchanneling is a useful tool for facilitating effective human communication. It helps convey active listening and engagement, signaling to the speaker that their message has been heard and understood. This feedback loop makes the conversation flow more smoothly, helps build connections, and encourages people to continue talking and sharing their thoughts. Designers may want their avatars to attempt to use backchanneling mechanisms to make the avatars seem more human and the interactions more natural, so supporting backchanneling mechanisms can be desirable.

[0059] Some designers may wish to support non-linear interactions. Designers often try to avoid the feeling of predictable, guided, or simple interactions, which can make users feel like they are following a pre-determined or set route that lacks spontaneity or freedom. Even if the required customer journey inherently contains a degree of linearity, it is desirable to support interactions in a way that allows users to break free from strict logic.

[0060] Proactivity can be a useful feature to achieve. Nowadays, many users are accustomed to voice assistants, but the conversation patterns with these digital assistants are usually very simple. Users initiate a conversation with a wake word and ask questions or provide commands. The voice assistant responds to this prompt by directly performing an action, answering a question, or following up with a clarifying question. While this interaction pattern may be effective for retrieving information or setting a timer, it is not very engaging and is generally not suitable for more complex use cases. Instead, designers may wish for their avatars to be proactive, rephrase questions if the user doesn't understand, guide them back to a process if they deviate from the conversation, or provide alternative ways to complete a task. Proactivity is very helpful in preventing interactions from becoming boring (in cases where the user disengages from the conversation or doesn't know how to continue).

[0061] Some designers may wish to leverage the capabilities of language models (e.g., LLMs, VLMs, etc.). For example, designers may wish for an avatar or chatbot to use an LLM to make its interactions with users more natural and adapt to the current interaction context. Some LLM uses may help avoid common pitfalls in the avatar or chatbot experience, such as when the bot repeats the same answer over and over again, or when a simple question doesn't elicit the expected response. In an interactive avatar setup, designers may wish to use an LLM to help create verbal and / or non-verbal responses, such as gestures or facial expressions, or may even wish to use an LLM to help provide useful information on a GUI. Therefore, supporting various LLM uses can be desirable.

[0062] Interactive Modeling Language and Interaction Classification Scheme . Generally, human-computer interactions and related events can be represented and communicated in various ways within an interactive system or an interactive agent platform that hosts the development and / or deployment of the interactive system.

[0063] One possible way to represent and / or convey an interaction is to use an interaction modeling language that uses a standardized interaction classification scheme to specify user and / or bot interactions and related events. Existing dialogue management techniques (such as flowcharts, state machines, and frame-based systems) do not have the ability to model highly flexible dialogue flows (such as those that might be expected from a realistic interactive avatar). In contrast, a standardized interaction classification scheme can provide a semantically meaningful way to classify, specify, and convey the desired interactions and interaction flows. For example, an interactive agent platform can provide an interpreter or compiler to interpret or execute code written in the interaction modeling language, and a designer can provide custom code written in the interaction modeling language for the interpreter to execute. An interaction modeling language using a standardized interaction classification scheme facilitates many technical advantages, including reducing the designer's workload by reducing the cognitive load on the designer when developing an interactive system, supporting various interactions or features (such as those described above) that a designer can use to customize the interactive system, and promoting interoperability through standardized interaction representations.

[0064] Consider the possible goal of reducing the cognitive load on developers when writing code to implement an interactive system. Existing programming languages require developers to write functions implementing interactions using common keywords and commands. However, some embodiments abstract some of the lower-level programming to support a more semantically intuitive representation of interactions: interaction flows. Interactions typically occur in the form of a flow, so an interaction modeling language can be used to define the flow of an interaction. A flow can be considered similar to a function, but can be composed of primitives that contain semantically meaningful (e.g., natural language) keywords and commands that use an interaction classification scheme to specify events (e.g., something happened) and actions (e.g., something needs to happen). Thus, an interaction flow can be used as a mechanism to indicate which actions or events an interpreter (e.g., an event-driven state machine) should generate in response to a detected and / or executed sequence of human-machine interactions.

[0065] In some embodiments, an interaction classification scheme can classify interactions according to standardized interaction modalities (e.g., BotUpperBodyMotion) and / or corresponding standardized action categories or types (e.g., BotPose, BotGesture) using standardized action keywords. The scheme can support any number and type of interactions or communication methods (e.g., user-system interactions, bot-user interactions, bot expected actions and expected signaling, scenario actions, etc.). Standardized event keywords, commands, and / or syntax can be used to represent the state of an action (e.g., the state of an observed user action, the current state of a bot or scenario action) and / or commands to change the state of a bot or scenario action. For example, an event specifier with standardized syntax (e.g., an event name and / or identifier including keywords identifying a standardized action category or type, and a specifier of the user or bot action state) can be used to represent an action event (e.g., the start or stop of a user or bot action).

[0066] In some embodiments, an interaction modeling language can use keywords, commands, and / or syntax that incorporate or classify the standardized modalities, action types, and / or event syntax defined by the interaction classification scheme. For example, an instruction line in a flow can include: an event trigger (e.g., using a keyword such as send), which causes an interpreter to generate a specified event when a certain specified condition is met (e.g., an event representing a command to execute a bot action can trigger the action to be performed, and an event representing a change in the state of a detected user action can trigger a corresponding bot action); or an event matcher (e.g., using a keyword such as match), which causes the interpreter to interrupt the flow and monitor a specified event before resuming the flow. The event trigger and event matcher can use an event specifier to specify the corresponding trigger and match conditions, the event specifier including a standardized event name or identifier (e.g., a keyword identifying a standardized action category or type, paired with a corresponding action state specifier or a command to change the action state) and arguments specifying one or more conditions that the specified event must meet (e.g., using predefined parameters and supported values, or a natural language description). In some embodiments, when the event specifier includes an action but omits the state (e.g., the name of the action can be specified as a shortcut for specifying action completion), the interpreter can infer the specified action state (e.g., Finished).

[0067] Take the UserSpeech (user speech) modality and the corresponding Utterance User Action as an example. Suppose the user issues an utterance that is recognized by the interaction system. Possible examples of this type of action include the user typing in a text interface to interact with the bot or the user speaking to an interaction avatar. This action can be classified as a user utterance, and the action events supported by this action can include UtteranceUserActionStarted (the user starts generating an utterance) or UtteranceUserActionFinished (the user's utterance has been completed). An example flow instruction for waiting for the user to say a specific content might be "Match UtteranceUserActionFinished (text = "How are you?", speed = "slow", volume = "normal")". In this example, the event identifier is a camelcase keyword that links the standardized action category (UtteranceUserAction) with the representation of the specified action state (Finished).

[0068] In some embodiments, the interaction modeling language and the corresponding interpreter may support any number of keywords for parallelizing action and flow execution and matching (e.g., send, match, start, stop, await, activate). Compared with conventional dialogue modeling languages where statements are always considered in sequential order, some embodiments may support a keyword (e.g., start) that instructs the interpreter to start a specified action in a specified (e.g., standardized) action category or flow and continue iterating its parent flow without waiting for the started action or sub-flow to complete, some embodiments may support a keyword (e.g., stop) that instructs the interpreter to stop a started action or sub-flow, and some embodiments may support a keyword (e.g., await) that instructs the interpreter to wait for a started action or sub-flow to complete and then advance the parent flow. In some embodiments, the interpreter may use other keywords (e.g., send, match) to implement some keywords (e.g., start, await) to send or wait for events to occur. In some implementations, once the flow has started, the interpreter executes all actions in the specified flow until the first matching statement. Subsequently, when the statement matches, the interpreter may execute the subsequent actions in the specified flow until the next matching statement or the end of the flow, and so on, until the flow is completed.

[0069] In some cases, a designer may wish for a sub - flow to automatically restart after completion. This can be useful for certain types of flows, such as those that attempt to trigger certain actions that depend on recurring events. Thus, some embodiments may support a keyword (e.g., "activate") that instructs the interpreter to automatically restart the flow after completion. In some embodiments, if the activated flow does not contain an event matcher, the interpreter will run the flow only once but keep it in an active state, so any sub - flows will also remain active.

[0070] Some embodiments may support a keyword that instructs the interpreter to complete a flow (e.g., "return") or abort a flow (e.g., "abort"), and the flow may instruct the interpreter to determine and return a certain value. Since some embodiments support multiple active flows, some implementations of the interpreter start a top - level, root, or main flow (e.g., at startup), which acts as the parent of all other flows. This hierarchy enables better abstraction and encapsulation capabilities than prior art. In some embodiments, an event matcher command can accept a specified name or identifier of a flow and a specified flow event (e.g., "start", "completed", "failed", "paused", "resumed") as arguments, and the interpreter can use these arguments as instructions to match the corresponding flow event.

[0071] Thus, in some embodiments, all flows represent corresponding interaction patterns. In some such embodiments, flows can be used to model bot intents or inferred user intents, and designers can use these bot intents or inferred user intents to build more complex interaction patterns. In some such implementations, the flow effectively describes the expected interaction pattern. If the interpreter starts a flow, it can designate that flow as active and attempt to match the patterns of the contained event matcher statements with events representing the ongoing interaction. Whenever the interpreter determines that a matching statement is satisfied by an event, the interpreter can advance the corresponding flow head to the next matching statement, executing all non - matching statements in between. Thus, the interpreter can be programmed to sequentially execute the instructions specified in the flow, generate any events specified by the event triggers, and stop when the flow head reaches an event matcher, an exception, or the end of the flow. To illustrate how flows can be used to implement various types of interaction patterns and features, consider the following example use cases.

[0072] Multimodal Interaction. In some embodiments, one or more flows may specify multimodal interaction sequences. While conventional chatbots use turn-based conversations, interactive avatars (e.g., animated digital characters) or other bots may support any number of interaction modalities and corresponding interaction channels to interact with users, such as channels for character or bot actions (e.g., speech, gestures, poses, movements, sound bursts, etc.), scene actions (e.g., two-dimensional (2D) GUI overlays, 3D scene interactions, visual effects, music, etc.), and user actions (e.g., speech, gestures, poses, movements, etc.). Thus, a flow may use any number of supported interaction modalities and corresponding interaction channels to specify multimodal action sequences (e.g., different types of bot or user actions).

[0073] For example, consider the following example flow that wraps a start bot utterance action command to improve programming readability and simplicity:

[0074] flow bot say $utterance

[0075] SendBotUtteranceStartAction(action_id, "$utterance")

[0076] MatchBotUtteranceActionFinished(action_id)

[0077] Now, the start bot utterance action command can be triggered with an instruction that simply writes its name and specifies the desired utterance: bot say "Hello world". In this example, when the interpreter executes this instruction, it looks for a flow named "bot say" that defines: an event trigger that generates an event which, when executed, will start a bot utterance action with the specified text (in this example, "Hello world"); and an event matcher that waits for the bot utterance action to complete. The following is an example flow that similarly wraps a start bot gesture action command:

[0078] flow bot gesture $gesture

[0079] SendBotGestureStartAction(action_id, "$gesture")

[0080] MatchBotGestureActionFinished(action_id)

[0081] By defining such a wrapper stream, a designer can use a single instruction to have the bot display a gesture (e.g., trigger a start bot gesture action command) that simply writes the name of the wrapper stream and specifies the desired gesture, e.g., the bot gesture "Wave with both hands".

[0082] Conceptually, actions based on different modalities can occur sequentially or in parallel (e.g., waving and saying hello). Therefore, it is desirable to provide the designer with precise timing control over the supported actions and their mutual alignment. For example, consider bot actions such as bot speech and bot gestures. In some embodiments, the stream can specify calling these actions in sequence as follows:

[0083] The bot says "Hello, world"

[0084] The bot gesture "Wave with both hands"

[0085] In this example, since the wrapper stream for the bot say action includes an event matcher that tells the interpreter to wait for the action to complete before advancing the stream, the bot gesture only starts after the "bot says" (or "wait for bot says") action is complete.

[0086] Since these two actions are in two different modalities, some embodiments can allow them to be executed simultaneously. One way to trigger the simultaneous execution of these two actions is to combine them into an "and" group (e.g., defined with a keyword such as "and") to start them in parallel:

[0087] The bot says "Hello, world"

[0088] And the bot gesture "Wave with both hands"

[0089] More complex action groups can be defined using an "or" group definition (e.g., defined with a keyword such as "or"), such as the following example:

[0090] The bot says "Hello, world"

[0091] And (the bot gesture "Wave with both hands" or the bot gesture "Smile")

[0092] Therefore, the designer can use an "or" action group to specify alternative actions. In this example, the resulting actions will be the bot saying "Hello, world" and waving or saying "Hello, world" and smiling. Another example way to execute two actions in parallel is to use a keyword such as "start" to start the two actions in parallel:

[0093] The bot starts by saying "Hello, World"

[0094] The bot starts with the gesture "Wave hands"

[0095] In some implementations of these examples, the interpreter does not wait for any one action to complete before proceeding to the next statement. To explicitly wait for a started action to complete, the flow can specify a "match" statement on the completed event of a previously started action, as shown in the following example:

[0096] The bot starts by saying "Hello, World" as $action

[0097] Match $action.Finished()

[0098] Using a flow like this example, the designer can associate an action with the end of another action by stopping the action using a keyword such as "stop", thereby restricting the lifetime of the action. For example, the following example will stop the bot from making gestures after it has finished speaking.

[0099] The bot starts by saying "Hello, World" as $action_1

[0100] Associate with the bot gesture "Wave hands" as $action_2

[0101] Match $action_1.Finished()

[0102] Stop $action_2

[0103] The foregoing examples focus on actions initiated by the bot. However, to provide meaningful interaction with the user, it is desirable to react to user actions. For example, consider the following example flow that wraps an event matcher for the event indicating that a user utterance action has been completed:

[0104] Flow: User says $text

[0105] Match UtteranceUserAction.Finished(final_transcript = $text)

[0106] Now, this flow wrapper (or the wrapped event matcher) can be used to indicate that the flow is waiting for a specific user utterance:

[0107] User says "Hi there!"

[0108] As with some of the example bot actions above, in some embodiments, the flow that includes an event matcher for this user action will progress to the next statement only after the user speaks the specified parameter value ("Hi there"). In some embodiments, the event matcher can trigger a user action group, as shown in the following example:

[0109] The user says "Hi there!"

[0110] Or the user says "Hello"

[0111] Or the user says "Hi"

[0112] In some embodiments, the flow that includes this example event matcher will wait for one of the user actions to occur before proceeding to the next statement in the flow. The flow can additionally or alternatively use an "and" group to wait for multiple user actions to occur before proceeding.

[0113] In some embodiments, a flow can be defined with an instruction that includes a keyword (e.g., "flow"), a name or identifier for the flow (e.g., "How are you response"), and some parameters (e.g., starting with a $ symbol), and the values of these parameters can be specified and passed when the flow is called, as shown in the following example:

[0114] flow How are you response$text

[0115] The user says "How are you?"

[0116] The bot says $text

[0117] In some embodiments, each flow defines an action scope. For example, if the interpreter triggers the initiation of any action during the flow and these active actions are not completed when the interpreter finishes executing the flow, the interpreter can stop these active actions. Returning to the "Hello, World" example, in some embodiments, there is no need to stop the gesture action because it will automatically stop when the process is completed:

[0118] flow Hello, World example

[0119] Start The bot says "Hello, World" as $action_1

[0120] With the bot gesture "Wave hands" $action_2

[0121] Match $action_1.Finished()

[0122] BackchannelingConversations with traditional chatbots or avatars often feel rigid or unnatural because they typically enforce a strict turn-taking conversation. To make conversations with an avatar feel more natural, some embodiments employ a technique called a backchannel mechanism, where an interaction system (such as an interactive avatar) provides feedback to the user when the user is speaking or doing something detectable.

[0123] One way to implement the backchannel mechanism is to use poses. For example, a designer may want the avatar to maintain a certain pose depending on whether the user or the avatar is speaking, or when the avatar is waiting for a response from the user. The following is an example flow that can be used to implement a listening pose:

[0124] Flow bot listening pose

[0125] When true

[0126] The user starts speaking

[0127] Start the bot pose "listening" as $listening

[0128] The user said something

[0129] Send $istening.Stop()

[0130] In this example, "The user starts speaking" is a flow wrapper of an event matcher for an event indicating the start of a user speech action, and "The user said something" is a flow wrapper of an event matcher for an event indicating the completion of a user speech action. After enabling such an example flow in some implementations, whenever the user starts speaking, the avatar listens (e.g., shows a listening animation), and when the user stops speaking, the avatar stops the animation.

[0131] Another example may include various other poses, such as "speaking", "focused", and / or "idle", to give the user feedback on the current state of the avatar, as shown in the following example:

[0132] Flow manage bot pose

[0133] Start the bot pose "idle" as $current_posture

[0134] When true

[0135] When the user starts speaking

[0136] Send $current_posture.Stop()

[0137] Start the bot pose "listening" as $current_posture

[0138] Or when the bot says something

[0139] Send $current_posture.Stop()

[0140] Start the bot posture "talking" as $current_posture

[0141] Or when the bot says something

[0142] Send $current_posture.Stop()

[0143] Start the bot posture "focused" as $current_posture

[0144] In some implementations, such an example flow is enabled, and the avatar will have an idle posture until the user starts speaking (in which case it adopts a listening posture), the avatar starts speaking (in which case it adopts a speaking posture), or the avatar has just finished saying something (in which case, it adopts a focused posture).

[0145] In some embodiments, the backchannel mechanism can be implemented using short bursts of sound (such as "yes", "aha", or "hmm") when the user is speaking. This can signal to the user that the avatar is listening and can make the interaction seem more natural. In some embodiments, a non-verbal backchannel mechanism can be used to enhance this effect, where the avatar reacts to something the user has said, for example, by making a gesture. The following is an example flow of implementing the backchannel mechanism using sound bursts and gestures:

[0146] The flow where the bot reacts to something sad

[0147] When True

[0148] The user mentions something sad

[0149] The bot gestures "shakes head" and the bot says "oh"

[0150] The flow where the bot reacts to something nice

[0151] When True

[0152] The user mentions something nice

[0153] The bot gestures "celebrates things going well" and the bot says "good"

[0154] In some implementations, whenever the user mentions something nice or sad, both of these streams produce brief vocal bursts and small hand gestures. In this example, unlike the "user said something" stream that waits for a complete utterance, the "user mentioned something" stream can be defined as a partial transcription that matches what was said while the user was still speaking (and thus reacts to it).

[0155] The following are example streams that use these two bot backchannel mechanism streams in an interaction sequence:

[0156] Start bot reaction to sad things

[0157] Start bot reaction to nice things

[0158] Start managing bot pose

[0159] The bot says "How's your day going?"

[0160] When the user says r".*terribl|horribl|bad.*"

[0161] The bot says "I hope you're feeling better now?" with the bot gesture "show concern for the user"

[0162] Otherwise

[0163] The bot says "That's great. Do you have any plans for the rest of the day?"

[0164] User said something

[0165] The bot says "Thanks for sharing"

[0166] Here, after activating the example bot backchannel mechanism stream, the bot asks the user how their day is going. If the user tells the bot that something good or bad has happened, the bot immediately reacts with a vocal burst and a short animation. These are some high-level examples of example implementations based on the interpreter, and other variations can be implemented within the scope of this disclosure. Other examples and features of possible interaction modeling languages and interaction classification schemes will be described in more detail below.

[0167] Event - Driven Architecture and Interaction Modeling API. In some embodiments, a development and / or deployment platform for an interaction system (e.g., an interactive agent platform) may use a standardized interaction modeling API and / or an event-driven architecture to represent and / or communicate human-computer interactions and related events. In some embodiments, the standardized interaction modeling API standardizes the way components represent multimodal interactions, thus achieving a high degree of interoperability between components and the applications that use them. In an example implementation, the standardized interaction modeling API serves as a common protocol, where components use a standardized interaction classification scheme to represent all activities of bots and users as actions in a standardized form, represent the states of multimodal actions of users and bots as events in a standardized form, implement a standardized mutually exclusive modality that defines how to resolve conflicts between standardized action categories or types (e.g., it is impossible to say two things at the same time, while it is possible to say something and make a gesture at the same time), and / or implement a standardized protocol for any number of standardized modalities and action types, regardless of the implementation.

[0168] In some embodiments, an interactive agent platform that hosts the development and / or deployment of an interaction system may implement an architectural pattern that separates components that implement decision logic (e.g., interpreters) from components that perform (e.g., multimodal) interactions. For example, an interaction manager may implement an interpreter for an interaction modeling language as a different event-driven component (e.g., an event-driven state machine). The interface of the interaction manager may use a standardized interaction modeling API that defines a standardized form for representing action categories, specifying action instances within an action category, events, and contexts. A sensing server for a corresponding input interaction channel may convert input or non-standard technical events into a standardized format and generate corresponding interaction modeling API events (also referred to as interaction modeling events). The interaction manager may process these incoming interaction modeling API events, determine what actions should be taken (e.g., based on code written in the interaction modeling language for the interpreter to execute), and generate outgoing interaction modeling API events that represent commands for taking certain actions (e.g., in response to instructions in the interaction modeling language such as "send"). An action server for a corresponding output interaction channel may interpret these outgoing interaction modeling API events and execute the corresponding commands. Decoupling these components enables interchangeability and interoperability, facilitating development and innovation. For example, one component can be replaced with another design, or another interaction channel can be connected with little impact on the operability of the existing system.

[0169] The architecture pattern and API design can provide a purely event-driven and asynchronous way to handle multimodal interactions. Compared with previous solutions, in some embodiments, there is no strict concept of rotation (e.g., bot speaks, user speaks, bot speaks). Instead, the participants in the interaction can simultaneously participate in the multimodal interaction, independently and concurrently acting and reacting to incoming events, thus enhancing the realism of human-computer interaction.

[0170] In some embodiments using this architecture pattern, the interaction manager does not need to know which specific action servers are available in the interaction system. It is sufficient for the interaction manager to know the supported modalities. Similarly, the action server and / or the sensing server can be independent of the interaction manager. Therefore, any of these components can be upgraded or replaced. Thus, the same platform and / or interaction manager can support different types of interaction systems, all of which are controlled through the same API and can be swapped in and out or customized for a given deployment. For example, one implementation can provide a text-based user interface, while another implementation can provide a voice-based system, and a third implementation can provide a 2D / 3D avatar.

[0171] Management of Multiple Streams The above example illustrates how to program the example interpreter to iterate through any specific flow until it reaches the event matcher. In some embodiments, the top-level flow can specify instructions to activate any number of flows containing any number of event matchers. Therefore, the interpreter can use any suitable data structure to track the active flows and the corresponding event matchers (e.g., using a tree or other representation of nested flow relationships), and can employ an event-driven state machine to listen for various events and trigger the corresponding actions specified in the matching flows (which have event matchers that match the incoming interaction modeling API events).

[0172] Since a flow can specify human-computer interaction, a designer may wish to activate multiple flows that specify conflicting interactions triggered under different conditions, and / or multiple flows that specify the same interaction (or different but compatible interactions) triggered based on the same or similar conditions. In some scenarios, multiple active flows specifying various interactions can be triggered by different conditions satisfied by the same event. Thus, an interpreter can process incoming interaction modeling API events (e.g., from a queue) sequentially, and for each event, test whether the event matcher specified by each active flow matches the event. If an event matcher in an active flow matches the event (matching flow), the interpreter can advance that flow (e.g., generate an outgoing interaction modeling API event to trigger an action). If there are multiple matching flows, the interpreter can determine whether the matching flows agree on the action. If they agree, the interpreter can advance both matching flows. If they do not agree, the interpreter can apply conflict resolution to determine which action should take precedence, advance the matching flow with the preferred action, and abort the other matching flows (e.g., because the interaction modalities represented by these flows will no longer apply). If there is no active flow that matches the event, the interpreter can generate a matching and trigger a predefined flow to dispose of internal events for unmatched or unhandled events, can run one or more unhandled event handlers, and / or can use other techniques to handle unhandled events. After checking for matches and advancing flows, the interpreter can check the flow states of any completed or aborted flows and can stop any active flows activated by these completed or aborted flows (e.g., because the interaction modalities represented by these flows should no longer apply). Thus, the interpreter can iterate through the events in the queue, advance flows, perform conflict management to determine which interactions to execute, and generate outgoing interaction modeling API events to trigger these interactions.

[0173] Accordingly, the interpreter can execute a main processing loop that processes incoming Interactive Modeling API events and generates outgoing Interactive Modeling API events. Compared to a simple event-driven state machine, the interpreter can use a set of flow heads. A flow can be regarded as a program containing a sequence of instructions, and a flow head can be regarded as an instruction pointer that advances through the instructions and indicates the current position within the corresponding flow. Depending on the instruction, the interpreter can advance any given flow head to the next instruction, jump to another flow referenced by a label or other flow identifier, fork into multiple heads, merge multiple flow heads together, and / or otherwise. Thus, the interpreter can use flow heads to construct and maintain a hierarchy of flow heads. If a parent flow head in a branch of the hierarchy of flows or flow heads is stopped, paused, or resumed, the interpreter can stop, pause, or resume all child flow heads of that parent flow head or branch. In some embodiments, any flow can specify any number of scopes, and the interpreter can use these scopes to generate events that indicate that the corresponding action server will have started an action and limit the lifecycle of the flow to the corresponding scope.

[0174] In some embodiments, advancing a flow can indicate to the interpreter to generate Interactive Modeling API events indicating certain actions. Additionally or alternatively, advancing a flow can indicate to the interpreter to generate Interactive Modeling API events that notify listeners that certain events have occurred. Accordingly, the interpreter can emit these events, and / or the interpreter can maintain an internal event queue, place these events in the internal event queue, and process any internal events in the internal event queue in sequence (e.g., test whether an activity flow matches an internal event) before advancing to process the next incoming Interactive Modeling API event.

[0175] Example Interpreter Language Model Use . In some embodiments, the Interactive Modeling language and the corresponding interpreter can support the use of natural language descriptions and the use of one or more language models (e.g., LLM, VLM, multimodal LLM, etc.) to reduce the cognitive burden on programmers and facilitate the development and deployment of more complex and nuanced human-machine interactions.

[0176] For example, each flow can be specified with a corresponding natural language description that summarizes the interaction pattern represented by the flow. In some embodiments, the interpreter does not require the designer to specify these flow descriptions, but can use the flow descriptions in certain cases (e.g., used by an unknown event handler that prompts the LLM to determine whether an unmatched event representing an unrecognized user intent semantically matches the natural language description of an activity flow representing the target user intent). Thus, in some embodiments, the interpreter can parse one or more specified flows (e.g., at design time), identify whether any of the specified flows lack a corresponding flow description, and if so, prompt the LLM to generate a flow description based on the name and / or instructions of the flow. Additionally or alternatively, the interpreter can (e.g., prompt the LLM) determine whether any specified flow description is inconsistent with its corresponding flow description, and if so, prompt the LLM to generate a new flow description (e.g., as a suggestion or for automatic replacement) based on the name and / or instructions of the flow.

[0177] In some embodiments, the designer can specify a flow description (e.g., a natural language description of what the flow should do) without an instruction sequence, or can call a flow by name without defining it. Thus, in some embodiments, the interpreter can parse one or more specified flows (e.g., at design time), identify whether any of the specified flows lack an instruction sequence, and if so, prompt the LLM to generate an instruction sequence (e.g., based on the name and / or description of the flow). For example, the interpreter can provide the LLM with one or more example flows, the specified name and / or description of the flow, and a prompt to complete the flow based on the name and / or description of the flow. These are just a few examples of possible ways the interpreter can call the LLM.

[0178] In an example implementation, flow instructions (e.g., including any encountered event triggers) can be executed until an event matcher is reached, at which point the flow can be interrupted. When there are no more flows to advance, incoming or internal events can be processed by executing the event matcher in each interrupted flow and comparing the event with the target event parameters and parameter values specified by the event specifier of the event matcher. Generally, any suitable matching technique can be used to determine whether an event matches the active event matcher of any active flow (e.g., comparing the target event parameters and parameter values with the parameters and parameter values of the incoming or internal event to generate some representation of whether the event matches).

[0179] Typically, a designer can use the name or identifier of an event to specify the event to be matched or triggered and one or more target event parameters and / or parameter values. The target event parameters and / or parameter values can be explicitly specified using positional or named parameters, or specified as a natural language description (NLD) (e.g., a docstring), and the interpreter can use this natural language description to infer the target event parameters and / or values (e.g., based on a single NLD for all target event parameters and values, based on the NLDs of individual parameter values). The following are some example event specifiers for the following:

[0180] · Event with named parameters and explicit values:

[0181] StartUtteranceBotAction(text = "How are you?", volume = "whisper", speed = "slow")

[0182] · Event with positional parameters:

[0183] StartUtteranceBotAction("How are you?", "whisper", "slow")

[0184] · Event with parameters specified using NLD values:

[0185] StartUtteranceBotAction(text = """Asking how it is going""", volume = "whisper", speed = "slow")

[0186] · Event with a single NLD parameter:

[0187] StartUtteranceBotAction("""Ask the user how it is going, very slowly and in a whisper""")

[0188] In some embodiments that support event specifiers using NLDs, before executing an instruction (e.g., an event matcher or event trigger) that includes an event specifier, the interpreter can (e.g., at runtime) determine whether the instruction includes an NLD parameter, and if so, prompt the LLM to generate the corresponding target event parameters and / or parameter values. In this way, the interpreter can use the generated target event parameters and / or parameter values to execute the instruction (e.g., an event trigger or event matcher).

[0189] Additionally or alternatively, the interpreter can prompt the LLM (e.g., at runtime) to determine whether an event (e.g., an interaction modeling API event) matches the flow description of the active flow. Generally, interaction modeling API events can represent user interactions or intents, bot interactions or intents, scenario interactions, or some other type of event using a standardized interaction classification scheme that classifies actions, action events, event parameters, and / or parameter values using standardized (e.g., natural language, semantically meaningful) keywords and / or commands. Thus, the interpreter can execute an event matcher by determining whether the received actions, action events, event parameters, and / or parameter values of an incoming or internal event match (e.g., exactly or approximately) the event specified by the event matcher. Additionally or alternatively, the interpreter can prompt the LLM to determine whether the representation of an incoming or internal event matches the (e.g., specified or generated) flow description of the active flow. Depending on the implementation, the LLM can provide a more nuanced or semantically understanding match than conventional express or fuzzy matching algorithms.

[0190] For example, assume the user makes some gesture indicating consent, such as giving a thumbs up, nodding, or saying something informal like "yeah". The designer may have written a flow that is designed to match the scenario where the user indicates consent, but only provided a few examples of verbal responses for exact matching. In this scenario, even without an exact match, the LLM may be able to determine that the standardized and semantically meaningful representation of the detected user response (e.g., GestureUserActionFinished("thumbs up")) is a semantic match for the flow description (e.g., "user indicates consent"). This is another example where the designer has specified a flow that is designed to match (via "user has selected option" and "user said" flow wrappers) the event that the user has selected option B from a list of options:

[0191] Flow user selection multimodal display

[0192] “““User has selected display (B).””””

[0193] User has selected option "multimodal"

[0194] or the user says "Show me the multimodal display"

[0195] or the user says "multimodal"

[0196] or the user says "Display B"

[0197] or the user says "The second display"

[0198] If the user uses option B for a certain interaction choice not anticipated by the designer, the LLM can be used to determine that the standardized representation of the detected user gesture matches the flow (e.g., the natural language description of the flow, the natural language description of the parameters, the natural language description of the parameter values, etc.). In this way, the specified flow can match multiple gestures, text responses, or other events that the designer may not have explicitly specified.

[0199] In some implementations (e.g., in some embodiments, the interpreter checks the event matchers of all active (e.g., interrupted) flows to find a match and determines that no active flow matches the incoming or internal event), the interpreter can (e.g., at runtime) prompt the LLM to determine whether the representation of the incoming or internal event and / or the recent interaction history matches the name and / or instructions of the active flow. For example, some flows may represent the target user intent, and the interpreter can implement an event handler for the unknown user action by providing the LLM with example interactions between the user and the bot, a list of some possible target flows for the target user intent, the corresponding list of target user intents, the recent interaction history, the unknown user action, and a hint for the LLM to predict whether the unknown user action matches one of the target user intents. Thus, the interpreter can use the LLM to implement an unknown event handler that provides a more granular or semantic understanding of matching the specified target user intent.

[0200] In some scenarios, there may be no matching flow defined for the bot's response to a specific user interaction. Thus, in some implementations (e.g., in some embodiments, the interpreter determines that there is no active flow that matches an incoming or internal event representing a user interaction), the interpreter can prompt the LLM to generate a flow (e.g., at runtime). For example, in some embodiments, the interpreter can first use the LLM to attempt to match an unknown incoming or internal event with the name, instructions, and / or other representations of one or more active flows that listen for a corresponding target user intent (and define the corresponding bot response), and if the LLM determines that there is no matching flow (target user intent), the interpreter can prompt the (same or some other) LLM to generate a response proxy (e.g., bot) flow. In some embodiments, the interpreter can prompt the LLM to generate one or more intents as an intermediate step. For example, if the unknown event is a user action, the interpreter can apply any number of prompts to instruct the LLM to classify the unknown user action as a user intent, generate a response proxy intent, and / or generate a flow that implements the response proxy intent. By way of non-limiting example, the interpreter can implement an event handler for an unknown user action by providing the LLM with sample interactions between the user and the bot, the most recent interaction history, the unknown user action, and prompts for the LLM to predict one or more intents (e.g., user, bot) and / or prompts for the LLM to generate a corresponding flow. Thus, the interpreter can use the LLM to implement an unknown event handler that intelligently responds to unknown events without the designer having to specify the code for the response flow.

[0201] Generally, neural networks operate like black boxes, which poses an obstacle to controlling the generated responses. The lack of transparency makes it challenging to ensure that the generated content is accurate, appropriate, and ethical. However, using the LLM to autocomplete event parameters or parameter values, perform event matching, or generate flows using a standardized and structured interaction modeling language and / or interaction classification scheme helps impose structure and interpretability on what the LLM does, thereby enhancing the ability to control the LLM output. Thus, embodiments that use the LLM to autocomplete event parameters or parameter values, perform event matching, or generate flows relieve the cognitive burden on the designer when developing an interactive system by providing an intuitive way to specify the human-machine interactions and events to be matched or triggered, while preventing the generation of unexpected content, making the designer's job easier.

[0202] Sensing Processing and Action Execution. According to embodiments and configurations, an interactive agent platform that hosts the development and / or deployment of interactive agents (e.g., chatbots, voicebots, digital assistants, interactive avatars, non-player characters (NPCs), digital humans, robots, etc.) can support any number of input and output interaction channels. In some embodiments where sensing processing, interaction decision-making, and action execution are decoupled, the interactive agent platform can support a sensing server for each input interaction channel and an action server for each output interaction channel. The sensing server for a corresponding input interaction channel can convert input or non-standard technical events into a standardized format and generate corresponding interaction modeling API events, the interaction manager can process these incoming interaction modeling API events and generate outgoing interaction modeling API events representing commands to take certain actions, and the action server for the corresponding output interaction channel can interpret these outgoing interaction modeling API events and execute the corresponding commands. Communicating between these components using the interaction modeling API enables distributing the responsibility for handling different types of input processing to different types of sensing servers and distributing the responsibility for different types of actions to different types of action servers. For example, each action server can be responsible for a corresponding group of actions and action events (e.g., associated with a common interaction modality), thus avoiding the complexity of having to manage events associated with different interaction modalities.

[0203] A combination of asynchronous event loops and processes can be used to implement the sensing server and / or the action server to ensure that multiple user sessions and system pipelines can be served in parallel. This architecture allows programmers to add different services that can handle different types of actions and events (corresponding to different types of interaction modalities) supported by the interaction modeling API actions. In some embodiments, an event gateway can be used to communicate and distribute events to the corresponding components, whether through synchronous interactions (e.g., via REST API, Google Remote Procedure Call (RPC), etc.) or asynchronous interactions (e.g., using a message or event broker). Thus, each sensing server can emit interaction modeling API events to the event gateway for any incoming input or non-standard technical event, and the interaction manager can be subscribed to or otherwise configured to pick up these events from the event gateway. The interaction manager can generate outgoing interaction modeling API events and forward them to the event gateway, and each action server can be subscribed to or otherwise configured to pick up those events that it is responsible for executing (e.g., one interaction modality per action server).

[0204] To handle all support actions for at least one interaction modality, the action server can be equipped with action handlers for each standardized action category and / or action event supported by the interaction modeling language and / or defined by the interaction classification scheme of a given interaction modality. For example, the action server can implement: a chat service that handles all interaction modeling API events for bot utterance actions; an animation service that handles all interaction modeling API events for bot gesture actions; a graphical user interface (GUI) service that handles all interaction modeling API events indicating the arrangement of visual information, such as visual information scene actions, visual selection actions, and / or visual form actions; and / or a timer service that handles all interaction modeling API events for timer actions; just to name a few examples.

[0205] Each action server can manage the lifecycle of all actions within its purview. Interaction modeling API events can specify commands for the action server to initiate, modify, or stop an action. Thus, a common action identifier (e.g., action_uid) can be used to represent all events related to the same action, such that the individual events associated with the same action identifier can represent different states in the lifecycle of the corresponding action. Thus, an action server for a particular interaction modality can start a particular action (e.g., a bot gesture or utterance) and can track the active action and its corresponding state. Each action server can enforce a modality policy that determines how to handle an action triggered during the execution of another action of the same interaction modality (e.g., multiple sound effects may be allowed to run simultaneously, but a new body animation can replace or temporarily override the active body animation). Some implementations can support commands to modify a running action, which can be useful for longer-running actions (e.g., avatar animations) that can dynamically adjust their behavior. For example, a nodding animation can be modified to change its speed based on the detected level of speech activity. Some implementations can support commands to stop a running action, which can be used to proactively stop actions that may run for a long period of time (e.g., gestures). In some embodiments, the action server can synchronize action state changes with prescribed conditions (e.g., waiting to start an action until a previous action of the same modality is completed, aligning the completion of two different actions of different modalities, aligning the start of one action with the end of some other action, etc.). When the action server implements an action state change, it can generate an interaction modeling API event reflecting the update and forward the interaction modeling API event to the event gateway so that any component listening for or waiting for that state change can respond to it.

[0206] Interactive Visualization GUI Elements. In some scenarios, a designer may wish to customize an interactive system, such as a system with an interactive avatar, that synchronizes conversational AI with supplementary visual content (such as visual representations of relevant information (e.g., text, images), choices presented to the user, or fields or forms for the user to complete).

[0207] Thus, in some embodiments, an interaction modeling API can use a standardized interaction classification scheme that defines a standardized format (e.g., standardized and semantically meaningful keywords) to specify events related to interactive visual content actions of standardized categories (e.g., actions indicating an overlay or other visual content arrangement that supplements a conversation with an interactive agent), such as visual information scenario actions, visual selection actions, and / or visual form actions. Some embodiments can incorporate an interaction modeling language that supports specifying visual designs using natural language descriptions (e.g., for an alert message, "attention-grabbing, bold, and professional"), and a corresponding interpreter can translate the specified description into a standardized representation of the corresponding design elements (e.g., color scheme, typography, layout, image) and generate outgoing interaction modeling API events using the standardized format of interactive visual content action events. Thus, an action server can implement a graphical user interface service that generates robust and visually appealing GUIs that can be synchronized with the spoken responses of conversational AI or otherwise facilitate human-machine interaction.

[0208] In some embodiments, the interaction modeling API defines a way to represent a specific GUI (e.g., the configuration or arrangement of visual elements) using an interaction classification scheme that defines standardized categories of interactive visual content actions and corresponding events with payloads specifying standardized GUI elements. For example, the interaction classification scheme can classify interactive visual content actions and / or GUI elements into semantically meaningful groups such that an interpreter or action server can generate the content of a given GUI element based on the current context of the interaction (e.g., generate a text block using an LLM, retrieve or generate an image based on a specified description). Each group of interactive visual content actions and / or GUI elements can be used to define a corresponding subspace of possible GUIs that represent different ways the bot can visualize information for the user and / or different ways the user can interact with that information. Example interaction classification schemes can classify interactive visual content actions as visual information scenario actions, visual selection actions, and / or visual form actions.

[0209] Visual information scenario actions can include displaying information to a user for informational purposes (e.g., text with background information about a topic or product, an image illustrating a situation or problem), e.g., where it is not expected that the user can interact with the information in any other way than by reading it. Visual selection actions can include displaying visual elements or interacting with visual elements that present choices to the user and / or describe the type of choices (e.g., multiple-choice versus single-choice, small or limited set of options versus large set of options). Visual form actions can include displaying visual elements or interacting with visual elements that request input from the user in some form or field (e.g., an avatar may wish to ask the user to provide their email address) and / or describe the type of input requested (e.g., email, address, signature).

[0210] In some embodiments, an interaction classification scheme may define a standardized format for specifying supported GUI interaction elements (e.g., a list of buttons, a grid of selectable options, an input text field, a prompt carousel) such that a sensing server (e.g., its corresponding action handler) can translate detected interactions with these interaction elements (e.g., the state when a button list element is released, e.g., after a click or a touch, the state when a user types a character into an input field, the state when a user presses the enter key or clicks away from a text box) into standardized interaction modeling API events representing possible interactions with these elements in the standardized format. In some embodiments, for each of multiple different input interaction channels (e.g., GUI interaction, user gesture, speech input, etc.), there may be a sensing server, each sensing server being configured to generate standardized interaction modeling API events representing detected interaction events in the standardized format. In some embodiments, a sensing server may translate detected interaction events (e.g., "user clicks button 'chai-latte', scrolls down and clicks button 'confirm'") into corresponding standardized interaction-level events (e.g., "user selects option 'Chai Latte'"). The standardized interaction-level events may depend on the type of interactive visual content actions defined by the scheme. Example standardized interaction-level events may include updates representing the confirmation state of the user and / or events when an update is detected (e.g., if there is a single input requested as part of a VisualForm, the "enter" keyboard event may be translated into a "confirmed" state update), updates representing the selection of the user and / or events when an update is detected (e.g., the detected selection of item "chai-latte" from a list of multi-select elements may be translated into a selection update), updates representing the form input of the user and / or events when an update is detected, and / or other events. In this way, standardized interaction modeling API events can be generated and forwarded to an event gateway and processed by an interpreter to generate an outgoing interaction modeling API event, which may specify a command for performing a response GUI update, and the outgoing interaction modeling API event may be forwarded to the event gateway for execution by a corresponding action server.

[0211] In some embodiments, an interaction modeling API event that specifies a command for a GUI update can be converted into a corresponding GUI and presented to the user. To achieve this, in some embodiments, an action server implementing a GUI service can convert a standardized representation of a particular GUI specified by a particular interaction modeling API event into a representation (e.g., JavaScript Object Notation (JSON)) of a modular GUI configuration that specifies content blocks such as paragraphs, images, buttons, multi-selection fields, and / or other types. Thus, the GUI service can use these content blocks to populate a visual layout (e.g., a HyperText Markup Language (HTML) layout that can be rendered in any modern web browser) covered by the GUI. For example, any number of template or shell visual layouts can define the corresponding arrangement of the respective content blocks, and the UI service can select a template or shell visual layout (e.g., based on which content blocks have been generated or specified by the interaction modeling API event) and use the corresponding generated content to fill placeholders for these blocks in the template or shell. In some embodiments, various features of the template or shell visual layout can be customized (e.g., resizing or arranging of blocks, appearance options such as the color palette of the GUI overlay, etc.). Thus, a visual layout representing the GUI specified by the interaction modeling API event can be generated and presented (e.g., via a user interface server) to the user.

[0212] Taking an interactive avatar as an example, an animation service can be used to animate the avatar (described in more detail below), and a GUI service can be used to synchronize the representation of related visual elements (e.g., visual information scenes, visual selections, visual forms). For example, the screen of the user's device can include some areas that render the avatar across the entire web page (e.g., using as much of the height and width of the browser window as possible while maintaining the same aspect ratio for the avatar stream), and the visual elements generated by the GUI service can be rendered in an overlay manner on top of the avatar stream. In an example embodiment, the avatar stream can maintain a fixed aspect ratio (e.g., 16:9) and use padding around the stream to maintain the aspect ratio when necessary. In some embodiments, the overlay can remain in the same relative position on the screen regardless of the size of the stream. In some embodiments, the overlay can be scaled with the size of the avatar. In some embodiments, the overlay can maintain a fixed configurable size relative to the size of the avatar (e.g., 10% of the avatar width and 10% of the avatar height).

[0213] In some embodiments, each GUI (e.g., a page of visual elements) can be configured as part of a stack, from which GUI pages can be pushed and popped. This configuration is particularly useful in AI-driven interaction environments because the context during a series of interactions can change in a non-linear manner. GUI stack overlays can be used to ensure that the visual content on the GUI remains relevant throughout the series of interactions. These stacked GUIs can be at least partially transparent to facilitate visualization of the stacked information, enabling conversational AI to combine GUIs or shuffle the stack at different stages of the conversation (e.g., the stacked overlay title can describe the overall customer journey, such as "Support Ticket XYZ", while the stacked pages within the overlay can represent different steps in the journey, such as "Please enter your email"). In some embodiments, the GUI can be part of a rendered 3D scene (e.g., a tablet held by an avatar), the GUI can be 3D (e.g., buttons can be rendered with corresponding depth), and / or otherwise. These are just a few examples, and other variations can be implemented within the scope of the present disclosure. For example, although the foregoing examples are described in the context of 2D GUIs, those of ordinary skill in the art will understand how to adapt the foregoing guidance to present avatars and / or overlays in augmented and / or virtual reality (AR / VR).

[0214] Interactive Agent Animation . In some embodiments, an interaction modeling API event that specifies a command to make a bot expression, pose, gesture, or other interaction or movement (e.g., executed by an interpreter to run code written in an interaction modeling language) can be generated and converted into a corresponding bot animation, and the bot animation can be presented to the user. More specifically, in some embodiments, an action server implementing an animation service can use a standardized representation of the target bot expression, pose, gesture, or other interaction or movement specified by a particular interaction modeling API event to identify and trigger or generate the corresponding animation.

[0215] Taking the standardized bot gesture action category (e.g., GestureBotAction) as an example type of bot actions, in some embodiments, the animation service can handle all events related to actions in the GestureBotAction category, can apply a modal policy (which overrides an active gesture with any subsequently indicated gesture), and can use the incoming StartGestureBotAction event to create an action stack when there is an active GestureBotAction. Thus, the animation service can implement an action state machine and an action stack for all GestureBotActions, connect to the animation graph of the state machine that implements the transitions between animation states and animations, and instruct the animation graph to set the corresponding state variables based on commands that change the state of an instance of GestureBotAction (e.g., initialize, stop, or resume a gesture) represented by an interaction modeling API event.

[0216] In some embodiments, the animation graph can support a certain number of clips that will make different expressions, poses, gestures, or other interactions or action avatars or other bot animations. Thus, the animation service can receive commands that change the state of GestureBotAction (e.g., initialize, stop, or resume a gesture) represented in the standardized interaction classification scheme to identify the corresponding supported animation clips. In some cases, the designer may wish to use natural language descriptions to specify bot expressions, poses, gestures, or other interactions or movements. Thus, in some embodiments, the animation service can use natural language descriptions (e.g., manually specified or generated by an interpreter using an LLM / VLM / etc., used as an argument to describe an instance of a standardized type of bot action in an interaction modeling API event) to select the best or generate animation clips. For example, the animation service can generate or access sentence embeddings for natural language descriptions of bot actions (e.g., bot gestures), use it to perform a similarity search on the sentence embeddings to obtain descriptions of available animations, and use some similarity metric (e.g., nearest neighbor, within a threshold) to select an animation. In some embodiments, if the best match is within the threshold similarity (e.g., the distance is below a specified threshold), then that animation can be played. If no animation matches within the specified threshold, then a fallback animation (e.g., a less specific version of the best-matching animation) can be played. If the animation service cannot identify a suitable match, the animation service can generate an interaction modeling API event indicating that the gesture failed (e.g., ActionFinished(is_success = false, failure_reason = "gesture not supported")) and forward that interaction modeling API event to the event gateway.

[0217] Anticipated Actions and Anticipated Signaling 。In various cases, it may be beneficial to inform an interaction system or one of its components (e.g., a sensing server that controls input processing, an action server that implements bot actions) what events the interaction manager (e.g., an interpreter) expects to receive next from the user or the system. For example, when the interaction manager expects the user to start speaking (e.g., UtteranceUserActionStarted event), the interaction system can configure itself to listen or improve its listening capabilities (e.g., by turning down the speaker volume, increasing the microphone sensitivity, etc.). In a noisy environment, the interaction system can be configured to turn off the listening capabilities (e.g., automatic speech recognition) and only activate listening when the interaction manager expects the user to speak. In a chatbot system, the designer may want to display a thinking indicator when the chatbot (e.g., the interaction manager) is processing a request, and once it expects a response (e.g., a text answer), the interaction manager can communicate this expectation to the action server to update the display with a visual indication that the chatbot is waiting for a response. Additionally, running computer vision algorithms is typically resource-intensive. Therefore, the interaction manager can communicate a representation of the currently expected visual event type at any given point during the interaction, and the interaction system can disable or enable the visual algorithms on the fly. Some example scenarios where disabling and enabling computer vision can be useful include quick response (QR) code reading, object recognition, user movement detection, and so on.

[0218] To facilitate these preparatory actions, instances of what is expected can be represented as instances of standardized action types (expected actions) with corresponding expected states, and the interaction modeling API events associated with a particular instance of an expected action can include one or more fields that represent, using a standardized interaction classification scheme, the expectation that a specified target event will occur. The standardized interaction classification scheme identifies the expectation as a supported action type (e.g., ExpectationBotAction) and uses standardized (e.g., natural language, semantically meaningful) keywords and / or commands to represent the corresponding expected events (e.g., indicating the expected states such as start, stop, and completed) and the expected target events (e.g., UtteranceUserActionStarted). Example standardized expected events can include: an event indicating that the bot expects a specified event on the event gateway in the near future (e.g., StartExpectationBotAction), which can indicate that the sensing or action server optimizes its functionality (e.g., a sensing server responsible for processing camera frames can enable or disable certain vision algorithms based on what the interaction manager expects); an event indicating that the sensing or action server acknowledges that the bot expects or acknowledges that the sensing or action server has updated its functionality in response to the expectation (e.g., ExpectationBotActionStarted); an event indicating that the expectation has stopped (e.g., StopExpectationBotAction), which can occur when the expectation has been met (e.g., an event has been received) or something else has occurred that changes the interaction process; an event indicating that the sensing or action server acknowledges that the bot expectation has been completed (e.g., ExpectationBotActionFinished), and / or other events.

[0219] In addition to or as an alternative to communicating to the sensing or action server the information that the interaction manager (e.g., interpreter) expects certain events to occur, some embodiments signal to the user that the bot is waiting for input (e.g., in a certain user interaction modality). Thus, the standardized interaction classification scheme can classify this expectation signaling as a supported action type (e.g., ExpectationSignalingAction). This action can allow the interaction system to provide the user with subtle (e.g., non-verbal) cues about what the bot expects to receive from the user (e.g., if an avatar is waiting for user input, the avatar's ears can get larger or the avatar can assume a listening pose).

[0220] For example, in a chatbot system, the user may need to input certain information (e.g., "Please enter your date of birth to confirm the order.") before the interaction is considered complete. In such a case, the designer may want the chatbot to signal to the user that it is actively waiting for a response from the user. Thus, the designer can specify code that triggers the generation of a StartExpectationSignalingBotAction (modality = UserSpeech) event. In another example, an interactive avatar may be waiting for a specific gesture from the user. In such a case, the designer may want the avatar to actively communicate this to the user (e.g., by displaying some specified animation). Thus, the designer can specify code that triggers the generation of a StartExpectationSignalingBotAction (modality = UserGesture) event. If there is a conflict with some other ongoing action in the corresponding output interaction channel (e.g., an active upper body animation), the action server can resolve the conflict based on a predefined modality policy.

[0221] To facilitate these expectation signaling actions, interaction modeling API events can use a standardized interaction classification scheme to represent expectation signaling events. This standardized interaction classification scheme classifies expectation signaling into supported action types (e.g., ExpectationSignalingBotAction) and uses standardized (e.g., natural language, semantically meaningful) keywords and / or commands to represent the corresponding expectation signaling events (e.g., which indicate the expected state, such as start, stop, complete) and the target or input interaction modality expected by the bot (e.g., UserSpeech). Example standardized expectation signaling events can include: an event that indicates the bot expects an event on a specified interaction modality at the event gateway in the near future (e.g., StartExpectationSignalingBotAction); an event that indicates the sensing or action server acknowledges an expectation signaling event or acknowledges that the sensing or action server has started actively waiting for an event on a specified interaction modality (e.g., ExpectationSignalingBotActionStarted); an event that indicates the expectation has stopped (e.g., StopExpectationSignalingBotAction); an event that indicates the sensing or action server acknowledges that the expectation has been completed or has stopped actively waiting (e.g., ExpectationSignalingBotActionFinished), and / or other events.

[0222] Accordingly, the present technology can be used to develop and / or deploy interactive agents, such as bots or robots (e.g., chatbots, voicebots, digital assistants, interactive avatars, non-player characters, etc.), which engage in more complex, nuanced, multimodal, discontinuous, and / or lifelike conversational AI and / or other types of human-machine interactions than the prior art. Additionally, various embodiments that implement or support an interaction modeling language and / or an interaction modeling API using a standardized interaction classification scheme facilitate many technical advantages, from making the designer's job easier by reducing the cognitive load on the designer in developing an interaction system, to supporting various interactions or functions that the designer can leverage to customize the interaction system, to promoting interoperability by standardizing the representation of interactions.

[0223] Reference Figure 1 , Figure 1 is an example interaction system 100 according to some embodiments of the present disclosure. It should be understood that such arrangements and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, function groupings, etc.) can be used in addition to the arrangements and elements shown, and some elements can be omitted entirely. Moreover, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and can be implemented in any suitable combination and location. The various functions performed by the entities described herein can be executed by hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in a memory. For example, in some embodiments, the systems and methods described herein can use one or more generative language models (e.g., as Figures 28A - 28C described in), one or more computing devices or their components (e.g., as Figure 30 described in) and / or one or more data centers or their components (e.g., as Figure 31 described in) to implement.

[0224] At a high level, the interaction system 100 can execute, control, or otherwise provide an interactive agent (e.g., a chatbot, voicebot, digital assistant, interactive avatar, non-player character (NPC), digital human, interactive television or other appliance, some other type of interactive robot, etc.). Some example interaction systems that can provide an interactive agent include digital kiosks, automotive infotainment systems, digital assistant platforms, smart TVs or other smart appliances, video game or animation environments, virtual or augmented reality environments, video conferencing systems, and / or others. Figure 1An example implementation is shown, in which a client device 101 (e.g., a smartphone, a tablet, a smart TV, a gaming console, a digital kiosk, etc.) provides an interface for human-computer interaction through any number and type of interaction channels, one or more sensing servers 160 convert the input into events (e.g., standardized interaction modeling API) representing the detected interaction states, an interaction manager 190 determines what actions the interactive agent should take and generates events (e.g., standardized interaction modeling API) representing the corresponding commands, and one or more action servers 170 interpret these commands and trigger the interactive agent to take corresponding actions through the corresponding interaction channels.

[0225] Depending on the implementation, Figure 1 the components of can be implemented on any number of physical machines (e.g., which can include components, features, and / or functions similar to those of Figure 30 the example computing device 3000 of). Taking a digital kiosk as an example. In some embodiments, the physical kiosk can correspond to the client device 101, which is connected to one or more remotely hosted components. In some embodiments, Figure 1 some or all of the components in can be implemented as corresponding microservices and / or physical devices, deployed in a cluster of nodes in a data center (e.g., which can include components, features, and / or functions similar to those of Figure 31 the example data center 3100 of), on one or more edge devices, on dedicated hardware, and / or elsewhere. In some implementations, some or all of the components run locally on certain physical machines (e.g., on a digital kiosk, a robot, or some other interactive system), which has various types of interface hardware managed by an operating system, firmware, and / or other software. In some such embodiments, the client device 101 corresponds to various hardware interfaces, and Figure 1 some or all of the other components in (e.g., the sensing server 160, the action server 170, the interaction manager 190, etc.) represent the functions of the operating system, firmware, and / or other software that send commands or requests to the various hardware interfaces.

[0226] In an example virtual or augmented reality environment, Figure 1The components shown in [the figure] can be implemented on a local device (e.g., an AR / VR headset, a smartphone running a VR / AR application), a cloud server, an edge computing device, dedicated hardware, and / or other devices. In some embodiments, there is one sensing server per input interaction channel (e.g., one sensing server for processing video input, one for processing audio input, one for processing touch input), and / or one action server per output interaction channel (e.g., one action server for processing bot animations, one for processing bot speech, one for processing interactive visual content). In some implementations, some or all of the sensing servers 160 and / or action servers 170 are combined into a single machine and / or microservice that uses the corresponding services to handle the corresponding interaction channels. These are just a few examples, and other configurations and implementations are possible within the scope of the present disclosure.

[0227] In some embodiments, Figure 1 Some or all of the components shown in [the figure] are part of, or at least partially hosted by, a development and / or deployment platform (e.g., an interactive agent platform) for an interaction system. For example, a platform such as (and / or another platform or system, e.g., a platform or system using the Universal Scene Description (USD) data format, e.g., OpenUSD) can host the infrastructure and various functions that provide a framework for developing and / or deploying interactive agents. The platform can provide various creation tools that enable users to create and customize interactive agents, a real-time rendering engine, integration with various services (e.g., computer vision, speech recognition, natural language understanding, avatar animation, speech generation, simulation software, recommendation engines), and / or other components. In some embodiments, Figure 1 Some or all of these tools and / or components shown in [the figure] can be integrated into an application and processed in real time (e.g., using a framework for developing and deploying cloud-native applications, such as Unified Cloud Service Tools). Thus, Figure 1 Some or all of these tools and / or components shown in [the figure] can be deployed as microservices and managed using a platform for orchestrating containerized applications (e.g., NVIDIA FLEET COMMAND TM ). Thus, in some embodiments, these tools and / or components can be used to customize and / or deploy the interaction system 100.

[0228] For example, in some embodiments, the interaction manager 190 may implement an interpreter for an interaction modeling language, and the code implementing the decision logic of the interactive agent may be written in the interaction modeling language, loaded onto the interaction manager 190, or otherwise accessed by the interaction manager 190 and executed by the interaction manager 190. Depending on the desired interactive agent, the corresponding sensing server 160 and / or action server 170 may be connected, configured, and support any number and type of interaction channels. Thus, in some embodiments, a development and / or deployment platform may be used to host the interaction system 100, and the interaction system 100 may implement (e.g., customizable) interactive agents.

[0229] At a high level, a user may operate the client device 101 or some other interaction system including any number of input and / or output interaction channels to interact with or otherwise interface with it. By way of non-limiting example, Figure 1 shown are: a video input interaction channel that includes a camera (not shown) and a vision microservice 110 that uses any known computer vision techniques to detect user gestures; an audio input interaction channel that includes a microphone (not shown) and a speech detection microservice 120 that uses any known speech detection and / or recognition techniques to identify user speech; a video output interaction channel that includes a display screen (not shown) and an animation microservice 140 that uses any known animation techniques to animate the bot (e.g., bot poses, bot gestures, blend shapes, text-to-motion, text-to-animation); an audio output interaction channel that includes a speaker (not shown) and a speech generation microservice 150 that uses any known speech synthesis techniques to synthesize bot speech; a graphical user interface (GUI) that has a GUI input interaction channel that accepts user GUI inputs (e.g., touch, click), a GUI output interaction channel that displays interactive visual content, and a user interface server 130 that uses any known techniques to manage and / or serve the user interface of the GUI.

[0230] In passing Figure 1In an example flow of the interaction system 100, certain representations of user input (e.g., gestures detected by the vision microservice 110, voice commands detected by the voice detection microservice 120, or touch or click inputs detected by the UI server 130) can be forwarded to the corresponding sensing server among one or more sensing servers 160 responsible for the corresponding interaction channels. Thus, the sensing server 160 can convert the user input into a standardized representation of a corresponding event and place the event on the event gateway 180. The event gateway 180 can be used to communicate and distribute events to the corresponding components, either through synchronous interactions (e.g., via REST APIs, Google Remote Procedure Calls (RPCs), etc.) or asynchronous interactions (e.g., using a message or event broker). The interaction manager 190 can be subscribed to or otherwise configured to pick up or receive those events from the event gateway 180. In this way, the interaction manager 190 can process the events (e.g., using an event-driven state machine), determine the interactions to participate in, and generate commands and forward them as corresponding events in the standardized representation to the event gateway 180. The action server 170 responsible for the corresponding interaction channel can be subscribed to or otherwise configured to pick up or receive those events for which it is responsible for execution from the event gateway 180. In this way, the action server 170 can execute, schedule, and / or otherwise dispose of the events of the corresponding interaction modality, engaging with the corresponding services that control the corresponding output interfaces. For example, according to the indicated action, a corresponding one of the action servers 170 in the action server 170 can schedule and trigger (e.g., the voice generation microservice 150 to generate) bot speech in the audio interface, (e.g., the animation microservice 140 generates) bot animations on the display screen or headphones, (e.g., the UI server 130 presents) interactive visual content on the display screen or headphones, and / or others.

[0231] In some embodiments, the interaction system 100 uses a standardized interaction modeling API and / or an event-driven architecture to represent and / or communicate human-computer interactions and related events. In some embodiments, the standardized interaction modeling API standardizes the way components (e.g., the sensing server 160, the action server 170, the interaction manager 190) represent multimodal interactions. In an example implementation, the standardized interaction modeling API serves as a common protocol, where the various components of the interaction system 100 use a standardized interaction classification scheme to represent all activities of the bot, the user, and / or the interaction system 100 as actions in a standardized form, represent states (e.g., the states of multimodal actions from the user and the bot) as events in a standardized form, support standardized mutually exclusive interaction modalities and define how to resolve conflicts between standardized action categories or types, and / or implement a standardized protocol for any number of standardized modalities and action categories independent of the implementation.

[0232] Figure 2FIG. 0 illustrates an example interaction modeling API 220 in accordance with some embodiments of the present disclosure. Generally, different types of interaction systems may include different types of interaction channels 230. For example, a chatbot may use a text interface that supports an input interaction channel for inputting text and an output channel for outputting text. A voice assistant may use an audio interface that supports an input interaction channel for inputting speech and an output channel for outputting speech. An interactive avatar may use a video input interface that supports an input interaction channel for detected gestures, an audio input interface that supports an input interaction channel for detected speech, a video output interface that supports an output interaction channel for avatar animation (e.g., poses, gestures), an audio output interface that supports an output interaction channel for avatar output speech, and / or a graphical user interface that supports an input interaction channel for touch input and / or an output channel for interactive visual content. A non-player character may use a game controller interface that supports an input interaction channel for controller input, a video output interface that supports an output interaction channel for non-player character animation, and an audio output interface that supports an output interaction channel for non-player character output speech. These are only examples, and within the scope of the present disclosure, other types of interaction systems, interactive agents, and / or interaction channels may be implemented.

[0233] Figure 2 FIG. 4 illustrates an example interaction modeling API 220 between an interaction manager 190 and an interaction channel 230. In some embodiments, the interaction modeling API 220 defines a standardized format for specifying user and / or bot interactions, system events, and related events using a standardized interaction classification scheme. The interaction classification scheme may use standardized (e.g., semantically meaningful) keywords, commands, and / or syntax that combine or classify standardized interaction modalities, action types, and / or event syntax. Taking standardized interaction modalities as an example, the interaction classification scheme may be used to classify interactions (e.g., bot actions) according to standardized interaction modalities and / or corresponding standardized action categories (e.g., bot utterances, bot poses, bot gestures, bot gazes) using standardized action keywords. Figure 2 This is illustrated by using separate lines to represent events for different interaction modalities (e.g., bot utterance events, bot pose events, bot gesture events, bot gaze events, scene or interactive visual content events). Additionally, in some embodiments, the interaction modeling API 220 defines a standardized format for specifying changes in action states as corresponding events to support an event-driven architecture. Figure 2 This is illustrated by using different start and stop times for different actions (e.g., the bot starts with a tense pose, then initiates an utterance, and before completing the utterance, initiates gesture and gaze actions, etc.).

[0234] In some embodiments, to facilitate configurability and interoperability, the interaction modeling API 220, the corresponding interaction modeling language supported by the interaction manager 190, and / or the corresponding interaction classification scheme supported by the interaction channels 230 (e.g., the sensing server and / or the action server therein) may provide a way to classify, specify, and represent the interactions of various different interaction systems and the corresponding interaction channels, which may enable designers to customize the interaction systems using standardized components. Figure 3 Illustrates some features of an example interaction system that may be supported by an example interaction modeling API and / or an example interaction modeling language according to some embodiments of the present disclosure. In some embodiments, the interaction system relies on an interaction modeling API and / or an interaction modeling language that supports more interaction and action keywords than those used by the interaction system itself. For example, the interaction modeling API and / or the interaction modeling language may support keywords for bot gestures (e.g., MakeGesture), even though an interaction system (e.g., a chatbot) using the API and / or the modeling language may not use this type of interaction. However, by supporting various multimodal interactions, the interaction modeling API and / or the interaction modeling language can support various interactions or features that designers can utilize to customize the interaction systems, promote interoperability by standardizing the representation of interactions, and make their work easier by reducing the cognitive burden on designers when developing interaction systems.

[0235] In some embodiments, the interaction modeling API and / or the interaction modeling language may support a standardized representation of actions and events for interaction modalities such as speech, gesture, emotion, movement, scene, and / or others. In some embodiments, the interaction modeling API and / or the language may define mutually exclusive interaction modalities such that actions in different interaction modalities can be executed independently of each other (e.g., by the corresponding action server) (e.g., a bot can say something independently of making a gesture). The possibility of simultaneous or conflicting actions in the same interaction modality can be addressed by implementing modality policies for the same interaction modality (e.g., by the corresponding action server). Thus, the action server implementing the interaction modality can use the prescribed modality policies to determine how to execute, schedule, and / or otherwise handle the events of the interaction modality. Figure 4 Illustrates some example modality policies according to some embodiments of the present disclosure. Thus, the interaction modeling API and / or the interaction modeling language may support an interaction classification scheme that defines the supported interaction modalities and the standardized representation of the corresponding actions and events, such as Figure 5The example interaction classification scheme shown in. As shown in this example, some modality groups (e.g., motion) can be subdivided into sets of interaction modalities that can be executed independently of each other (e.g., BotExpression can be animated on the BotFace modality independently of BotPose on the BotUpperBody modality).

[0236] In an interaction system that supports multimodal interaction, information can be exchanged between the user and the interaction system through multiple interaction modalities. Each interaction modality can be implemented through a corresponding interaction channel between the interaction system and the user. In some embodiments, an interaction classification scheme can classify any given action as part of a single interaction modality, although depending on the interaction system, the action server for that interaction modality can map the action to multiple output interfaces (e.g., audio, video, GUI, etc.). For example, the BotUtterance action (indicating that the bot communicates verbally with the user) can be classified as part of the BotVoice modality. In an interaction system that represents the bot as a 3D avatar (e.g., on a 2D screen, on an AR or VR device), the BotVoice modality and / or the BotUtterance action can trigger different types of outputs, such as audio output (e.g., synthesized speech), lip movement (e.g., lip synchronization with speech), and / or text on the user interface (e.g., dialogue subtitles). In another example, the BotMovement action can be classified as part of the BotLowerBody modality and can trigger lower body animation (e.g., walking animation) and audio output (e.g., footsteps).

[0237] Now turning to Figure 6 , Figure 6 FIG. shows an example event-driven interaction system 600 according to some embodiments of the present disclosure. Figure 6 FIG. shows an example implementation of an architectural pattern that separates the component (e.g., interaction manager 640) that implements the decision logic for determining what action to perform from the components (e.g., sensor server 620 and action server 670) that handle the interaction.

[0238] At a high level, a detected input event 610 (e.g., representing some user input such as a detected gesture, voice command, or touch or click input; representing some detected feature or event associated with the user input such as the presence or absence of detected voice activity, the presence or absence of detected typing, detected transcribed speech, detected changes in typing volume or speed; etc.) can be forwarded to a sensing server 620, and the sensing server 620 can convert the detected input event 610 into a standardized input event 630. An interaction manager 640 can process the standardized input event 630 and generate an event representing an indicated bot action (indicated bot action event 650), and an action server 670 can execute the action represented by the indicated bot action event 650. In some embodiments, the interaction manager 640 can generate an internal event 660 representing a change in internal state (e.g., flow state change) or an indicated bot action, and / or the action server 670 can generate an event 665 representing an acknowledgement of a change in action state, any of which can be evaluated by the interaction manager 640 to determine what action to take.

[0239] The interaction manager 640 (which can correspond to Figure 1 and / or Figure 2 the interaction manager 190) can be responsible for determining what action the interaction system 600 should perform in response to a user action or other event (e.g., standardized input event 630, internal event 660, event 665 representing an acknowledgement of a change in action state). The interaction manager 640 can (but need not) interact with the rest of the interaction system 600 (e.g., specifically) through an event-driven mechanism. In practice, while the interaction manager 640 is busy processing an event (e.g., deciding on the next action), other parts of the interaction system 600 can generate other events. Thus, depending on the implementation, the interaction manager 640 can process multiple events one by one or all events at once. In a stateful approach, the interaction manager 640 can maintain the state or context of the user's interaction with the interaction system 600 across multiple interactions in a given session. In a stateless approach, the history of the state or context can be represented with each new event. In some embodiments, there is no shared state between the interaction manager 640 and the rest of the interaction system 600.

[0240] Typically, the interaction system 600 may include any number of interaction managers (e.g., interaction manager 640). In some implementations, the interaction system 600 may include a primary interaction manager with an internal or secondary interaction manager. In an example involving an interactive avatar experience, the primary interaction manager may manage the high-level flow of the human-machine interaction (e.g., various stages such as greeting, collecting data, providing data, obtaining confirmation, etc.), and the primary interaction manager may transfer decision-making authority to one or more secondary interaction managers when applicable (e.g., for complex authentication flows, for interactive Q&A scenarios, etc.). In some implementations, the interaction system 600 may include multiple peer interaction managers, each peer interaction manager handling different types of events. For example, one interaction manager may handle the dialogue logic (e.g., what the bot should say), while a second interaction manager may animate the avatar based on what is said.

[0241] In some embodiments, the interaction between the interaction manager 640 and the rest of the interaction system 600 occurs through different types of (e.g., standardized) events, such as events representing detected input events (e.g., detected input event 630), indicated bot action events (e.g., indicated bot action event 650), and system or context events. Typically, detected input events can be used to represent any event that may be relevant to the interaction, such as the user saying something (e.g., UserSaid), making a gesture (e.g., UserGesture), or clicking on a GUI element (e.g., UserSelection). Bot action events can define what the interaction system 600 should do, such as saying something, playing a sound, displaying something on a display, changing the appearance or pose of the avatar, calling a third-party API, etc. Bot action events can represent a transition in the action lifecycle (e.g., via an instruction to do something (e.g., StartAction)), an indication of when the action starts (e.g., ActionStarted), or an indication of when it is completed (e.g., ActionFinished). System or context events can represent a change to the associated interaction data contained in the interaction system 600 (e.g., ContextUpdate), such as the username, user rights, selected product, device information, etc.

[0242] Thus, the interaction manager 640 can evaluate various types of events (e.g., the normalized input event 630, internal event 660, event 665 representing confirmation of a change in action state), determine which actions to perform, and generate the indicated bot action event 650. In this way, the action server 670 can perform the actions represented by the indicated bot action event 650. For example, the interaction manager 640 can decide that the interaction system 600 should say "Hello!", and after that utterance (e.g., the Say action) is completed, make a specific gesture (e.g., point to the screen and ask something). In some such examples, the interaction manager 640 can generate an event that specifies that when the interaction system 600 finishes saying hello (e.g., by specifying a condition such as ActionFinished(Say)) the gesture should start (e.g., using a keyword such as StartAction(MakeGesture)). As another example, the interaction manager 640 can decide to start a waving animation when the Say(Hello) action starts and stop the animation when the Say(Hello) ends. In some such examples, the interaction manager 640 can specify the conditions (e.g., ActionStarted(Say) and ActionFinished(Say)) when specifying the corresponding instructions for starting and stopping the gesture (e.g., StartAction(MakeGesture(Wave)) and StopAction(MakeGesture(Wave))).

[0243] In some embodiments, the interaction manager 640 implements an interpreter or compiler that interprets or executes code written in an interaction modeling language that uses a normalized interaction classification scheme (e.g. Figure 5The standardized interaction classification scheme shown in specifies user and / or bot interactions and related events. Generally, an interpreter and an interaction modeling language can support any number of keywords that are used to parallelize action and flow execution and matching (e.g., send, match, start, stop, wait, activate). The interaction modeling language can be used to define interaction flows using primitives that include semantically meaningful (e.g., natural language) keywords and commands that specify events (e.g., something happens) and actions (e.g., something needs to happen) using the interaction classification scheme. In some embodiments, events (e.g., standardized input event 630, internal event 660, event 665 indicating confirmation of a change in action state) can be represented using event specifiers having a standardized syntax defined by the interaction classification scheme, the interaction modeling language, and / or the interaction modeling API and supported by the interpreter. In some embodiments, an event can include certain representations of corresponding (e.g., standardized) fields and values (e.g., a payload specifying these representations) that the interpreter (and other components) may be able to understand. Thus, the interpreter can execute code implementing the interaction flow of the interaction modeling language, where the interaction flow can indicate what actions or events the interpreter generates in response to which events.

[0244] For example, events can be represented and / or communicated within the interaction system 600 in various ways. By way of non-limiting example, an event (e.g., a payload) can include fields specifying or encoding values representing: an action type (e.g., which identifies a standardized interaction modality or corresponding action type, e.g., UserSaid), an action state (e.g., the state of an observed user action, e.g., Finished, the current or confirmed state of a bot or scenario action, e.g., Started, the indicated state of a bot or scenario action, e.g., Start), the detected or indicated action content (e.g., a transcription or indication of speech, e.g., "Hello", a description of a detected or indicated gesture, a description of a detected or indicated pose or expression, etc.), a unique identifier (UID) for identifying the event, a timestamp (e.g., indicating when the event was created, when the action was updated), a unique source identifier identifying the event source, one or more tags (e.g., specifying that the event was generated as part of a particular flow or session, or is associated with a particular user or account), context, and / or other attributes or information.

[0245] In some embodiments, each action may be identified by a unique identifier (action_uid), and all events associated with the same action may reference the same action_uid. Thus, individual events that reference the same action_uid can be used to represent the lifecycle of the corresponding action from start to end (e.g., including updated action states in between). In some embodiments, the component that emits the StartAction and ActionStarted events may generate the action_uid for a new instance of the action, and the particular component involved may depend on the type of action (e.g., bot vs. user action). For example, the Interaction Manager 640 may be responsible for generating the action_uid for new instances of bot actions initiated by the Interaction Manager 640, and the Sensing Server 620 may be responsible for generating the action_uid for new instances of observed user actions. Thus, individual events can be associated with corresponding instances of a particular type of action.

[0246] Taking an interaction classification scheme such as Figure 5 the interaction classification scheme shown as an example, actions can be classified into corresponding interaction modalities, such as speech, gesture, emotion, movement, scene, etc. Taking speech as an example, the Interaction System 600 can use the speech modality to support various events and actions related to dialogue management. For example, the user can use the UserSpeech modality (e.g., via the UserUtterance action), or the bot provided by the Interaction System 600 can use the BotSpeech modality (e.g., via the BotUtterance action). In an example user utterance action, the user can emit an utterance that is recognized by the Interaction System 600. Examples of this action include the user typing in a text interface to interact with the bot or the user speaking to an interactive avatar. Examples of possible events associated with this action include UtteranceUserActionStarted, StopUtteranceUserAction (e.g., indicating to the Action Server 670 to reduce the automatic speech recognition hold time), UtteranceUserActionTranscriptUpdated (e.g., providing an updated transcript during the UtteranceUserAction), UtteranceUserActionIntensityUpdated (e.g., providing changes in the detected speaking intensity level, typing rate, volume, or pitch, etc.), UtteranceUserActionFinished (e.g., providing the final transcript), and / or others.

[0247] In an example bot discourse action, the bot can generate a discourse (e.g., say something) to the user through some form of verbal communication (e.g., through a chat interface, a voice interface, brain-computer communication, etc.). Examples of possible events associated with this action include StartUtteranceBotAction (e.g., which indicates that the bot generates a discourse, and its payload can include the transcription of the indicated discourse of the bot, a representation of the intensity (e.g., the speaking intensity level), the output text rate, changes in volume or pitch, etc.), UtteranceBotActionStarted (e.g., which indicates that the bot has started generating a discourse), ChangeUtteranceBotAction (e.g., which indicates adjusting the volume or other properties after the action has started), UtteranceBotActionScriptUpdated (e.g., providing an updated transcription during the UtteranceBotAction), StopUtteranceBotAction (e.g., which indicates that the bot discourse stops), UtteranceBotActionFinished (e.g., confirming or reporting that the bot discourse has been completed, e.g., because it has completed or due to the user stopping the discourse), and / or others.

[0248] Taking motion as an example, the interaction system 600 can support various events and actions related to the motion modality. A motion action can represent a movement or a set of movements with a prescribed meaning. For example, the user can make a gesture or pose detected by computer vision, or a bot provided by the interaction system 600 can make a gesture or pose. In some embodiments, the user and / or the bot can use any suitable motion modality (e.g., face, upper body, lower body). In some embodiments, these modalities can be governed by an "override" modality strategy, and the action server 670 can interpret it as an instruction to handle concurrent actions by temporarily overriding the currently running action with a newly started action. As a non-limiting example, if the interaction manager 640 starts a BotPosture ("cross arms") action, which indicates that the avatar keeps its arms crossed until the action stops, and two seconds later the interaction manager 640 starts a BotGesture ("wave") action, the action server 670 can execute the wave action by overriding the "cross arms" posture with the wave action (e.g., so the avatar waves to the user). Once the wave action is completed, the action server 670 can restore the avatar to the "cross arms" posture (e.g., restoring the overridden action).

[0249] In an example facial expression bot action, a corresponding event may instruct the bot to make a facial expression (e.g., a smiley face in a chatbot's text message, a facial expression of a digital avatar in an interactive body experience) using a specified expression or emotion (e.g., happy, surprised, contemptuous, sad, fearful, disgusted, angry, etc.). Examples of possible events associated with this action include StartExpressBotAction (e.g., indicating a change in the bot's facial expression and specifying the type of expression), ExpressionBotActionStarted (e.g., indicating that the bot has started the action), StopExpressBotAction (e.g., indicating that the bot stops the facial expression), ExpressionBotActionFinished (e.g., indicating that the bot has stopped the facial expression), and / or others.

[0250] In some embodiments, the interaction system 600 may support facial expression user actions and corresponding events representing the detected user expressions. Examples of possible events associated with this action include ExpressionUserActionStarted (e.g., indicating that a user's facial expression has been detected, including a representation of the expression content, such as happy, surprised, contemptuous, sad, fearful, disgusted, angry, etc.) and ExpressionUserActionFinished (e.g., indicating that the detected user facial expression has returned to a neutral expression).

[0251] In an example gesture bot action, a corresponding event may instruct the bot to make a specified gesture. In some embodiments, the event associated with this action may include a payload that includes a natural language description of the gesture, which may include basic gestures, one or more gesture modifiers, and / or other features. Example basic gestures include talking, idling (e.g., spontaneous body movements or actions during an inactive period), affirming (e.g., a non-verbal cue or action indicating agreement, confirmation, or affirmation), negating (e.g., a non-verbal cue or action indicating disagreement, contradiction, or rejection), attracting (e.g., a specific movement, action, or behavior designed to attract a user's or audience's attention and draw it to a specific object, location, or activity), and / or others. Some example hierarchies of basic gestures include: talking mood (e.g., "talking excitedly"), idling excitability (e.g., "idling nervously"), attracting intensity (e.g., "attracting subtly"). Examples of possible events associated with this action may include StartGestureBotAction, GestureBotActionStarted, StopGestureBotAction, GestureBotActionFinished, and / or others.

[0252] In some embodiments, the interaction system 600 may support gesture user actions and corresponding events representing detected user gestures. Examples of possible events associated with such an action include GestureUserActionStarted (e.g., indicating that a user gesture has been detected, including a representation of the gesture content) and GestureUserActionFinished (e.g., indicating the completion of a detected user gesture).

[0253] In an example bot position change or bot movement action (e.g., on the BotLowerBody motion modality), the corresponding event may indicate that the bot has moved to a specified position (e.g., on the screen, in a simulated or virtual environment). The specified position may include a base position, one or more position modifiers, and / or other characteristics. In an example implementation, the supported base positions may include front and back, and the supported position modifiers may include left and right. Examples of possible events associated with such an action include StartPositionChangeBotAction (e.g., identifying the specified position to which the bot is to move) and PositionChangeBotActionFinished.

[0254] In an example user position change or user movement action (e.g., on the BotLowerBody motion modality), the corresponding event may indicate a detected change in the position of the user's lower body. Examples of possible events associated with such an action include PositionChangeUserAction (e.g., indicating that a detected user movement has started, including a representation of the detected movement direction or characteristics, such as active, approaching, passive, leaving, lateral, etc.); PositionChangeUserActionDirectionUpdated (e.g., indicating when the user changes direction during the detected movement), PositionChangeUserActionFinished (e.g., indicating that the detected movement has been completed).

[0255] In some embodiments, the interaction system 600 supports interactive visual content actions and events representing the presentation of different types of visual information and / or interaction therewith (e.g., in a 2D or 3D interface). Example interactive visual content actions (also referred to as visual actions) include visual selection actions, visual information scene actions, and visual form actions.

[0256] In an example visual selection action, the corresponding event can indicate the visualization of a selection with which the user can interact. The interaction system 600 can support different types of interactions with the visual selection (e.g., by presenting a website on a display that accepts touch or click options, accepting voice input for a selection option). For example, the StartVisualChoiseSceneAction event can include a payload with a prompt that describes the selection to be presented to the user; an image that describes what should be shown to the user, one or more supporting prompts that support or guide the user in making a selection (e.g., "Just say 'yes' or 'no' to continue"), or a recommended selection ("I can recommend a cheeseburger"); a list of options for the user to choose from (e.g., each option can have a corresponding image); the type of selection ("select", "search", etc.); and / or an indication of whether multiple selections are allowed. Other examples of possible events associated with this action include the VisualChoiceSceneActionUpdated event (e.g., indicating user interaction with a selection presented in the scene detected while the user has not yet confirmed the selection), StopVisualChoiceSceneAction (e.g., indicating the visual selection to be removed), VisualChoiceSceneActionFinished (e.g., indicating the finally confirmed selection), and / or others.

[0257] In an example visual information scene action, the corresponding event can indicate the visualization of information specified by the user. The visual information scene action can be used to show the user detailed information about a specific topic associated with the interaction. For example, if the user is interested in detailed information about a specified or displayed product or service, the visual information scene action can indicate the presentation of information about that product or service. Examples of possible events associated with such an action include StartVisualInformationSceneAction (e.g., indicating the visualization; a description of what is to be visualized; specifying one or more chunks of content to be visualized, such as a title, a content summary, and / or a description of one or more images to be visualized; one or more supporting prompts, etc.); VisualInformationSceneActionStarted (e.g., indicating that the visual information scene action has started), StopVisualInformationSceneAction (e.g., indicating the visualization has stopped), VisualInformationSceneActionFinished (e.g., indicating that the user has closed the visualization or the visual information scene action has stopped), and / or others.

[0258] In an example visual form action, a corresponding event can indicate the visualization of a specified visual form having one or more form fields (e.g., email, address, name, etc.) for a user to complete. Examples of possible events associated with this type of action include StartVisualFormSceneAction (e.g., indicating the visualization; specifying one or more inputs, prompts to the user, one or more supporting prompts, one or more images, etc.), VisualFormSceneActionStarted (e.g., indicating that the user has started entering information into the form), VisualFormSceneActionInputUpdated (e.g., indicating that the user has entered information into the form but has not confirmed the selection), StopVisualFormSceneAction (e.g., indicating the stop of the form visualization), VisualFormSceneActionFinished (e.g., indicating that the user has confirmed or canceled the form input), and / or others.

[0259] In some embodiments, the interaction system 600 can support actions and events representing aspects of a scene in which a human-machine interaction is taking place. For example, the interaction system 600 can support a sound modality (e.g., specifying sound effects or background sounds), an object interaction modality (e.g., specifying an interaction between a bot and a virtual object in the environment), a camera modality (e.g., specifying a camera cut, action, transition, etc.), a visual effects modality (e.g., specifying visual effects), a user presence modality (e.g., indicating whether the presence of a user has been detected), and / or other example actions. These examples and others are described in more detail in U.S. Provisional Application No. 63 / 604,721, filed Nov. 30, 2023, the contents of which are incorporated herein by reference in their entirety.

[0260] Having described some example events associated with standardized types of actions and interaction modalities, and some possible ways of representing such events and actions, the following discussion turns to some possible ways in which the interaction manager 640 (e.g., an interpreter) can use a prescribed interaction flow (or simply a flow), e.g., written in an interaction modeling language, to evaluate such events (e.g., incoming and / or queued instances of standardized input events 630, internal events 660, events 665 representing confirmations of changes in action state), determine what action or event to generate in response, and generate the corresponding events (e.g., outgoing instances of indicated bot action events 650, internal events 660).

[0261] Typically, a flow can use primitives from an interaction modeling language to specify instructions, which includes semantically meaningful (e.g., natural language) keywords and commands that use an interaction classification scheme to specify events (e.g., something happened) and actions (e.g., something needs to happen). The state of an action (e.g., the state of an observed user action, the current state of a bot or scene action) and / or commands to change the state of a bot or scene action can be represented using standardized event keywords, commands, and / or syntax. For example, an action event (e.g., a user or bot action starts or stops) can be represented using an event specifier with standardized syntax (e.g., an event name and / or identifier (which includes keywords identifying standardized action categories), and a specifier of the user or bot action state). Instruction lines in a flow can include: an event trigger (e.g., using a keyword such as send) that causes the interpreter to generate a specified event when certain specified conditions are met (e.g., an event representing a command to execute a bot action can trigger the action to be performed, and an event representing a change in the user state can trigger a corresponding bot action); or an event matcher (e.g., using a keyword such as match) that causes the interpreter to interrupt the flow and monitor for a specified event before resuming the flow. The event trigger and event matcher can use an event specifier to specify the corresponding triggering and matching conditions, which includes a standardized event name or identifier (e.g., a keyword identifying a standardized action category paired with a corresponding action state specifier or a command to change the action state) and arguments specifying one or more conditions that the specified event must meet (e.g., using predefined parameters and supported values, or a natural language description).

[0262] Accordingly, the interaction manager 640 (e.g., its interpreter) can be equipped with logic to interpret corresponding keywords, commands, and / or syntax such as these. In some embodiments, the interaction manager 640 can support any number of keywords for parallelizing action and flow execution and matching (e.g., any of the above keywords such as send, match, start, stop, wait, activate, return, abort, and / or others). Accordingly, the interaction manager 640 can be programmed to sequentially execute the instructions specified in a prescribed flow, generate any events specified by event triggers, and stop when the flow head reaches an event matcher, an exception, or the end of the flow. In some embodiments, the interaction manager 640 can support and track multiple active flows (e.g., active flows interrupted at corresponding event matchers), listen (e.g., using an event-driven state machine) for incoming events that match the event matchers of the active flows, and trigger the corresponding events and actions specified in the matching flows.

[0263] Figure 7An example interaction manager 700 in accordance with some embodiments of the present disclosure is shown. In this example, the interaction manager 700 includes an interpreter 710, one or more interaction flows 780, and an internal event queue 790. At a high level, the interaction flows 780 specify corresponding sequences of instructions in an interaction modeling language and can be loaded or otherwise made accessible to the interpreter 710, and the interpreter 710 can include an event handling component 730 that sequentially executes the instructions specified in the interaction flows 780 to process incoming events and generate outgoing events (e.g., in a standardized form).

[0264] The event handling component 730 can execute a main processing loop that processes incoming events and generates outgoing events. At a high level, the event handling component 730 includes a flow execution component 750 and a flow matcher 740. The flow execution component 750 can sequentially execute the instructions specified in the flows (e.g., parent flows, matching flows) in the interaction flow 780, generate any events specified by event triggers, and stop when the flow head reaches an event matcher, an exception, or the end of the flow. The flow matcher 740 can evaluate incoming events to determine whether they match the event matchers of the active flows, direct an action conflict resolver 760 to resolve any conflicts between multiple matching flows, and direct the flow execution component 750 to advance (e.g., non-conflicting) matching flows.

[0265] In an example embodiment, the flow execution component 750 can perform lexical analysis on the instructions specified in the interaction flow 780 (e.g., tokenizing; identifying keywords, identifiers, arguments, and other elements), iterate over the flow instructions, execute each instruction in order, and include a mechanism for handling exceptions. In some embodiments, the flow execution component 750 uses a different flow head for each (e.g., active) interaction flow 780 to indicate the current position and advance through the instructions in the corresponding interaction flow. Depending on the instruction, the flow execution component 750 can advance any given flow head to the next instruction, jump to another flow referenced by a specified label or other flow identifier, fork into multiple heads, merge multiple flow heads together, and / or otherwise. Thus, the flow execution component 750 can coordinate with a flow tracking and control component 770 to build and maintain a hierarchy of flow heads. If a parent flow head in a branch of the hierarchy of flows or flow heads is stopped, paused, or resumed, the flow execution component 750 can coordinate with the flow tracking and control component 770 to stop, pause, or resume all child flow heads of that parent flow head or branch, respectively. In some embodiments, any flow can specify any number of scopes, and the flow execution component 750 can use these scopes to generate stop events that direct the corresponding action server to stop previously started actions within the corresponding scope.

[0266] For example (e.g., at startup), the flow execution component 750 may execute a top-level flow (e.g., the top-level flow of the interaction flow 780), which specifies instructions to activate any number of flows (e.g., interaction flows 780) including any number of event matchers. The flow tracking and control component 770 may use any suitable data structure to track the active flows and the corresponding event matchers (e.g., using a tree or other representation of nested flow relationships), and may employ an event-driven state machine to listen for various events and trigger the corresponding actions specified in the matching flows (using the event matchers that match the incoming events). Thus, the flow execution component 750 may iterate through the active flows, generate any events specified by the event triggers, and stop when the flow head reaches an event matcher, an exception, or the end of the flow.

[0267] In some embodiments, an advancing flow may instruct the flow execution component 750 to generate outgoing events indicating certain actions. Additionally or alternatively, an advancing flow may instruct the flow execution component 750 to generate events that notify a listener (e.g., the flow execution component 750 itself) that an event has occurred. Thus, the flow execution component 750 may send out these events, and / or the interpreter 710 may maintain an internal event queue 790 and place these events in the internal event queue 790 (e.g., in case another flow is listening for the generated events).

[0268] Once the flow heads of all the advanced flows reach an event matcher, an exception, or the end of the flow, the flow matcher 740 may sequentially process the incoming events (e.g., from the internal event queue 790, from some other queue or event gateway, such as Figure 1 the event gateway 180), and for each event, test whether the event matchers specified by each active flow match the event. In some embodiments, the flow matcher 740 sequentially processes any internal events in the internal event queue 790 (e.g., testing whether the active flows match the internal events) before advancing to process the next incoming event (e.g., from the event gateway). The internal events may represent the updated state of the interaction flow 780 that has been advanced in response to a particular incoming event (e.g., indicating that a particular flow has started, completed, aborted, etc.). Thus, a designer may create flows that depend on the evolution or state of other flows.

[0269] When processing an event, the flow matcher 740 can compare the event with the event matchers of each active (e.g., interrupt) flow to determine whether the event matches any active flow (e.g., using any known matching techniques and / or as described in more detail below). In some cases, multiple active flows specifying various interactions can be triggered by different conditions that may be satisfied by the same event. If there is an event matcher from an active flow (matching flow) that matches the event, the flow matcher 740 can instruct the flow execution component 750 to advance that flow (e.g., and generate an outgoing event to trigger any actions specified by the advancing flow).

[0270] If there are multiple matching flows, the flow matcher 740 can instruct the action conflict resolver 760 to determine whether the matching flows agree on an action. If they agree, the action conflict resolver 760 (or the flow matcher 740) can instruct the flow execution component 750 to advance both matching flows. If they do not agree, the action conflict resolver 760 can apply conflict resolution to identify which action should take precedence, instruct the flow execution component 750 to advance the matching flow with the precedence action, and abort the other matching flows (e.g., because the interaction patterns represented by those flows will no longer apply). If there are no active flows that match the event, the flow matcher can generate an internal event that matches a prescribed flow to handle the unmatched or unhandled event, can run one or more unhandled event handlers (e.g., unhandled event handler 744), and / or can use some other technique to handle the unhandled event.

[0271] After checking for matches and advancing the flows, the flow tracking and control component 770 can check the flow states of any completed or aborted flows and can stop any active flows activated by those completed or aborted flows (e.g., because the interaction patterns represented by those flows will no longer apply). Thus, the interpreter 710 can iterate through the events, advance the flows, perform conflict management to determine which actions to execute, and generate outgoing events to trigger those actions.

[0272] For example, in some embodiments, the interpreter 710 uses an event-driven state machine (e.g., Figure 8 the event-driven state machine 800 therein) to process incoming action events 805 and internal events 820. In some embodiments, the event-driven state machine 800 can place the incoming action events 805 (e.g., which can correspond to Figure 3 the standardized input event 630, and can be routed via an event gateway (e.g., Figure 1 the event gateway 180 therein)) in the interaction event queue 810. The event-driven state machine 800 can place the internal events 820 (e.g., which can correspond to Figure 6Internal events are placed in an internal event queue 815 (e.g., which may correspond to Figure 7 the internal event queue 790), and events from the internal event queue 815 may be processed preferentially over events from the interaction event queue 810.

[0273] For each event, the event-driven state machine 800 may execute at least some of the steps shown in block 825. For example, at block 830, the event-driven state machine 800 may test whether the event matcher specified by each active flow matches the event. If there is an event matcher in the active flow that matches the event (matching flow), the event-driven state machine 800 may proceed to block 835 and advance that flow (e.g., generate an outgoing interaction event 870 to trigger an action). If there are multiple matching flows, the event-driven state machine 800 may proceed to block 840 and determine whether the matching flows agree on an action. If they agree, the event-driven state machine 800 may proceed to block 850 and advance both matching flows. If they do not agree, the event-driven state machine 800 may proceed to block 855 and may apply conflict resolution to identify which action should be prioritized, advance the matching flow with the prioritized action, and abort the other matching flows. If there is no active flow that matches the event, the event-driven state machine 800 may proceed to block 835 and run one or more unhandled event handlers (or generate internal events that match a specified flow to handle unmatched or unhandled events). After checking for matches and advancing the flows, the event-driven state machine 800 may proceed to block 860, may check the flow state of any completed or aborted flows, may stop any active flows activated by those completed or aborted flows, and may proceed to the next event at block 865. Thus, the event-driven state machine 800 may iterate through the internal events 820 in the internal event queue 815 and / or the incoming action events 805 in the interaction event queue 810, advance the flows, perform conflict management to determine which interactions to execute, and generate outgoing interaction events 870 to trigger those interactions.

[0274] Return Figure 7 , in some embodiments, the interpreter 710 may support the use of natural language descriptions and the use of one or more LLMs, such as Figure 28A the example generative LLM system 2800 or Figure 28A 、 Figure 28B or Figure 28C the generative LLM 2830.

[0275] For example, each interaction flow 780 can be specified with a corresponding natural language description that summarizes the interaction pattern represented by the flow, and the interpreter 710 uses such a flow description in some cases (e.g., the prescribed flow for handling unknown events and / or the unknown event handler 744 can prompt the LLM to determine whether a mismatch event representing an unrecognized user intent semantically matches the natural language description of the activity flow representing the target user intent). Thus, in some embodiments, the interpreter 710 can include a flow description generator 720 that parses one or more specified interaction flows 780 (e.g., at design time), performs lexical analysis to identify whether any specified flow lacks a corresponding flow description, and if so, prompts the LLM to generate a flow description (e.g., based on the name and / or instructions of the flow). Additionally or alternatively, the flow description generator 720 can (e.g., prompt the LLM) determine whether any specified flow description is inconsistent with its corresponding flow description, and if so, prompt the LLM to generate a new flow description (e.g., as a suggestion or for automatic replacement) (e.g., from the name and / or instructions of the flow). Thus, the flow description generator 720 can determine whether to generate a description for any interaction flow 780 and can generate the corresponding flow description.

[0276] In some embodiments, a designer can specify a flow description for the interaction flow 780 without an instruction sequence (e.g., a natural language description of what the flow should do), or can call one of the interaction flows 780 by name without defining it. Thus, in some embodiments, the interpreter 710 can include a flow autocomplete component 725 that parses the interaction flow 780 (e.g., at design time, at runtime), identifies whether the interaction flow 780 lacks an instruction sequence, and if so, prompts the LLM to generate an instruction sequence (e.g., based on the name and / or description of the flow). For example, the flow autocomplete component 725 can provide one or more prompts to the LLM that include one or more example flows, the specified name of the interaction flow 780, and / or the (e.g., specified or generated) natural description of the interaction flow 780, as well as a prompt to complete the interaction flow 780.

[0277] For example, the flow autocomplete component 725 can use a template prompt with placeholders to construct the prompt, such as the following:

[0278] Content: |-

[0279] # Example flows:

[0280] {{example}}

[0281] # Complete the following flow based on its instructions:

[0282] Flow {{flow_name}}

[0283] """{{Natural language description of the flow}}"""

[0284] This example template prompt includes placeholders for an example flow, a specified name for the flow, and a specified natural language description of the flow. The flow autocomplete component 725 can generate one or more prompts, populate the placeholders with corresponding content (e.g., a predefined example flow, the specified name for the flow, the specified natural language description of the flow, and / or other content), and can provide this constructed prompt to an LLM (e.g., via an API request). Thus, the LLM can generate and return an autocomplete flow with the generated instructions, and the flow autocomplete component 725 can insert it into the corresponding interaction flow 780 or otherwise associate it with the corresponding interaction flow 780.

[0285] In an example implementation, the flow execution component 750 can execute the instructions specified in the interaction flow 780 (e.g., including any encountered event triggers) until it reaches an event matcher, at which point the flow execution component 750 can interrupt the interaction flow 780. The flow matcher 740 can process each event by executing the event matcher in each interrupted flow, comparing the event with the target event parameters and parameter values specified by the event specifier of the event matcher. Depending on the implementation, the flow matcher 740 can support various matching techniques to determine whether the event matches any active event matcher of an active flow. Generally, the flow matcher 740 can use any known technique to compare the target event parameters and parameter values with the parameters and parameter values of the event to generate some representation of whether the event matches (e.g., a binary indication quantifying an exact or fuzzy match or a match score).

[0286] However, in some implementations, an event trigger or event matcher in one of the interaction flows 780 can use a natural language description to specify target event parameters and / or parameter values. Thus, in some embodiments, the syntax generator 752 can infer the target event parameters and / or values from the specified natural language description in the interaction flow 780 (e.g., descriptions of all target event parameters and values, descriptions of individual parameter values), and the syntax generator 752 can insert the generated target event parameters and values into the corresponding event specifier in the interaction flow 780 (or otherwise associate it with the corresponding event specifier). For example, before the flow execution component 750 executes an instruction that includes an event specifier (e.g., an event trigger), the flow execution component 750 can (e.g., at runtime) instruct the syntax generator 752 to determine whether the instruction includes parameters specified using a natural language description (e.g., using lexical analysis). Additionally or alternatively, before the flow matcher 740 executes an instruction that includes an event specifier (e.g., an event matcher), the flow matcher 740 can (e.g., at runtime) instruct the syntax generator 752 to determine whether the instruction includes parameters specified using a natural language description (e.g., using lexical analysis). If so, the syntax generator 752 can prompt the LLM to generate the corresponding target event parameters and / or parameter values for the event specifier and update the event specifier in the corresponding one of the interaction flows in the interaction flow 780 with the generated target event parameters and / or parameter values.

[0287] Taking the example of a sample prompt for generating a target event parameter value (or any other variable value), the syntax generator 752 can use a templated prompt with placeholders to construct a prompt such as the following:

[0288] Content: |-

[0289] """

[0290] {{general_instructions}}

[0291] """

[0292] # The conversation between the user and the bot can proceed as follows:

[0293] {{sample_conversation}}

[0294] # This is the current conversation between the user and the bot:

[0295] {{history|colang}}

[0296] # {{Natural language description of the parameter value}}

[0297] ${{var_name}} =

[0298] This example template prompt includes placeholders for general instructions, example conversations (or a series of interactions), the history of the current conversation (or a series of interactions), the name of the variable whose value is being generated, and a prompt for generating the value ("${{var_name}}="). The syntax generator 752 can generate one or more prompts, fill the placeholders with corresponding content (e.g., specified instructions, specified example conversations or interaction histories, the recorded history of the current conversation or a series of interactions, the extracted natural language description of the parameter values to be generated for the corresponding variable, the name of the variable, and / or other content), and can provide the constructed prompt to the LLM (e.g., via an API request). Thus, the LLM can generate and return a prompt value, and the syntax generator 752 can insert the prompt value into the event specifier in the corresponding instruction. This example only represents one possible way in which the LLM can be used to generate target event parameter values from a specified natural language description of the values. Other types of prompts and prompt content can be implemented within the scope of the present disclosure. Additionally, those of ordinary skill in the art will understand how to adapt the above example prompts to generate other types of content described herein (e.g., generating the name of a target event parameter from a natural language description of the target event parameter and / or parameter values, a list of supported variable names, etc.). Thus, the flow execution component 750 can use the target event parameters and parameter values generated by the LLM to execute event triggers, and / or the flow matcher 740 can use the target event parameters and parameter values generated by the LLM to execute event matching.

[0299] Thus, in some embodiments, the flow matcher 740 generates and / or quantifies a representation of whether an event matches (e.g., explicitly or implicitly) by comparing the specified or generated target event parameters / parameter values of the event matcher (e.g., which represent keywords or commands for the target interaction modality, action, action state, and / or other event parameter values) with the corresponding parameters / parameter values of the event being tested (e.g., which represent keywords or commands for the indicated or detected interaction modality, action, action state, and / or other event parameter values). Additionally or alternatively, the flow matcher 740 can include a flow description matcher 742 that (e.g., at runtime) prompts the LLM to determine whether the event matches the flow description of one of the interaction flows 780 and / or the specified natural language description of one or more parameters or parameter values to be matched.

[0300] At a high level, events can represent user actions or intents, bot actions or intents, scenario interactions, or some other type of event using a standardized interaction classification scheme that classifies actions, action events, event parameters, and / or parameter values using (e.g., standardized, natural language, semantically meaningful) keywords and / or commands and / or natural language descriptions (e.g., GestureUserActionFinished(“thumbs up”))). Accordingly, the flow description matcher 742 of the flow matcher 740 can perform event matching by prompting the LLM to determine whether the keywords, commands, and / or natural language descriptions of an incoming or internal event match a (e.g., specified or generated) flow description of one of the interaction flows 780. For example, the flow description matcher 742 can use a template prompt to construct a prompt that includes a prompt for determining whether an event matches a flow description, fills placeholders with corresponding content (e.g., prescribed instructions, prescribed sample dialogues or interaction histories, the recorded current dialogue or history of a series of interactions, the specified or generated flow description of the interaction flow 780, the keywords and / or commands and / or other content represented by the incoming or internal event), and can provide the constructed prompt to the LLM (e.g., via an API request). In this way, the LLM can return an indication of whether the event matches the flow description of the interaction flow 780. In many cases, the LLM can provide a more nuanced match or semantically understanding match compared to conventional exact or fuzzy matching algorithms.

[0301] Additionally or alternatively, the flow matcher 740 can include a flow instruction matcher 746 that prompts the LLM to determine whether an incoming or internal event matches the instructions of the active flow in the interaction flow 780. For example, in response to the flow matcher 740 applying one or more matching techniques (e.g., using exact matching, fuzzy matching, flow description matching, and / or others) and determining that there is no active flow that matches the incoming or internal event, the flow matcher 740 can trigger a prescribed flow (e.g., which is used to handle unknown events) or execute an unhandled event handler 744 that includes the flow instruction matcher 746. In an example implementation, the unhandled event handler 744 includes the flow instruction matcher 746 and a bot interaction flow generator 748, but this is only an example. Generally, any number of matching techniques can be applied in any order, whether as an initial test, as part of the unhandled event handler 744, or otherwise.

[0302] In an example embodiment, the flow instruction matcher 746 may prompt the LLM to determine whether a representation of an incoming or internal event and / or a recent interaction history matches specified content of an active flow in the interaction flow 780. The flow instruction matcher 746 may achieve this by inferring the user's intent (e.g., matching an incoming or internal event with an instruction of a flow that listens for the corresponding user intent). In an example embodiment, the flow instruction matcher 746 may perform an event matcher by prompting the LLM to determine whether keywords, commands, and / or natural language descriptions of an incoming or internal event match an (e.g., specified or generated) instruction of one of the interaction flows 780.

[0303] For example, the flow instruction matcher 746 may use a templated prompt with placeholders to construct a prompt, such as the following:

[0304] Content: |-

[0305] """

[0306] {{general_instructions}}

[0307] """

[0308] # The conversation between the user and the bot can proceed as follows:

[0309] {{sample_conversation}}

[0310] # These are the most likely user intents:

[0311] {{sample_flows_and / or_intents}}

[0312] # This is the current conversation between the user and the bot:

[0313] {{history}}

[0314] User action: {{incoming_or_internal_event}}

[0315] User intent:

[0316] This example template prompt includes placeholders for general instructions, an example conversation (or series of interactions), some possible flows (or corresponding list of possible user intents) representing the target user intent to match, the history of the current conversation (or series of interactions), keywords and / or commands represented by incoming or internal events, and a hint for predicting the matching flow or user intent ("User intent:"). The flow instruction matcher 746 can generate one or more hints, populate the placeholders with corresponding content (e.g., prescribed instructions, a prescribed sample conversation or interaction history, the name and / or instructions of a prescribed interaction flow 780 representing a possible user intent, the corresponding list of possible user intents, the recorded history of the current conversation or series of interactions, keywords and / or commands represented by incoming or internal events, and / or other content), and can provide the constructed hint to the LLM (e.g., via an API request). Thus, the LLM can return an indication of whether the event matches one of the interaction flows 780 and / or the corresponding user intent, and if so, an indication of which one.

[0317] In some cases, there may be no matching flow defined for the bot's response to a particular user interaction, or the flow matcher 740 may not identify a matching flow. Thus (e.g., in some embodiments where the flow matcher 740 determines that there is no active flow that matches an incoming or internal event representing a user interaction), the bot interaction flow generator 748 can prompt the LLM to generate a flow (e.g., at runtime). For example, in some embodiments, the flow matcher 740 (e.g., the flow instruction matcher 746) can first attempt to use the LLM to match an unknown incoming or internal event with the name, instructions, and / or other representations of one or more prescribed flows that listen for the corresponding target user intent (and define the bot response), and if the LLM determines that there is no matching flow or target user intent, the bot interaction flow generator 748 can prompt the (same or some other) LLM to predict the user intent represented by the unknown incoming or internal event, generate a response proxy intent, and / or generate a response flow. For example, if the unknown event represents a user action, the bot interaction flow generator 748 can apply any number of hints to instruct the LLM to classify the unknown user action as a user intent, generate a response proxy intent, and / or generate a flow that implements the response proxy intent.

[0318] For example, the bot interaction flow generator 748 can use a template prompt with placeholders to construct a first hint, such as the following:

[0319] Content: |-

[0320] """

[0321] {{general_instructions}}

[0322] """

[0323] # The conversation between the user and the bot can be carried out as follows:

[0324] {{sample_conversation}}

[0325] # This is the current conversation between the user and the bot:

[0326] {{history}}

[0327] bot intention:

[0328] And construct a second prompt as follows:

[0329] bot action:

[0330] These example template prompts include placeholders for general instructions, sample conversations (or a series of interactions), the history of the current conversation (or a series of interactions) (including keywords and / or commands represented by incoming or internal events), a prompt for predicting the response agent's intention ("bot intention:"), and a prompt for predicting the response agent's flow ("bot action:"). The bot interaction flow generator 748 can generate one or more prompts, fill the placeholders with corresponding content (e.g., prescribed instructions, prescribed sample conversations or interaction histories, the recorded history of the current conversation or a series of interactions, keywords and / or commands represented by incoming or internal events, and / or other content), and can provide the constructed prompt to the LLM (e.g., via an API request). Thus, the LLM can generate and return a response agent flow, which the bot interaction flow generator 748 can specify as a match and provide to the flow execution component 750 for execution.

[0331] The following examples can be used to illustrate some possible prompt content. For example, the interpreter 715 can use general instructions to construct a prompt (e.g., by filling a template prompt with a placeholder for the prescribed general instructions), such as:

[0332] The following is a conversation between Emma (a helpful AI assistant (bat)) and the user.

[0333] The bot is designed to generate human-like actions based on the user actions it receives.

[0334] The bot is very talkative and provides a lot of specific details.

[0335] The bot likes to chat with the user, and the topics include but are not limited to sports, music, hobbies, NVIDIA, technology, food, weather, and animal themes.

[0336] When the user remains silent, the bot encourages the user to try different presentations, prompting the user to choose one from different options by voice or by clicking on the presented options.

[0337] When the user asks a question, the bot gives an appropriate answer.

[0338] When the user gives an instruction, the bot acts according to the instruction.

[0339] bot Appearance:

[0340] Emma is wearing a dark green dress and a white shirt. There is a small card on the shirt with the logo of the AI company NVIDIA printed on it. Emma is wearing glasses and white earrings and has medium-length brown hair. Her eyes are greenish-brown.

[0341] These are the available presentations:

[0342] A) Simple number guessing game

[0343] B) Multimodal presentation case demonstrating how the interaction modeling language handles multiple parallel actions

[0344] C) Demonstrating how the bot uses the backchannel mechanism to communicate with the user

[0345] D) Presentation case showing different bot postures depending on the current interaction state

[0346] E) Demonstrating how the bot proactively responds to unanswered user questions by repeating the question

[0347] Important note:

[0348] The bot uses "bot gestures" actions as much as possible

[0349] If the user remains silent, the bot must not repeat

[0350] User actions:

[0351] The user says "text"

[0352] bot actions:

[0353] The bot says "text"

[0354] The bot notifies "text"

[0355] The bot asks "text"

[0356] The bot expresses "text"

[0357] The bot responds "text"

[0358] The bot clarifies "text"

[0359] The bot suggests "text"

[0360] The bot gesture "gesture"

[0361] In some embodiments, the interpreter 715 can use sample conversations or a series of interactions to construct prompts (e.g., by filling a template prompt with placeholders for a prescribed sample conversation or series of interactions), such as:

[0362] # The conversation between the user and the bot can proceed as follows:

[0363] User action: The user says "Hello!"

[0364] User intent: The user expresses greetings

[0365] Bot intent: The bot expresses greetings

[0366] Bot action: The bot says "Hello! What can I do for you today?" and the bot gesture "smile"

[0367] User action: The user says "What can you do for me?"

[0368] User intent: The user asks about functions

[0369] Bot intent: The bot responds about functions

[0370] Bot action: The bot says "As an AI assistant, I can help you with various tasks." and the bot gesture "open hands in an inviting gesture"

[0371] User action: The user says "ddsf poenwrtbj vhjhd sfd dfs"

[0372] User intent: The user says something unclear

[0373] Bot intent: The bot informs the user of unclear input

[0374] Bot action: The bot says "As an AI assistant, Bot action: The bot informs "Sorry! I didn't hear clearly! Can you say that again?" and the bot gesture "make a sad expression"

[0375] User action: The user says "Tell me something about the history of NVIDIA."

[0376] User intent: The user asks about the history of NVIDIA

[0377] Bot intent: The bot responds to provide information about the history of NVIDIA

[0378] Bot action: Bot responds with "NVIDIA is a technology company that specializes in designing and manufacturing graphics processing units (GPUs) and other computer hardware. The company was founded in 1993 by Jen-Hsun Huang, Chris Malachowsky, and Curtis Priem." and Bot gesture "Continue to explain gesture"

[0379] User action: User says "So, do you also produce keyboards?"

[0380] User intent: User asks about keyboards

[0381] Bot intent: Bot clarifies about keyboards

[0382] Bot action: Bot clarifies with "No, we actually only manufacture CPUs and do not make keyboards." and Bot gesture "Move index finger in a shaking manner"

[0383] User action: User says "Tell me more about NVIDIA"

[0384] User intent: User requests more information about NVIDIA

[0385] Bot intent: Bot provides more information about NVIDIA

[0386] Bot action: Bot informs "Initially, the company focused on developing 3D graphics processing technology for the PC gaming market. In 1999, NVIDIA released the world's first GPU, GeForce 256, which was a major breakthrough in the gaming industry. The company has continued to innovate in the GPU field, releasing new products and expanding into other markets such as professional graphics, mobile devices, and artificial intelligence." and Bot gesture "Continue to explain gesture"

[0387] User action: User says "Thank you"

[0388] User intent: User expresses gratitude

[0389] Bot intent: Bot expresses gratitude and offers more help

[0390] Bot action: Bot expresses "You're welcome. If you have any other questions or if I can assist you further, please don't hesitate to ask." Bot gesture "Place both hands together flat and bow slightly"

[0391] User action: User says "Turn around and count to 3"

[0392] User intent: User instructs to turn around and count to 10

[0393] Bot intent: Bot turns around and counts to 10

[0394] Bot action: The bot gesture "turn around" and the bot says "one, two, three"

[0395] In some embodiments, the interpreter 715 can construct a prompt (e.g., by using the recorded current conversation or the history of a series of interactions, filling a template prompt with placeholders for the recorded current conversation or the history of a series of interactions), such as:

[0396] # This is the current conversation between the user and the bot:

[0397] User intention: The user remains silent for 8.0

[0398] The prompt content described herein is only an example, and variations can be implemented within the scope of the present disclosure.

[0399] Back to Figure 6 , the sensing server 620 can convert the detected input event 610 (e.g., which represents some detected user input, such as a detected gesture, voice command, or touch or click input; which represents some detected feature or event associated with the user input, such as the presence or absence of detected voice activity, the presence or absence of detected typing, detected transcribed speech, detected changes in typing volume or speed; etc.) into a standardized input event 630. In some embodiments, different sensing servers can handle the detected input events 610 for different interaction modalities (e.g., one sensing server for converting detected gestures, one sensing server for converting detected voice commands, one sensing server for converting detected touch inputs, etc.). Thus, any given sensing server 620 can operate as an event responder, acting as a mediator between the corresponding input source and one or more downstream components (e.g., Figure 1 the event gateway 180 in), for example, by converting the input event into a standardized format.

[0400] Taking the sensing server for GUI input events as an example, the sensing server can effectively convert GUI input events (e.g., "the user clicks the button 'chai-latte', scrolls down and clicks the button 'confirm'") into standardized interaction-level events (e.g., "the user selects the option 'Chai Latte'"). A possible example of a standardized interaction-level event is a confirmation status update event (e.g., which indicates the detected status or change in status of the presented confirmation status, such as confirmed, cancelled, or unknown). For example, the sensing server can convert different types of GUI inputs into corresponding confirmation status update events, and the conversion logic can vary depending on the type of interaction element being presented or interacted with. For example, a button press can be converted into a "confirmed" status update event, or if a visual form presents a single form field input, the sensing server can convert an "Enter" keyboard event into a "confirmed" status update event. Another possible standardized interaction-level event is a selection update event (e.g., which indicates a detected change in the current option selected by the user). For example, if the user selects the item "chai-latte" from a list of multi-select elements, the sensing server can convert the corresponding detected GUI input event (e.g., a click or tap on a button or icon) into a standardized selection update event, which indicates the detected change in the current option selected by the user. Another example of a possible standardized interaction-level event is a form input update event, which indicates an update to the requested form input. These are just a few examples, and other examples are also within the scope of this disclosure. Other examples of standardized interaction-level GUI input events (e.g., which represent detected GUI gestures, such as swiping, pinch zooming, or rotating for a touchscreen device), standardized interaction-level video input events (e.g., which represent detected visual gestures, such as face recognition, pose recognition, object detection, presence detection, or motion tracking events), standardized interaction-level audio input events (e.g., which represent detected speech, detected voice commands, detected keywords, other audio events, etc.) and / or others are contemplated within the scope of this disclosure.

[0401] Figure 9 FIG. 930 shows an example action server 930 according to some embodiments of the present disclosure. The action server 930 can correspond to Figure 1 action server 170 and / or Figure 6 action server 670. At a high level, the action server 930 can be subscribed to or otherwise configured to pick up and execute those events that the action server 930 is responsible for executing from the event bus 910 (e.g., which can correspond to Figure 1 event gateway 180). In Figure 9In the illustrated embodiment, the action server 900 includes one or more event workers 960 (which forward incoming events to the corresponding modality services), and an event interface manager 940 that manages the event workers 960.

[0402] For example, the event interface manager 940 can be subscribed to a global event channel of the event bus 910 that carries (e.g., standardized) events that indicate when an interaction channel between the connection interaction manager and the end user device has been acquired (e.g., PipelineAcquired) or released (e.g., PipelineReleased). Thus, the event interface manager 940 can create a new event worker (e.g., event worker 960) in response to an event indicating that a new interaction channel has been acquired, and / or can delete an event worker in response to an event indicating that the corresponding interaction channel has been released. In some embodiments, the event interface manager 940 performs periodic health checks (e.g., using any known technique, such as inter - process communication) to ensure that the event workers 960 are healthy and running. If the event interface manager 940 discovers that one of the event workers 960 is not responsive, the event interface manager 940 can restart that event worker.

[0403] The event workers 960 can subscribe to one or more per - flow event channels of the event bus 910 (e.g., per - flow event channels dedicated to a specific interaction modality that the action server 930 is responsible for), and can forward incoming events to different modality services registered for the corresponding events. In some embodiments, the event workers can run in a separate (e.g., multi - processing) process (e.g., process 950), and can manage incoming and outgoing events (e.g., using an asycnio event loop).

[0404] Modality services (e.g., Figure 9The modal services A and B) in can implement action-specific logic for each standardized action category and / or action event supported by an interaction modeling language and / or defined by an interaction classification scheme for a given interaction modality. Thus, a given modal service can be used to map actions of the corresponding interaction modality to a specific implementation within the interaction system. In an example implementation, all supported actions in a single interaction modality are handled by a single modal service. In some embodiments, a modal service can support multiple interaction modalities, but different actions of the same interaction modality are not handled by different modal services. For example, in some embodiments involving an interactive visual content modality, different actions in that interaction modality (e.g., VisualFormSceneAction, VisualChoiceSceneAction, VisualInformationSceneAction) are handled by the same GUI modal service.

[0405] Figure 10 shows an example event flow through an example action server 1000 according to some embodiments of the present disclosure. In this example, flow XY (e.g., which can correspond to Figure 9 one of the per-flow event channels of the event bus 910) can carry various events (e.g., events 1-7), and an event worker 1010 (e.g., which can correspond to Figure 9 the event worker 960) can be subscribed to flow XY and provide a corresponding event view (e.g., event view A, event view) to the subscribed modal service (e.g., modal service A), which is populated with a subset of the events in flow XY that the modality is subscribed to receive. Thus, the modal service can execute the indicated actions represented by the events it subscribes to, apply the corresponding modal policy to manage the corresponding action stack, and call the corresponding action handler to execute the action. Thus, the action handler can execute the action, generate internal events (e.g., which indicate a timeout) and place them in the corresponding event view (so that the modal service can take appropriate actions and maintain the action stack and lifecycle), and / or generate (e.g., standardized) interaction modality (IM) events (e.g., which indicate that certain actions have started, been completed, or been updated) and place them in flow XY.

[0406] In some embodiments, each modality service can register itself with an event worker (e.g., event worker 1010) with a list (e.g., type) of events of interest (e.g., events handled by that modality service). Thus, the event worker 1010 can provide the service with an event view (e.g., event view A) that is a subset of all events in the stream. The modality service can process the events within the corresponding event view sequentially. In some embodiments where the action server 1000 includes multiple modality services, different modality services can process events in parallel (e.g., using an asynchronous event loop).

[0407] In some embodiments, each modality service implements a prescribed modality policy (e.g., the modality policy shown in Figure 4 ). In an example of a modality policy that allows concurrent actions in a given modality, the corresponding modality service can trigger, track, and / or otherwise manage concurrent actions and can perform any number of actions simultaneously. The modality service can assign a common action identifier (e.g., action_uid) that uniquely identifies a particular instance of an action and can track the lifecycle of that action instance in response to action events generated by the corresponding action handler and referencing the same action identifier.

[0408] In an example of a modality policy that overrides overlapping actions in a given modality, the corresponding modality service can manage an action stack, and the modality service can pause or hide the currently executing action in response to a subsequently indicated action. Once the action is complete and the corresponding (e.g., internal) action event representing that event is again relevant to the modality service, the modality service can trigger the topmost remaining action in the stack to resume or become unhidden. For example, an animation modality service can initially start a StartGestureBotAction(gesture = "talking") event, and if it subsequently receives a StartGestureBotAction(gesture = "pointing down") event before the talking gesture (animation) ends, the modality service can pause the talking gesture, trigger the pointing down gesture, and resume the talking gesture when the pointing down gesture ends.

[0409] In some embodiments, the modality service can synchronize action state changes with prescribed conditions (e.g., wait until a previous action in the same modality is complete before starting an action, align the completion of two different actions in different modalities, align the start of one action with the end of some other action, etc.). By way of illustration, Figure 11 an example action lifecycle 1100 is shown. For example, a modality service can receive a StartAction event 1110, which indicates that an action on the corresponding interaction modality handled by that modality service should start. In Figure 1In the illustrated embodiment, at decision block 1120, the modality service may determine whether a modality is available. For example, the modality service may enforce a modality policy that waits for any ongoing actions on the modality to complete before starting a new action on the modality. The modality service may track the lifecycle of the initiated actions on the modality and may thus determine that there are some other pending actions that have started but not yet completed. Thus, the modality service may wait until it receives an event indicating that the action has completed (e.g., Figure 11 the modality event shown in Figure 11 ), at which point the modality service may proceed to decision block 1130. At decision block 1130, the modality service may determine whether a specified start condition is met (e.g., an instruction to synchronize starting a new action with starting or completing some other action). Thus, the modality service may wait for the designed start condition to occur (e.g., indicated by a synchronization event from

[0410] ) and the interaction modality remains idle before initiating the action, and at block 1140 may generate an event indicating that the action has started.

[0411] Returning to Figure 10 , the action handlers (e.g., Action1Handler, Action2Handler) may be responsible for executing a single category of supported (e.g., standardized) actions and may implement the corresponding action state machines. For example, the action handlers may receive events representing instructions to change the action state (e.g., start, stop, change) and may receive internal events (e.g., API callback calls, timeouts, etc.) from the modality service or from themselves. In some embodiments, the action handlers may directly publish (e.g., standardized interaction modeling) events (e.g., those indicating a change in the action state, such as started, completed, or updated) to stream XY.

[0412] The following sections describe some example implementations of some example modality services, namely an example GUI service that disposes of interactive visual content actions and an example animation service that disposes of bot gesture actions.

[0413] Example GUI Service。In some embodiments, a GUI service (e.g., which may correspond to the modal service B in Figure 9 ) handles interactive visual content actions (e.g., VisualInformationSceneAction, VisualChoiceSceneAction, VisualFormSceneAction) and corresponding events. In an example implementation, the GUI service converts a standardized event representing an indicated interactive visual content action (e.g., an indicated GUI update) into a call to an API of a user interface server, applies a modal policy that overrides an active action with the subsequently indicated action, and manages a stack of corresponding visual information scene actions (e.g., in response to receiving an event indicating a new interactive visual content action when there is at least one ongoing interactive visual content action). Thus, the GUI service can implement a GUI update that synchronizes the interactive visual content (e.g., visual information, a choice the user is prompted to make, or a field or form for the user to complete) with the current state of the interaction with the conversational AI.

[0414] In some embodiments, the GUI service can operate in cooperation with a user interface server (e.g., on the same physical device, on connected or networked physical devices, etc.), such as Figure 1 the user interface server 130. Generally, the user interface server can be responsible for managing the user interface and providing the user interface to a client device (e.g., Figure 1 the client device 101), and can provide the front-end components that make up the user interface of a web application (e.g., HTML files for constructing content, cascading style sheets for styling, JavaScript files for interaction). The user interface server can provide static assets for the user interface, such as images, fonts, and other resources, and / or can use any known technology to serve the user interface. The user interface server can act as a mediator between the client device and the GUI service, converting GUI inputs into standardized GUI input events and converting standardized GUI output events into corresponding GUI outputs.

[0415] The GUI service can manage an action state machine and / or an action stack for all interactive visual content actions. In an example implementation, the GUI service includes action handlers for each supported event of each supported interactive visual content action. Figures 12A - 12C Some example action handlers for some example interactive visual content action events according to some embodiments of the present disclosure are shown. More specifically, Figure 12A some example action handlers for some example visual information scene action events are shown, Figure 12BShows some example action handlers for some example visual selection action events, Figure 12C Shows some example action handlers for some example visual form action events.

[0416] For example, an interactive visual content event (e.g., generated by an interaction manager (e.g., Figure 1 interaction manager 190 or Figure 7 interaction manager 700)) can indicate the visualization of different types of visual information (e.g., in a 2D or 3D interface). In some embodiments, an interactive visual content event (e.g., a payload) includes a field that specifies or encodes a value representing a supported action type, which classifies the indicated action (e.g., VisualInformationSceneAction, VisualChoiceSceneAction, VisualFormSceneAction), action state (e.g., "init (initialize)", "scheduled", "starting", "running", "paused", "resuming", "stopping", or "finished"), some representation of the indicated visual content, and / or other attributes or information. The type of visual content specified by the event can depend on the action type.

[0417] For example, an event (e.g., a start event) for a visual information scene action (e.g., the payload of the event) can include fields that specify corresponding values, such as a specified title, a specified summary of the information to be presented, specified content to be presented (e.g., a list of information chunks to be shown to the user, where each chunk can contain specified text, a specified image (e.g., a description or identifier, such as a uniform resource locator), or both), one or more specified support cues that support or guide the user in making a selection, and / or others. Thus, an action handler for the corresponding (e.g., start) event for a visual information scene action can convert the event into a (e.g., JSON) representation of a modular GUI configuration that specifies content chunks, such as a carousel chunk of one or more specified support chunks, a title chunk for the specified title, an image and / or text chunk for the specified content, (e.g., continue, cancel) buttons, and / or other elements. Thus, the action handler can use these content chunks to generate a custom page by populating a visual layout (e.g., a prescribed template or shell visual layout with corresponding placeholders) of a GUI overlay (e.g., HTML), and can call a user interface server endpoint with the custom page to trigger the user interface server to render the custom page.

[0418] In some embodiments, an event for a visual selection action (e.g., a start event) (e.g., the payload of the event) may include fields specifying corresponding values, such as a specified prompt (e.g., describing the selection to be presented to the user), a specified image (e.g., a description or identifier of the image to be presented along with the selection, such as a Uniform Resource Locator), one or more specified support prompts to support or guide the user in making the selection, one or more specified options for the user to select (e.g., the text, image, and / or other content of each option), a specified selection type (e.g., configuring the type of selection the user can make, such as a selection, search bar, etc.), a specification of whether multiple selections are allowed, and / or others. Thus, an action handler for the corresponding (e.g., start) event for a visual selection action can convert the event into a (e.g., JSON) representation of a modular GUI configuration that specifies content blocks, such as a hint carousel block for one or more specified support blocks, a title block for a specified title, an image block for a specified image, an optional option grid block for specified options, (e.g., cancel) buttons, and / or other elements. Thus, the action handler can use these content blocks to generate a custom page by filling a visual layout of a GUI overlay (e.g., HTML) layout (e.g., a prescribed template or shell visual layout with corresponding placeholders), and can call a user interface server endpoint with the custom page to trigger the user interface server to render the custom page.

[0419] Figures 13A - 13F Illustrates some example interactions with visual selection in accordance with some embodiments of the present disclosure. For example, Figure 13A and Figure 13D illustrates a visual selection being presented between four captioned images, where an interactive avatar asks the user which image they like best. Figure 13B Illustrates a scenario where the user indicates the third image using touch input, Figure 13E and illustrates the same selection being made using verbal input. In these scenarios, the verbal input can be detected, routed to the corresponding sensing server and converted into a corresponding standardized event, routed to the interaction manager, and used to generate events indicating corresponding GUI updates, events indicating verbal bot responses, and / or events indicating response agent gestures. Thus, the events can be routed to the corresponding action server and executed. Figure 13C and Figure 13F illustrates an example bot response (e.g., visually emphasizing the selected option and responding with a verbal confirmation).

[0420] In some embodiments, an event for a visual form action (e.g., a start event) (e.g., the payload of the event) may include fields that specify corresponding values, such as a specified prompt (e.g., which describes the desired user input), a specified image (e.g., a description or identifier of an image that should be presented with the selection, such as a Uniform Resource Locator), one or more specified support prompts that support or guide the user in making a selection, one or more specified user inputs (e.g., where each specified user input may include a specified input type (e.g., number or date), a specified description (e.g., "personal email address" or "place of birth", etc.)) and / or others. Thus, an action handler for a corresponding (e.g., start) event for a visual form action may convert the event into a (e.g., JSON) representation of a modular GUI configuration that defines content blocks specified or otherwise represented by the event (e.g., the corresponding fields of the event), such as a carousel block of one or more specified support blocks, a title block for a specified prompt, an image block for a specified image, a list of input blocks for corresponding form fields representing specified inputs, (e.g., cancel) buttons, and / or other elements. Thus, the action handler may use these content blocks to generate a custom layout or page by populating the visual layout of a GUI overlay (e.g., HTML) page (e.g., a predefined template or shell visual layout having placeholders for the corresponding content blocks), and may call a user interface server endpoint with the custom layout or page to trigger the user interface server to render the custom layout or page.

[0421] In some embodiments, where an event for an interactive visual content action (e.g., a start event) (e.g., the payload of the event) specifies an image using a natural language description (e.g., "image of summer mountains"), the corresponding action handler for the event may trigger or perform an image search for the corresponding image. For example, the action handler may extract the natural language description of the desired image, engage any suitable image search tool (e.g., via a corresponding API), and send the natural language description of the desired image to the search tool. In some embodiments, the search tool returns an identifier (e.g., a Uniform Resource Locator) of a matching image, and the action handler may insert the identifier into a corresponding block in a custom page. Thus, the action handler may provide the custom page to a user interface server (which may retrieve the specified image using the inserted identifier) for rendering.

[0422] Figures 14A - 14L An example layout of visual elements of interactive visual content according to some embodiments of the present disclosure is shown. For example, Figure 14A An example GUI overlay 1420 presented over a scene with an interactive avatar is shown. Figures 14B - 14LShows some example layouts of visual element blocks that can be used as corresponding GUI overlays. These are only examples, and other layouts can be implemented within the scope of the present disclosure.

[0423] Example Animation Service . In some embodiments, an animation service (e.g., which can correspond to Figure 9 the modal service A in

[0424] can handle bot gesture actions (e.g., GestureBotAction) and corresponding events. In an example implementation, the animation service applies a modal policy that overrides the active action with the subsequently indicated action and creates a corresponding action stack in response to an incoming StartGestureBotAction event when one or more GestureBotActions are in progress. The animation service can manage the action state machine and action stack for all GestureBotActions, connect to the animation graph of the state machine that implements the transitions between animation states and animations, and instruct the animation graph to set the corresponding state variables. Figure 12D Shows some example action handlers for some example GestureBotAction events according to some embodiments of the present disclosure.

[0425] For example, a bot gesture action event (e.g., by an interaction manager (e.g., Figure 1 the interaction manager 190 of Figure 7The interaction manager 700) generates) can indicate a specified animation (e.g., in a 2D or 3D interface). In some embodiments, a bot gesture action event (e.g., a payload) can include a field that specifies or encodes a value representing a supported action type, which classifies the indicated action (e.g., GestureBotAction), action state (e.g., "init", "scheduled", "starting", "running", "paused", "resuming", "stopping", or "finished"), some representation of the indicated bot gesture, and / or other attributes or information. For example, an event for a bot gesture action (e.g., a start event) (e.g., the payload of the event) can include a field that specifies the bot gesture (e.g., a natural language description of the bot gesture or other identifier). Depending on the implementation, one or more categories or types of actions (e.g., bot expressions, postures, gestures, or other interactions or movements) can be standardized for the corresponding bot function, and the corresponding action event can specify the required action. Taking the bot gesture specified for the bot gesture action category (as a natural language description) as an example, the action handler for the corresponding (e.g., start) event for the bot gesture action category can extract the natural language description from the event, generate or access a sentence embedding of the natural language description of the bot gesture, perform a similarity search on the sentence embedding to obtain a description of the available animations, and select an animation using some similarity metric (e.g., nearest neighbor, within a threshold). In some embodiments, if the best match is above a certain specified threshold, the action handler can trigger the animation graph to play the corresponding animation for the user. In some embodiments, the action handler for the corresponding (e.g., start) event for the bot gesture action category can extract the natural language description from the event and generate an animation from the natural language description using any known generative technique (e.g., a text-to-motion model, text-to-animation technique, any other suitable animation technique).

[0426] Example Event Stream . The following discussion illustrates some possible event flows in an example implementation. For example, the following table represents a series of events that can be generated and distributed in an implementation where a bot is having a conversation with a user:

[0427]

[0428]

[0429] In this example, the event in the first row indicates the completion of detecting the user's utterance ("Hello!"), which triggers an event indicating that the bot starts to respond with an utterance ("Hello there!"). The event in the second row indicates that the bot has started the utterance, and the event in the third row indicates that the bot has completed the utterance.

[0430] The following table shows a series of events that can be generated and distributed in an implementation where the bot interacts with the user through gestures, emotions, and displays:

[0431]

[0432]

[0433] In this example, the event in the first row indicates that the bot has completed prompting the user "Which option?", which triggers an event indicating that the GUI presents a visual selection. The event in the second row indicates starting a two - second timer. The event in the third row indicates that the visual selection has been presented, and the event in the fourth row indicates that the timer has been started, which triggers an event indicating that the bot points to the display of the visual selection. The event in the fifth row indicates that the pointing gesture has started, and the event in the sixth row indicates that the pointing gesture has been completed. The event in the seventh row indicates that the two - second timer has completed, which triggers the bot's utterance ("Do you need more time?"). The event in the eighth row indicates that the bot's utterance has started. The event in the ninth row indicates the completion of detecting the detected user gesture (nodding), which triggers a responsive proxy gesture (leaning forward). The event in the tenth row indicates that the bot gesture has started. The event in the eleventh row indicates the completion of the bot's utterance ("Do you need more time?"), the event in the twelfth row indicates the completion of the bot's gesture (leaning forward), and the event in the last row indicates the detected start of the detected user expression (happy).

[0434] In various situations, it may be beneficial to indicate that an interaction system or one of its components (e.g., a sensing server that controls input processing, an action server that implements bot actions) takes some action to anticipate an event that the interaction manager (e.g., an interpreter) expects next from the user or the system, or otherwise signal an expectation. The following discussion illustrates some possible anticipatory actions and other example features in an example implementation.

[0435] For example, Figure 15Shows an example event flow 1500 of user speech actions in an implementation where user 1518 speaks to interactive avatar 1504. In this example, interactive avatar 1504 is implemented using user interface 1516 (e.g., microphone and audio interface), voice activity detector 1514, automatic speech recognition system 1512, and action server 1510 responsible for handling events of user speech actions (e.g., UtteranceUserAction 1508). In this example, action server 1510 acts as both a sensing server and an action server, converting sensing input into standardized events and executing standardized events that indicate certain actions. Interaction manager 1506 can make decisions for interactive avatar 1504. Although interaction manager 1506 and interactive avatar 1504 are shown as separate components, interaction manager 1506 can be considered part of interactive avatar 1504.

[0436] In the example flow, at step 1520, user 1518 starts speaking. At step 1522, voice activity detector 1514 picks up the speech and sends the speech stream to automatic speech recognition system 1512. At step 1524, voice activity detector 1514 notifies action server 1510 that voice activity has been detected, and at step 1526, automatic speech recognition system 1512 streams the transcribed speech to action server 1510. Thus, at step 1528, action server 1510 generates a standardized event indicating that the detected user speech has started (e.g., including the transcribed speech) and sends the event (e.g., UtteranceUserActionStarted) to event gateway 1502, which is picked up by interaction manager 1506 at step 1530.

[0437] The following steps 1532 - 1546 can be executed in a loop. In step 1532, the user has finished speaking a few words, and in step 1534, the automatic speech recognition system 1512 sends a partial transcript to the action server 1510. In step 1536, the action server 1510 generates a standardized event indicating an update to the detected user utterance (e.g., including the transcribed speech), and sends this event (e.g., UtteranceUserActionTranscriptUpdated) to the event gateway 1502. In step 1538, the interaction manager 1506 picks up this event. In step 1540, the user speaks louder, and in step 1542, the voice activity detector 1514 detects the increase in volume and notifies the action server 1510 of the detected volume change. In step 1544, the action server 1510 generates a standardized event indicating an update to the intensity of the user utterance (e.g., including the detected intensity or volume level), and sends this event (e.g., UtteranceUserActionIntensityUpdated) to the event gateway 1502. In step 1546, the interaction manager 1506 picks up this event.

[0438] In some embodiments, in step 1548, the interaction manager 1506 generates a standardized event indicating the expectation that the user is about to stop speaking and / or indicating that the interactive avatar 1504 should take some preparatory actions, and the interaction manager 1506 sends this event (e.g., StopUtteranceUserAction) to the event gateway 1502. In step 1550, the action server 1510 picks up this event. In response, in step 1552, the action server 1510 instructs the voice activity detector 1514 to reduce the audio hold time (e.g., the period during which the detected speech signal persists before being considered inactive or silent).

[0439] In step 1554, the user stops speaking. In step 1556, the voice activity detector 1514 detects speech inactivity and stops transmitting the speech stream to the automatic speech recognition system 1512. In step 1558, the automatic speech recognition system 1512 stops streaming the transcript to the action server 1510. In step 1560, the hold time expires. In step 1562, the voice activity detector 1514 notifies the action server 1510 of the detected speech inactivity. Thus, in step 1564, the action server 1510 generates a standardized event indicating the detected completion of the detected user utterance, and sends this event (e.g., UtteranceUserActionFinished) to the event gateway 1502. In step 1566, the interaction manager 1506 picks up this event.

[0440] Figure 16 An example event flow 1600 of user utterance actions in an implementation scenario where user 1618 converses with chatbot 1604 is shown. In this example, chatbot 1604 is implemented using user interface 1616 (e.g., a hardware or software keyboard and driver), timer 1612, and action server 1610 responsible for handling events of user utterance actions (e.g., UtteranceUserAction 1608). In this example, action server 1610 acts as both a sensing server and an action server, converting sensing inputs (e.g., detected text, typing rate) into standardized events and executing standardized events that indicate certain actions. Interaction manager 1606 can make decisions for chatbot 1604. Although interaction manager 1606 and chatbot 1604 are shown as separate components, interaction manager 1606 can be considered part of chatbot 1604.

[0441] In the example flow, at step 1620, user 1618 starts typing. At step 1622, user interface 1616 notifies action server 1610 that typing has started, and at step 1624, action server 1610 generates a standardized event indicating that the detected user utterance has started and sends this event (e.g., UtteranceUserActionStarted) to event gateway 1602, which is picked up by interaction manager 1606 at step 1626.

[0442] The following steps 1628 - 1640 can be executed in a loop. At step 1628, user interface 1616 sends the typed text to action server 1610, and at step 1630, action server 1610 generates a standardized event indicating a detected update to the detected user utterance (e.g., including the typed text) and sends this event (e.g., UtteranceUserActionTranscriptUpdated) to event gateway 1602, which is picked up by interaction manager 1606 at step 1634. At step 1632, the user starts typing faster, and at step 1636, user interface 1616 detects the increase in typing speed and notifies action server 1610 of the detected speed change. At step 1638, action server 1610 generates a standardized event indicating a detected update to the detected user utterance intensity (e.g., including the detected intensity or typing speed) and sends this event (e.g., UtteranceUserActionIntensityUpdated) to event gateway 1602, which is picked up by interaction manager 1606 at step 1640.

[0443] In some embodiments, at step 1642, the interaction manager 1606 generates a standardized event indicating the expectation that the user is about to stop typing and / or indicating that the chatbot 1604 should take some preparatory action, and the interaction manager 1606 sends this event (e.g., StopUtteranceUserAction) to the event gateway 1602, and at step 1644 the action server 1610 picks up this event. In response, at step 1646, the action server 1610 reduces the timeout after keystrokes (e.g., the period of detected inactivity or typing latency that is interpreted as the completion of an utterance).

[0444] At step 1648, the user stops typing. At step 1650, the user interface 1616 sends a notification of the typing stop to the action server 1610, and at step 1652, the action server 1610 instructs the timer 1612 to start. At step 1654, the timer 1612 notifies the action server 1610 that the timer has elapsed, and the action server 1610 notifies the user interface 1616 to block further input into the input field. At step 1658, the user interface 1616 sends the completed text input to the action server 1610. Thus, at step 1660, the action server 1610 generates a standardized event indicating the detected completion of the detected user utterance (e.g., including the completed text input), and sends this event (e.g., UtteranceUserActionFinished) to the event gateway 1602, and at step 1662 the interaction manager 1606 picks up this event.

[0445] Figure 17 An example event flow 1700 for bot expected actions in an implementation where a user 1718 is talking to an interactive avatar 1704 is shown. In this example, the interactive avatar 1704 is implemented using a client device 1716 (e.g., including a microphone and an audio interface), an automatic speech recognition system 1714, and an action server 1712, and the action server 1712 is responsible for handling events for user utterance actions (e.g., UtteranceUserAction 1710) and events for bot expected actions for user utterance actions (e.g., BotExecepectionAction 1708). In this example, the action server 1712 acts as both a sensing server and an action server, converting sensing inputs into standardized events and executing standardized events indicating certain actions. The interaction manager 1706 can make decisions for the interactive avatar 1704. Although the interaction manager 1706 and the interactive avatar 1704 are shown as separate components, the interaction manager 1706 can be considered part of the interactive avatar 1704.

[0446] In step 1720, the interaction manager 1706 generates a standardized event indicating that the user's utterance is expected to start soon and representing an instruction to take some preparatory actions in anticipation of the user's utterance, and sends this event (e.g., StartBotExpectionAction(UtteranceUserActionFinished)) to the event gateway 1702. In step 1722, the action server 1712 picks up this event. Note that in this example, the argument for identifying the expected keyword is the expected target event (e.g., the completion of the user's utterance), which can trigger a corresponding stop action that indicates that the expectation of the interaction manager 1706 has been met or is no longer relevant, which in itself can trigger a reversal of the preparatory actions, but this syntax is only for example and need not be used. In response, in step 1724, the action server 1712 notifies the client device 1716 to disable its audio output. In step 1726, it notifies the client device 1716 to enable its microphone. In step 1728, it notifies the automatic speech recognition system 1714 to enable automatic speech recognition. In step 1730, the action server 1712 generates a standardized event confirming that the bot's expected action has started and / or indicating that the preparatory action has been initiated, and sends this event (e.g., BotExpectionActionStarted(UtteranceUserActionFinished)) to the event gateway 1702. In step 1732, the interaction manager 1706 picks up this event.

[0447] In some embodiments, when user 1718 starts speaking, speech (not shown) is detected, and at step 1734, action server 1712 generates a standardized event indicating that the detected user utterance has started and sends the event (e.g., UtteranceUserActionStarted) to event gateway 1702, and at step 1736 interaction manager 1706 picks up the event. Once user 1718 stops speaking and the end of the utterance is detected (not shown), action server 1712 generates a standardized event indicating the detected completion of the detected user utterance and sends the event (e.g., UtteranceUserActionFinished) to event gateway 1702 (not shown), and at step 1738 interaction manager 1706 picks up the event. In this example, interaction manager 1706 is programmed to stop the bot expected action in response to receiving an event indicating the detected completion of the detected user utterance, so at step 1740, interaction manager 1706 generates a standardized event indicating that the expected user utterance has been completed and indicating the abandonment of the preparatory action and sends the event (e.g., StopBotExpectionAction(UtteranceUserActionFinished)) to event gateway 1702, and at step 1742 action server 1712 picks up the event. In response, at step 1744, action server 1712 instructs automatic speech recognition system 1714 to stop automatic speech recognition, and at step 1746, instructs client device 1716 to disable its microphone. At step 1748, action server 1712 generates a standardized event confirming that the bot expected action has been completed and / or indicating that the preparatory action has been abandoned and sends the event (e.g., BotExpectionActionFinished(UtteranceUserActionFinished)) to event gateway 1702, and at step 1750 action server 1712 picks up the event.

[0448] Flowchart . Now refer to Figures 18 - 27 , each block of the methods 1800 - 2700 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in a memory. The methods 1800 - 2700 can also be embodied as computer-usable instructions stored on a computer storage medium. The methods 1800 - 2700 can be provided by a stand-alone application, a service, or a hosted service (independently or in combination with another hosted service) or a plug-in of another product, to name a few. Additionally, the methods 1800 - 2700 are illustrated by example systems (e.g., Figure 1The interactive system 100) is described. However, these methods may be additionally or alternatively performed by any one system or any combination of systems, including but not limited to the systems described herein.

[0449] Figure 18 FIG. 1800 is a flowchart of a method for generating a representation of a response agent action classified using an interaction classification scheme, according to some embodiments of the present disclosure. At block B1802, method 1800 includes: receiving, by an interpreter of an interactive agent platform associated with an interactive agent, one or more representations of one or more detected user actions classified using an interaction classification scheme. For example, with respect to Figure 1 the interactive system 100, some representations of user input (e.g., gestures detected by the visual microservice 110, voice commands detected by the voice detection microservice 120, or touch or click inputs detected by the UI server 130) may be forwarded to the corresponding sensing server in the sensing server 160 responsible for the corresponding interaction channel. The sensing server 160 may convert the user input into a standardized representation of the corresponding event defined by the interaction classification scheme and place the event on the event gateway 180. The interaction manager 190 may implement an interpreter that is subscribed to or otherwise configured to pick up or receive these events from the event gateway 180.

[0450] At block B1804, method 1800 includes: generating, at least based on the interpreter executing one or more instruction lines of one or more interaction flows, one or more representations of one or more response agent actions classified using an interaction classification scheme, where the interaction flows are written in an interaction modeling language and indicate generation of one or more agent actions in response to one or more detected user actions. For example, with respect to Figure 1 the interactive system 100, the interpreter implemented by the interaction manager 190 may support the interaction modeling language, the code implementing the decision logic of the interactive agent may be written in the interaction modeling language, may be loaded onto or otherwise accessed by the interaction manager 190, and may be executed by the interaction manager 190. Thus, the interaction manager 190 may process events from the event gateway 180 (e.g., using an event-driven state machine), determine what interactions to participate in, and generate commands and forward the commands as corresponding events in the standardized representation to the event gateway 180.

[0451] Figure 19is a flowchart showing method 1900 for generating a representation of a response agent action based at least on performing one or more interaction flows. At block B1902, method 1900 includes: receiving, by an interpreter of an interactive agent platform that supports simultaneous execution of agent actions in different interaction modalities, one or more representations of one or more detected user actions. For example, with respect to Figure 1 the interaction system 100, the interaction manager 190 may implement an interpreter that is subscribed to or otherwise configured to pick up or receive events from the event gateway 180 that represent detected user actions. The interaction manager 190 may implement the decision logic of an interactive agent written in an interaction modeling language, and the interaction modeling API and / or language used by the interaction manager 190 may define mutually exclusive interaction modalities such that events indicating actions in different interaction modalities may be executed independently (e.g., simultaneously) of each other by corresponding action servers 170 dedicated to the respective interaction modalities.

[0452] At block B1904, method 1900 includes: generating, based at least on one or more instruction lines of one or more interaction flows that the interpreter executes in response to one or more detected user actions, one or more representations of one or more response agent actions. For example, with respect to Figure 1 the interaction system 100, the code that implements the decision logic of the interactive agent and defines one or more interaction flows may be written in an interaction modeling language, loaded onto or otherwise accessible to the interpreter implemented by the interaction manager 190, and executed by the interpreter. In this way, the interaction manager 190 may process events from the event gateway 180 (e.g., representing detected user actions) (e.g., using an event-driven state machine), determine which interactions to participate in, and generate commands and forward the commands as corresponding events in a standardized representation to the event gateway 180.

[0453] Figure 20 is a flowchart showing method 2000 for triggering an interactive avatar to provide backchannel mechanism feedback. At block B2002, method 2000 includes: receiving, by an interpreter associated with an interactive avatar that supports non-sequential human-computer interaction, one or more representations of one or more detected initiations of one or more user actions, and at block B2004 includes: triggering, based at least on one or more instruction lines of one or more interaction flows that the interpreter executes in response to one or more detected initiations, the interactive avatar to provide backchannel mechanism feedback during one or more user actions. For example, Figure 1The interaction system 100 can use various features described herein to support non-sequential interactions, such as an event-driven architecture, an interpreter that supports various keywords, and / or decoupling of sensing processing, interaction decision-making, and action execution. For example, to support the execution of a response agent action before a triggering user action (e.g., utterance) is completed, the sensing server 160 can generate an event representing the initiation of the detected user action and provide the event (via the event gateway 180) to the interaction manager 190, and the interaction manager 190 can check whether there is a matching active (e.g., interrupted) flow waiting for such an event. Since the sensing server 160, the interaction manager 190, and the action server 170 can operate independently of each other, the sensing server 160 can continue to process user input while the interaction manager 190 generates an event representing the response action and triggers the action server 170 to execute the response action (e.g., back-channel mechanism feedback).

[0454] Figure 21 is a flowchart showing a method 2100 for generating interaction modeling events that command an interactive agent to execute a response agent or scenario action according to certain embodiments of the present disclosure. At block B2102, the method 2100 includes: receiving, via one or more event gateways and by an interaction manager associated with the interactive agent, one or more first interaction modeling events that represent at least one of the following: one or more detected user actions, one or more indicated agent actions, or one or more indicated scenario actions. For example, with respect to Figure 1 the interaction system 100, the sensing server 160 can convert the detected user input into a standardized representation of a corresponding event and place the event on the event gateway 180. Additionally, with respect to Figure 6 , the interaction manager 640 can generate internal events 660 representing internal state changes (e.g., flow state changes) or indicated bot actions, and / or the action server 670 can generate events 665 representing confirmations of action state changes. Thus, Figure 1 the interaction manager 190 of Figure 6 and / or the interaction manager 640 of

[0455] can be subscribed to or otherwise configured to pick up or receive events from the event gateway 180. At block B2104, the method 2100 includes: generating, at least based on the interaction manager using an event-driven state machine to process the one or more first interaction modeling events, one or more second interaction modeling events that command the interactive agent to execute at least one of the following: one or more response agent actions or one or more response scenario actions. For example, with respect to Figure 6The event-driven interaction system 600, the interaction manager 640 (which may correspond to Figure 1 and / or Figure 2 the interaction manager 190) may be responsible for determining what actions the interaction system 600 should perform in response to user actions or other events (e.g., standardized input events 630, internal events 660, events 665 indicating confirmation of a change in action state). The interaction manager 640 may interact with the rest of the interaction system 600 through an event-driven mechanism. Thus, the interaction manager 640 may evaluate various types of events (e.g., standardized input events 630, internal events 660, events 665 indicating confirmation of a change in action state), determine which actions to perform, and generate corresponding bot action events 650 indicating the actions or event instructions for updating certain other aspects of the scenario (e.g., interactive visual content actions).

[0456] Figure 22 is a flowchart showing a method 2200 for triggering one or more response agents or scenario actions specified by one or more matching interaction flows according to some embodiments of the present disclosure. At block B2202, method 2200 includes: tracking one or more interrupted interaction flows representing one or more human-machine interactions. For example, with respect to Figure 6 the event-driven interaction system 600, the interaction manager 640 may support and track multiple active flows (e.g., interrupted at the corresponding event matchers).

[0457] At block B2204, method 2200 includes: checking one or more incoming interaction events to find one or more matching interaction flows among the one or more interrupted interaction flows. For example, with respect to Figure 6 the event-driven interaction system 600, the interaction manager 640 may use an event-driven state machine to listen for events that match the event matchers of the active flows. For example, Figure 7 the flow matcher 740 may evaluate incoming events to determine whether they match the event matchers of the active flows, process the incoming events sequentially (e.g., from an internal event queue 790, from some other queue or event gateway, such as Figure 1 the event gateway 180), and for each event, test whether the event matchers specified by each active flow match the event.

[0458] At block B2206, method 2200 includes: triggering one or more response agents or scenario actions specified by one or more matching interaction flows in response to identifying one or more matching interaction flows. For example, with respect to Figure 6 the event-driven interaction system 600, the interaction manager 640 may trigger the corresponding events and actions specified in the flow that matches the event being tested. For example,Figure 7 The flow matcher 740 can instruct the flow execution component 750 to advance (e.g., non-conflicting) matching flows, and the advancing flow can instruct the flow execution component 750 to generate an outgoing event indicating a certain action.

[0459] Figure 23 is a flowchart showing a method 2300 for generating a response agent or scenario action based at least on prompting one or more large language models according to some embodiments of the present disclosure. At block B2302, the method 2300 includes: receiving, by an interpreter of an interactive agent platform, one or more representations of one or more detected user actions. For example, regarding Figure 1 the interactive system 100, the sensing server 160 can convert the detected user input into a standardized representation of a corresponding event and place the event on the event gateway 180, and the interaction manager 190 can implement an interpreter that is subscribed to or otherwise configured to pick up or receive events from the event gateway 180.

[0460] At block B2304, the method 2300 includes: generating, at least based on the interpreter prompting one or more large language models (LLMs) and evaluating one or more matches of one or more representations of one or more detected user actions with one or more interrupted interaction flows, one or more representations of one or more response agents or scenario actions. For example, regarding Figure 7 , the interpreter 710 can support the use of natural language descriptions and the use of one or more LLMs. For example, the interpreter 710 can prompt the LLM to generate a natural language description defining one or more instruction lines of a flow, generate one or more instruction lines for a specified flow, determine whether an event matches the flow description of an active flow, determine whether a non-matching event matches the name and / or instructions of an active flow, generate a flow in response to a non-matching event, and / or otherwise.

[0461] Figure 24 is a flowchart showing a method 2400 for generating one or more outgoing interaction modeling events instructing one or more action servers to perform one or more response agents or scenario actions according to some embodiments of the present disclosure. At block B2402, the method 2400 includes: generating, by one or more sensing servers in one or more input interaction channels, one or more incoming interaction modeling events representing one or more detected user actions. For example, regarding Figure 1 the interactive system 100, the sensing server 160 can convert the detected user input into a standardized representation of a corresponding event and place the event on the event gateway 180.

[0462] At block B2404, method 2400 includes generating, by an interaction manager, one or more outgoing interaction modeling events based at least on one or more incoming interaction modeling events, the outgoing interaction modeling events indicating that one or more action servers in one or more output interaction channels perform one or more response agent actions or scenario actions associated with an interactive agent. For example, with respect to Figure 1 the interaction system 100, the interaction manager 190 may implement an interpreter that is subscribed to or otherwise configured to pick up or receive events from the event gateway 180, process the events (e.g., using an event-driven state machine), determine what interactions to participate in, and generate commands and forward the commands as corresponding events in a standardized representation to the event gateway 180. The action server 170 responsible for the corresponding interaction channel may be subscribed to or otherwise configured to pick up or receive those events that it is responsible for executing from the event gateway 180. Thus, the action server 170 may execute, schedule, and / or otherwise dispose of events for the corresponding interaction modality, engaging with the corresponding services that control the corresponding output interfaces.

[0463] Figure 25 is a flow diagram showing a method 2500 for generating a visual layout representing an update specified by an event according to some embodiments of the present disclosure. At block B2502, method 2500 includes receiving, by one or more action servers that handle one or more visual content overlays supplementing one or more conversations with an interactive agent, one or more events that represent one or more visual content actions classified using an interaction classification scheme and indicate one or more updates to one or more overlays in one or more GUIs. For example, with respect to Figure 9 , the action server 930 may include a GUI service (e.g., modal service B) that handles interactive visual content and corresponding events. Interactive visual content events (e.g., generated by an interaction manager (e.g., Figure 1 the interaction manager 190 of Figure 7The interaction manager 700) generates) visualizations that can indicate different types of visual information (e.g., in a 2D or 3D interface). In some embodiments, an interactive visual content event (e.g., a payload) includes a field that specifies or encodes a value representing a supported action type, which classifies the indicated action (e.g., VisualInformationSceneAction, VisualChoiceSceneAction, VisualFormSceneAction), action state (e.g., "init", "scheduled", "starting", "running", "paused", "resuming", "stopping", or "finished"), some representation of the indicated visual content, and / or other attributes or information.

[0464] At block B2504, method 2500 includes: generating, by one or more action servers, one or more visual layouts that represent one or more updates specified by one or more events. For example, regarding Figure 9 , action server 930 may include a GUI service (e.g., Modal Service B), the GUI service including action handlers for each supported event of each supported interactive visual content action, and the action handler for a corresponding (e.g., start) event of a visual information scene action may convert the event into a (e.g., JSON) representation of a modular GUI configuration that specifies content blocks, such as a hint carousel block for one or more specified support blocks, a header block for a specified title, an image and / or text block for specified content, (e.g., continue, cancel) buttons, and / or other elements. Thus, the action handler may use these content blocks to generate a custom page by populating a visual layout (e.g., a predefined template or shell visual layout with corresponding placeholders) for a GUI overlay (e.g., HTML) layout, and may call a user interface server endpoint with the custom page to trigger the user interface server to render the custom page.

[0465] Figure 26 is a flowchart of a method 2600 for triggering an animation state of an interactive agent according to some embodiments of the present disclosure. At block B2602, method 2600 includes: receiving, by one or more action servers that handle gesture animations of the interactive agent, one or more first interaction modeling events that indicate one or more target states of one or more agent gestures represented using an interaction classification scheme. For example, regarding Figure 9, the action server 930 may include an animation service (e.g., Modal Service A) that handles bot movement and / or gesture actions (e.g., GestureBotAction) and corresponding events. For example, a bot gesture action event (e.g., generated by an interaction manager (e.g., Figure 1 the interaction manager 190 or Figure 7 the interaction manager 700)) may use a field specifying or encoding a value representing the supported action type to indicate a prescribed animation (e.g., in a 2D or 3D interface), which classifies the indicated action (e.g., GestureBotAction), action state (e.g., start, started, updated, stop, finished), a certain representation of the indicated bot gesture, and / or other attributes or information.

[0466] At block B2604, method 2600 includes: triggering, by one or more action servers, one or more animation states of an interactive agent corresponding to one or more target states of one or more agent gestures indicated by one or more first interaction modeling events. For example, regarding Figure 9 , the action server 930 may include an animation service (e.g., Modal Service A) that includes action handlers for each supported event of each supported bot gesture action. Figure 12D Some example action handlers for some example GestureBotAction events according to some embodiments of the present disclosure are shown. Taking the bot gesture specified as a natural language description for a bot gesture action as an example, the action handler for the corresponding (e.g., start) event of the bot gesture action may extract the natural language description from the event, generate or access a sentence embedding of the natural language description of the bot gesture, perform a similarity search on the sentence embeddings of the available animation descriptions using it, and select an animation using a certain similarity metric (e.g., nearest neighbor, within a threshold).

[0467] Figure 27 is a flowchart of a method 2700 for performing one or more preparation actions according to some embodiments of the present disclosure. At block B2702, method 2700 includes: receiving, by one or more servers associated with an interactive agent, one or more first interaction modeling events that indicate one or more preparation actions associated with the expectation that one or more specified events will occur and are represented using an interaction classification scheme. For example, regarding Figure 17, at step 1720, the interaction manager 1706 generates a standardized event indicating that the user's utterance is expected to start soon and indicating a preparation action, and sends this event (e.g., StartBotExpectionAction(UtteranceUserActionFinished)) to the event gateway 1702, and at step 1722 the action server 1712 picks up this event.

[0468] At block B2704, method 2700 includes: performing one or more preparation actions by a first server. For example, regarding Figure 17 , at step 1724, the action server 1712 notifies the client device 1716 to disable its audio output, at step 1726, notifies the client device 1716 to enable its microphone, and at step 1728, notifies the automatic speech recognition system 1714 to enable automatic speech recognition. At step 1730, the action server 1712 generates a standardized event confirming that the bot expectation action has started and / or indicating that the preparation action has been initiated, and sends this event (e.g., BotExpectionActionStarted(UtteranceUserActionFinished)) to the event gateway 1702, and at step 1732 the interaction manager 1706 picks up this event gateway.

[0469] The systems and methods described herein can be used for various purposes, such as but not limited to, for machine (e.g., robot, vehicle, construction machinery, warehouse vehicle / machine, autonomous, semi-autonomous, and / or other machine types) control, machine movement, machine driving, synthetic data generation, model training (e.g., using real, augmented, and / or synthetic data, such as synthetic data generated using a simulation platform or system, synthetic data generation techniques (e.g., but not limited to the techniques described herein), etc.), perception, augmented reality (AR), virtual reality (VR), mixed reality (MR), robotics, security and surveillance (e.g., in smart city implementations), autonomous or semi-autonomous machine applications, deep learning, environment simulation, object or actor simulation, and / or digital twins, data center processing, conversational AI, optical transmission simulation (e.g., ray tracing, path tracing, etc.), distributed or collaborative content creation of 3D assets (e.g., using Universal Scene Description (USD) data, such as OpenUSD and / or other data types), cloud computing, generative artificial intelligence (e.g., using one or more diffusion models, transformer models, etc.), and / or any other suitable applications.

[0470] The disclosed embodiments may be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots or robotic platforms, aerial systems, medical systems, boating systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations (e.g., in driving or vehicle simulation, in robot simulation, in smart city or surveillance simulation, etc.), systems for performing digital twin operations (e.g., in combination with a collaborative content creation platform or system, such as but not limited to NVIDIA's OMNIVERSE and / or another platform, system, or service using USD or OpenUSD data types), systems implemented using edge devices, systems incorporating one or more virtual machines (VMs), systems for performing synthetic data generation operations (e.g., using one or more neural radiance fields (NERF), gaussian splat techniques, diffusion models, transformer models, etc.), systems implemented at least partially in a data center, systems for performing conversational AI operations, systems implementing one or more language models (e.g., one or more large language models (LLM), one or more vision language models (VLM), one or more multimodal language models, etc.), systems for performing optical transmission simulation, systems for performing collaborative content creation of 3D assets (e.g., using Universal Scene Description (USD) data, such as OpenUSD, computer-aided design (CAD) data, 2D and / or 3D graphics or design data, and / or other data types), systems implemented at least partially using cloud computing resources, and / or other types of systems.

[0471] In some embodiments, the systems and methods described herein may be executed within a 3D content collaboration platform (e.g., NVIDIA's OMNIVERSE) for 3D rendering, industrial digitization, generative physics AI, and / or other use cases, applications, or services. For example, the content collaboration platform may host a framework for developing and / or deploying interactive agents (e.g., interactive avatars), and may include systems for using or developing Universal Scene Description (USD) (e.g., OpenUSD) data for managing objects, features, scenes, etc. within digital environments, simulation environments, etc. The platform may include real physics simulation, such as using NVIDIA's PhysX SDK, to simulate real physical phenomena and physical interactions with virtual objects, characters, simulations, or other types of 3D content hosted by the platform. The platform may integrate OpenUSD and ray tracing / path tracing / light transport simulation (e.g., NVIDIA's RTX rendering technology) into software tools and rendering workflows. In some embodiments, the development and / or deployment of interactive agents (e.g., interactive bots or robots) may leverage one or more cloud services and / or machine learning models (e.g., neural networks, large language models). For example, NVIDIA's Avatar Cloud Engine (ACE) is a set of cloud-based AI models and services designed to create and manage interactive, realistic avatars using hosted natural language processing, speech recognition, computer vision, and / or conversational AI services. In some embodiments, interactive agents may be developed and / or deployed as part of an application hosted by a (e.g., streaming) platform (e.g., a cloud-based gaming platform (e.g., NVIDIA GeFORCE NOW)). Thus, interactive agents such as digital avatars may be developed and / or deployed for various applications such as customer service, virtual assistants, interactive entertainment or gaming, digital twins (e.g., for video conference participants), education or training, healthcare, virtual or augmented reality experiences, social media interactions, marketing and advertising, and / or other applications.

[0472] Example language model

[0473] In at least some embodiments, a language model, such as a large language model (LLM), a vision language model (VLM), a multimodal language model (MMLM), and / or other types of generative artificial intelligence (AI) may be implemented. These models may be capable of understanding, summarizing, translating, and / or otherwise generating text (e.g., natural language text, code, etc.), images, videos, computer-aided design (CAD) assets, OMNIVERSE, and / or METAVERSE file information (e.g., USD format such as OpenUSD) and / or the like based on the context provided in an input prompt or query. In an embodiment, these language models may be considered "large" as these models are trained on a vast dataset and have an architecture with a large number of learnable network parameters (weights and biases) - e.g., millions or billions of parameters. The LLM / VLM / MMLM / etc. may be implemented to summarize text data, analyze data (e.g., text, images, videos, etc.), and extract insights from data (e.g., text, images, videos, etc.), as well as generate new text / images / videos / etc. in a user-specified style, tone, and / or format. In an embodiment, the LLM / VLM / MMLM / etc. of the present disclosure may be specialized for text processing, while in other embodiments, a multimodal LLM may be implemented to accept, understand, and / or generate text and / or other types of content, such as images, audio, 2D, and / or 3D data (e.g., USD format) and / or videos. For example, a vision language model (VLM) or more specifically a multimodal language model (MMLM) may be implemented to accept images, videos, audio, text, 3D designs (e.g., CAD), and / or other input data types and / or generate or output images, videos, audio, text, 3D designs, and / or other output data types.

[0474] Various types of LLM / VLM / MMLM / etc. architectures can be implemented in various embodiments. For example, different architectures that use different techniques to understand and generate outputs (such as text, audio, video, images, 2D and / or 3D designs or asset data, etc.) can be implemented. In some embodiments, LLM / VLM / MMLM / etc. architectures (such as recurrent neural networks (RNNs) or long short-term memory networks (LSTMs)) can be used, while in other embodiments, transformer architectures (such as architectures that rely on self-attention and / or cross-attention (e.g., between context data and text data) mechanisms) can be used to understand and identify the relationships between words or tokens and / or context data (such as other text, videos, images, design data, USD, etc.). One or more generative processing pipelines including LLM / VLM / MMLM / etc. can also include one or more diffusion blocks (such as denoisers). The LLM / VLM / MMLM / etc. of the present disclosure can include encoder and / or decoder blocks. For example, discriminative or encoder-only models (such as BERT (Bidirectional Encoder Representations from Transformers)) can be implemented for tasks related to language understanding (such as classification, sentiment analysis, question answering, and named entity recognition). As another example, generative or decoder-only models (such as GPT (Generative Pretrained Transformer)) can be implemented for tasks related to language and content generation (such as text completion, story generation, and dialogue generation). LLM / VLM / MMLM / etc. that include encoder and decoder components (such as T5 (Text-to-Text Transformer)) can be implemented to understand and generate content, such as for translation and summarization. These examples are not intended to be limiting, and any architecture type (including but not limited to the architecture types described herein) can be implemented according to specific embodiments and the tasks performed using LLM / VLM / MMLM / etc.

[0475] In various embodiments, unsupervised learning can be used to train LLM / VLM / MMLM / etc., where the LLM / VLM / MMLM / etc. learns patterns from large amounts of unlabeled text / audio / video / image / design / USD / etc. data. Due to extensive training, in embodiments, the model may not require task-specific or domain-specific training. The LLM / VLM / MMLM / etc. that has been extensively pre-trained on large amounts of unlabeled data can be referred to as a base model and can be proficient in various tasks such as question answering, summarization, filling in missing information, translation, image / video / design / USD / data generation. Some LLM / VLM / MMLM / etc. can be customized for specific use cases using techniques such as prompt tuning, fine-tuning, retrieval-augmented generation (RAG), adding adapters (e.g., custom neural networks and / or neural network layers for adjusting or conditioning prompts or tokens to bias the language model towards a specific task or domain), and / or other fine-tuning or customization techniques for optimizing the model for a specific task and / or within a specific domain for a specific use case.

[0476] In some embodiments, the LLM / VLM / MMLM / etc. of the present disclosure can be implemented using various model alignment techniques. For example, in some embodiments, guardrails can be implemented to identify incorrect or unwanted inputs (e.g., prompts) and / or outputs of the model. In this process, the system can use the guardrails and / or other model alignment techniques to prevent the processing of specific unwanted inputs using the LLM / VLM / MMLM / etc., and / or to prevent the output or presentation (e.g., display, audio output, etc.) of information generated using the LLM / VLM / MMLM / etc. In some embodiments, one or more additional models (or their layers) can be implemented to identify problems with the inputs and / or outputs of the model. For example, these "protective" models can be trained to identify "safe" or otherwise okay or wanted inputs and / or outputs and / or "unsafe" or otherwise unwanted inputs and / or outputs for a specific application / implementation. Thus, the LLM / VLM / MMLM / etc. of the present disclosure is less likely to output language / text / audio / video / design data / USD data / etc. that may be offensive, vulgar, inappropriate, unsafe, out of scope, and / or unwanted for a specific application / implementation.

[0477] In some embodiments, an LLM / VLM / etc. can be configured to or capable of accessing or using one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc. For example, for certain tasks or operations that the model is not ideally suited for, the model can have instructions for accessing one or more plugins (e.g., third-party plugins) to assist in processing the current input (e.g., as a result of training, and / or based on the instructions in a given prompt). In such an example, when at least a portion of the prompt is related to a restaurant or the weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs) to retrieve relevant information. Another example is that if at least a portion of the response requires mathematical calculations, the model can access one or more math plugins or APIs to assist in solving the problem, and then the response from the plugin and / or API can be used in the output of the model. This process can be repeated (e.g., recursively) any number of times of iteration and use any number of plugins and / or APIs until a response can be generated that addresses each query / question / request / process / operation / etc. of the input prompt. Thus, the model can rely not only on its own knowledge obtained from training on large datasets but also on the expertise or optimized nature of one or more external resources (e.g., APIs, plugins, etc.).

[0478] In some embodiments, multiple language models (e.g., LLM / VLM / MMLM / etc., multiple instances of the same language model, and / or multiple prompts provided to the same language model or an instance of the same language model) can be implemented, executed, or accessed (e.g., using one or more plugins, user interfaces, APIs, databases, data stores, repositories, etc.) to provide outputs in response to the same query or in response to separate portions of a query. In at least one embodiment, the same input query and prompt (e.g., a set of constraints, condition generators, etc.) can be provided to multiple language models (e.g., language models with different architectures, language models trained on different (e.g., updated) data corpora). In one or more embodiments, the language models can be different versions of the same underlying model. In one or more embodiments, at least one language model can be instantiated as multiple agents, e.g., more than one prompt can be provided to constrain, guide, or otherwise influence the style, content, or character, etc. of the output provided. In one or more example non-limiting embodiments, the same language model can be required to provide outputs corresponding to different roles, perspectives, characters, or having different knowledge bases, etc. (as defined by the provided prompts).

[0479] In any such embodiment, the outputs of two or more (e.g., each) language models, two or more versions of at least one language model, two or more instantiation agents of at least one language model, and / or two or more prompts provided to at least one language model can be further processed, e.g., aggregated, compared, or filtered, or used to determine (and provide) a consensus response. In one or more embodiments, the output from one language model (or version, instance, or agent) can be provided as input to another language model for further processing and / or validation. In one or more embodiments, a language model can be required to generate or otherwise obtain an output with respect to input source material, where the output is associated with the input source material. Such an association can include, e.g., generating an embedding (e.g., as metadata) within the input source text or image, a caption, or a text portion. In one or more embodiments, the output of a language model can be used to determine the validity of the input source material for further processing or inclusion in a dataset. For example, a language model can be used to evaluate the presence (or absence) of a target word in a text portion or the presence (or absence) of an object in an image, where the text or image is annotated to indicate such presence (or its absence). Alternatively, the determination from the language model can be used to determine whether the source material should be included in a curated dataset, such as but not limited to this.

[0480] Figure 28A is a block diagram of an example generative language model system 2800 suitable for implementing at least some embodiments of the present disclosure. In Figure 28A the example shown, the generative language model system 2800 includes a Retrieval-Augmented Generation (RAG) component 2892, an input processor 2805, a tokenizer 2810, an embedding component 2820, a plug-in / API 2895, and a generative language model (LM) 2830 (which can include an LLM, a VLM, a multimodal LM, etc.).

[0481] At a high level, the input processor 2805 can receive an input 2801 that includes text and / or other types of input data (e.g., audio data, video data, image data, sensor data (e.g., LiDAR, RADAR, ultrasonic, etc.), 3D design data, CAD data, Universal Scene Description (USD) data (e.g., OpenUSD, etc.), depending on the architecture of the generative LM 2830 (e.g., LLM / VLM / MMLM / etc.). In some embodiments, the input 2801 includes plain text in the form of one or more sentences, paragraphs, and / or documents. Additionally or alternatively, the input 2801 can include sequences of numbers, pre-computed embeddings (e.g., word or sentence embeddings), and / or structured data (e.g., table format, JSON, or XML). In some implementations where the generative LM 2830 is capable of processing multimodal inputs, the input 2801 can combine text (or text can be omitted) with image data, audio data, video data, design data, USD data, and / or other types of input data (e.g., but not limited to the data described herein). Taking the original input text as an example, the input processor 2805 can prepare the original input text in various ways. For example, the input processor 2805 can perform various types of text filtering to remove noise from the relevant text content (e.g., special characters, punctuation marks, HTML tags, stopwords, parts of images, parts of audio, etc.). In an example involving stopwords (common words that often have little semantic meaning), the input processor 2805 can remove stopwords to reduce noise and focus the generative LM 2830 on more meaningful content. The input processor 2805 can apply text normalization, e.g., by converting all characters to lowercase, removing accent marks, and / or handling special cases (such as abbreviations or contractions) to ensure consistency. These are just a few examples, and other types of input processing can be applied.

[0482] In some embodiments, the RAG component 2892 (which can include one or more RAG models, and / or it can use the generative LM 2830 itself to perform) can be used to retrieve additional information to be used as part of the input 2801 or the prompt. RAG can be used to enhance the input to the LLM / VLM / MMLM / etc. with external knowledge so that the answer to a particular question or query or request is more relevant, e.g., in cases where specific knowledge is required. The RAG component 2892 can obtain this additional information from one or more external sources (e.g., underlying information such as underlying text / images / videos / audio / USD / CAD / etc.), and then it can feed this together with the prompt to the LLM / VLM / MMLM / etc. to improve the accuracy of the model's response or output.

[0483] For example, in some embodiments, in addition to the data retrieved using the RAG component 2892, a query or model input (e.g., a question, a request, etc.) can be used to generate the input 2801. In some embodiments, the input processor 2805 can analyze the input 2801 and communicate with the RAG component 2892 (or in an embodiment, the RAG component 2892 can be part of the input processor 2805) to identify relevant text and / or other data to provide it to the generative LM 2830 as additional context or an information source from which a response, answer, or output 2890 is typically identified. For example, when the input indicates that the user is interested in the required tire pressure for a specific make and model of vehicle, the RAG component 2892 can use a RAG model to perform a vector search, e.g., in an embedding space, to retrieve tire pressure information or its corresponding text from a digital (embedded) version of the user manual for that specific vehicle make and model. Similarly, when the user accesses the chatbot related to a specific product sale or service again, the RAG component 2892 can retrieve the previously stored conversation history (or at least a summary thereof) and provide the previous conversation history along with the current query / request as part of the input 2801 to the generative LM 2830.

[0484] The RAG component 2892 can use various RAG techniques. For example, a naive RAG ( RAG) can be used, where documents are indexed, chunked, and applied to an embedding model to generate embeddings corresponding to the chunks. The user query can also be applied to the embedding model and / or another embedding model of the RAG component 2892, and the embeddings of the chunks can be compared with the embedding of the query to identify the most similar / most relevant embeddings, which can be provided to the generative LM 2830 to generate an output.

[0485] In some embodiments, more advanced RAG techniques can be used. For example, before passing the chunks to the embedding model, the chunks can undergo a pre-retrieval process (e.g., routing, rewriting, metadata analysis, expansion, etc.). Additionally, a post-retrieval process (e.g., re-ranking, prompt compression, etc.) can be performed on the output of the embedding model before generating the final embedding, which is then used for comparison with the input query.

[0486] As a further example, modular RAG techniques can be used, such as techniques similar to naive RAG and / or advanced RAG, but also including features such as hybrid search, recursive retrieval and query engines, StepBack methods, subqueries, and hypothetical document embeddings.

[0487] As another example, a Graph RAG can use a knowledge graph as a source of context or factual information. The Graph RAG can be implemented using a graph database as a source of context information to be sent to an LLM / VLM / MMLM / etc. Instead of (or in addition to) providing the model with chunks of data extracted from a larger document (which can lead to lack of context, factual correctness, language accuracy, etc.), the Graph RAG can also provide structured entity information to an LLM / VLM / MMLM / etc. by combining structured entity text descriptions with their many attributes and relationships, enabling the model to gain deeper insights. In implementing the Graph RAG, the systems and methods described herein use a graph as a content store and extract relevant document chunks and require the LLM / VLM / MMLM / etc. to use them for answering. In such an embodiment, the knowledge graph can contain relevant text content and metadata about the knowledge graph and can also be integrated with a vector database. In some embodiments, the Graph RAG can use the graph as a subject matter expert, where descriptions of concepts and entities relevant to a query / prompt can be extracted and passed to the model as semantic context. These descriptions can include relationships between concepts. In other examples, the graph can be used as a database, where a portion of the query / prompt can be mapped to a graph query, the graph query can be executed, and the LLM / VLM / MMLM / etc. can summarize the results. In such an example, the graph can store relevant factual information and can use queries (natural language queries) and entity linking to a graph query tool (NL to graph query tool). In some embodiments, the Graph RAG (e.g., using a graph database) can be combined with a standard (e.g., vector database) RAG and / or other RAG types to benefit from multiple approaches.

[0488] In any embodiment, the RAG component 2892 can implement a plug-in, API, user interface, and / or other functionality to perform RAG. For example, an LLM / VLM / MMLM / etc. can use a Graph RAG plug-in to run queries against a knowledge graph to extract relevant information to feed into the model and can use a standard or vector RAG plug-in to run queries against a vector database. For example, the graph database can interact with the REST interface of the plug-in so that the graph database can be decoupled from the vector database and / or the embedding model.

[0489] Tokenizer 2810 can split (e.g., processed) text data into smaller units (tokens) for subsequent analysis and processing. Depending on the implementation, tokens can represent individual words, sub-words, characters, parts of audio / video / images / etc. Word-based tokenization divides the text into individual words, treating each word as a separate token. Sub-word tokenization breaks words into smaller meaningful units (e.g., prefixes, suffixes, stems), enabling the generative LM 2830 to understand morphological variations and handle out-of-vocabulary words more effectively. Character-based tokenization represents each character as a separate token, enabling the generative LM 2830 to process text at a fine-grained level. The choice of tokenization strategy can depend on factors such as the language being processed, the task at hand, and / or the characteristics of the training dataset. Thus, the tokenizer 2810 can transform (e.g., processed) text into a structured format according to the tokenization scheme implemented in a particular embodiment.

[0490] The embedding component 2820 can transform discrete tokens into a (e.g., dense, continuous vector) representation of semantic meaning using any known embedding technique. For example, the embedding component 2820 can use pre-trained word embeddings (e.g., Word2Vec, GloVe, or FastText), one-hot encoding, term frequency-inverse document frequency (TF-IDF) encoding, one or more embedding layers of a neural network, and / or others.

[0491] In some implementations where the input 2801 includes image data / video data / etc., the input processor 2801 may resize the data to a standard size compatible with the format of the corresponding input channel and / or may normalize the pixel values to a common range (e.g., 0 to 1) to ensure a consistent representation, and the embedding component 2820 may encode the image data using any known technique (e.g., using one or more convolutional neural networks (CNNs) to extract visual features). In some implementations where the input 2801 includes audio data, the input processor 2801 may resample the audio file to a consistent sampling rate for unified processing, and the embedding component 2820 may use any known technique to extract and encode audio features, e.g., in the form of a spectrogram (e.g., Mel spectrogram). In some implementations where the input 2801 includes video data, the input processor 2801 may extract frames or apply resizing to the extracted frames, and the embedding component 2820 may extract features such as optical flow embeddings or video embeddings and / or may encode temporal information or frame sequences. In some implementations where the input 2801 includes multimodal data, the embedding component 2820 may use techniques such as early fusion (concatenation), late fusion (sequential processing), attention-based fusion (e.g., self-attention, cross-attention), etc. to fuse the representations of different types of data (e.g., text, image, audio, USD, video, design, etc.).

[0492] The generative LM 2830 and / or other components of the generative LM system 2800 may use different types of neural network architectures depending on the implementation. For example, a Transformer-based architecture (e.g., the architecture used in models such as GPT) may be implemented, and it may include a self-attention mechanism that weights the importance of different words or tokens in the input sequence and / or a feed-forward network that processes the output of the self-attention layer, which applies a non-linear transformation to the input representation and extracts higher-level features. Some non-limiting example architectures include Transformer (e.g., encoder-decoder, decoder-only, multimodal), RNN, LSTM, fusion models, diffusion models, cross-modal embedding models that learn a joint embedding space, graph neural networks (GNNs), hybrid architectures that combine different types of architecture adversarial networks (such as generative adversarial networks or GANs or adversarial autoencoders (AAEs) for joint distribution learning), etc. Thus, depending on the implementation and architecture, the embedding component 2820 may apply the encoded representation of the input 2801 to the generative LM 2830, and the generative LM 2830 may process the encoded representation of the input 2801 to generate an output 2890, which may include response text and / or other types of data.

[0493] As described herein, in some embodiments, the generative LM 2830 can be configured to access or use (or be able to access or use) a plugin / API 2895 (which can include one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc.). For example, for certain tasks or operations that the generative LM 2830 is not ideally suited for, the model can have instructions (e.g., as a result of training, and / or based on instructions in a given prompt, e.g., retrieved using the RAG component 2892) to access one or more plugins / APIs 2895 (e.g., third-party plugins) to assist in processing the current input. In such an example, when at least a portion of the prompt is related to a restaurant or the weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs), send at least a portion of the prompt related to the specific plugin / API 2895 to the plugin / API 2895, the plugin / API 2895 can process the information and return an answer to the generative LM 2830, and the generative LM 2830 can use the response to generate an output 2890. This process can be repeated (e.g., recursively) any number of times of iteration and repeated with any number of plugins / APIs 2895 until an output 2890 can be generated that solves each query / question / request / process / operation / etc. from the input 2801. Thus, the model can rely not only on its own knowledge obtained from training on large datasets and / or data retrieved using the RAG component 2892, but also on the expertise or optimized nature of one or more external resources (e.g., the plugin / API 2895).

[0494] Figure 28B is a block diagram of an example implementation, where the generative LM 2830 includes a transformer encoder-decoder. For example, assume the input text (e.g., "Who discovered gravity") is tokenized (e.g., by Figure 28A the tokenizer 2810) into tokens such as words, and each token is encoded (e.g., by Figure 28A the embedding component 2820) into a corresponding embedding (e.g., of size 512). Since these token embeddings generally do not represent the position of the tokens in the input sequence, any known technique can be used to add positional encodings to each token embedding to encode the sequential relationship and context of the tokens in the input sequence. Thus, the (e.g., resulting) embeddings can be applied to one or more encoders 2835 of the generative LM 2830.

[0495] In an example implementation, the encoder 2835 forms an encoder stack, where each encoder includes a self-attention layer and a feed-forward network. In an example Transformer architecture, each token (e.g., word) flows through a separate path. Thus, each encoder can receive a sequence of vectors, pass each vector through the self-attention layer, then through the feed-forward network, and then pass it up to the next encoder in the stack. Any known self-attention technique can be used. For example, to compute the self-attention scores for each token (word), query vectors, key vectors, and value vectors can be created for each token, and the self-attention scores for token pairs can be computed by taking the dot product of the query vector with the corresponding key vector, normalizing the resulting scores, multiplying by the corresponding value vector, and summing the weighted value vectors. The encoder can apply multi-head attention, where the attention mechanism is applied multiple times in parallel with different learned weight matrices. Any number of encoders can be cascaded to generate a context vector that encodes the input. The attention projection layer 2840 can transform the context vector into an attention vector (keys and values) for the decoder 2845.

[0496] In an example implementation, the decoder 2845 forms a decoder stack, where each decoder includes a self-attention layer, an encoder-decoder self-attention layer that uses the attention vector (keys and values) from the encoder to attend to relevant parts of the input sequence, and a feed-forward network. Similar to the encoder 2835, in an example Transformer architecture, each token (e.g., word) flows through a separate path in the decoder 2845. During the first pass, the decoder 2845, the classifier 2850, and the generation mechanism 2855 can generate a first token, and the generation mechanism 2855 can apply the generated token as input during the second pass. This process can be repeated iteratively, generating tokens (e.g., words) in sequence and adding them to the output of the previous pass, and applying the token embeddings of the composite sequence with positional encoding as input to the decoder 2845 in subsequent passes, generating one token at a time (referred to as autoregressive) until a symbol or token representing the end of the response is predicted. In each decoder, the self-attention layer is typically restricted to attending only to previous positions in the output sequence by applying a masking technique (e.g., setting future positions to negative infinity) before the softmax operation. In an example implementation, the operation of the encoder-decoder attention layer is similar to the operation of the (e.g., multi-head) self-attention in the encoder 2835, except that it creates its queries from the layer below it and obtains the keys and values (e.g., matrices) from the output of the encoder 2835.

[0497] Accordingly, the decoder 2845 can output some decoded (e.g., vector) representations of the input applied during a particular pass. The classifier 2850 can include a multi-class classifier that includes one or more neural network layers and a softmax operation that converts logits to probabilities, and the neural network layers project the decoded (e.g., vector) representations onto corresponding dimensions (e.g., one dimension for each supported word or token in the output vocabulary). Thus, the generation mechanism 2855 can select or sample words or tokens based on the corresponding predicted probabilities (e.g., select the word with the highest predicted probability) and append them to the output of the previous pass, thereby generating each word or token in sequence. The generation mechanism 2855 can repeat this process, triggering successive decoder inputs and corresponding predictions until a symbol or token representing the end of the response is selected or sampled, at which point the generation mechanism 2855 can output the generated response.

[0498] Figure 28C is a block diagram of an example implementation, where the generative LM 2830 includes a decoder-only transformer architecture. For example, Figure 28C the decoder 2860 of Figure 28B can operate similarly to the decoder 2845 of Figure 28C except that each decoder 2860 of Figure 28B omits the encoder-decoder self-attention layer (since there is no encoder in this implementation). Accordingly, the decoder 2860 can form a decoder stack, where each decoder includes a self-attention layer and a feed-forward network. Additionally, a symbol or token representing the end of the input sequence (or the start of the output sequence) can be appended to the input sequence instead of encoding the input sequence, and the resulting sequence (e.g., the corresponding embedding with positional encoding) can be applied to the decoder 2860. Similar to Figure 28B the decoder 2845 of Figure 28B each token (e.g., word) can flow through a separate path in the decoder 2860, and the decoder 2860, classifier 2865, and generation mechanism 2870 can use autoregression to generate one token at a time in sequence until a symbol or token representing the end of the response is predicted. The classifier 2865 and generation mechanism 2870 can operate similarly to Figure 28B the classifier 2850 and generation mechanism 2855 of

[0499] Example content streaming system

[0500] Now refer to Figure 29 ,Figure 29 FIG. 2900 is an example system diagram of a content streaming system 2900 in accordance with some embodiments of the present disclosure. Figure 29 It includes an application server 2902 (which may include components, features, and / or functions similar to those of the Figure 30 example computing device 3000), a client device 2904 (which may include components, features, and / or functions similar to those of the Figure 30 example computing device 3000), and a network 2906 (which may be similar to the networks described herein). In some embodiments of the present disclosure, the system 2900 may support application sessions corresponding to game streaming applications (e.g., NVIDIA GeFORCE NOW), remote desktop applications, simulation applications (e.g., autonomous or semi-autonomous vehicle simulation), computer-aided design (CAD) applications, virtual reality (VR) and / or augmented reality (AR) streaming applications, deep learning applications, and / or other application types.

[0501] In the system 2900, for an application session, one or more client devices 2904 may receive only input data in response to input to one or more input devices, transmit the input data to one or more application servers 2902, receive encoded display data from one or more application servers 2902, and display the display data on a display 2924. Thus, computationally more intensive computing and processing may be offloaded to one or more application servers 2902 (e.g., rendering of the graphical output of an application session that may be performed by one or more GPUs of one or more application servers 2902, such as one or more game servers - particularly ray or path tracing). In other words, the application session is streamed from one or more application servers 2902 to one or more client devices 2904, thereby reducing the requirements for graphics processing and rendering on one or more client devices 2904.

[0502] For example, with respect to the instantiation of an application session, the client device 2904 may display frames of the application session on the display 2924 based on receiving display data from one or more application servers 2902. The client device 2904 may receive an input from one of the one or more input devices and generate input data in response. The client device 2904 may send the input data to the application server 2902 via the communication interface 2920 and via the network 2906 (e.g., the Internet), and the application server 2902 may receive the input data via the communication interface 2918. The CPU may receive the input data, process the input data, and transmit data to the GPU that causes the GPU to generate a rendering of the application session. For example, the input data may represent a user's movement of a character, firing of a weapon, reloading, passing a ball, steering a vehicle, etc. in a gaming session of a gaming application. The rendering component 2912 may render the application session (e.g., representing the result of the input data), and the render capture component 2914 may capture the rendering of the application session as display data (e.g., as image data capturing frames of the rendering of the application session). The rendering of the application session may include lighting and / or shadow effects computed using ray or path tracing by one or more parallel processing units (such as GPUs) of the one or more application servers 2902, and the one or more parallel processing units may further use one or more dedicated hardware accelerators or processing cores to perform ray or path tracing techniques. In some embodiments, one or more virtual machines (VMs) - e.g., including one or more virtual components such as vGPUs, vCPUs, etc. - may be used by the application server 2902 to support the application session. The encoder 2916 may then encode the display data to produce encoded display data, and the encoded display data may be sent to the client device 2904 via the communication interface 2918 over the network 2906. The client device 2904 may receive the encoded display data via the communication interface 2920, and the decoder 2922 may decode the encoded display data to produce display data. The client device 2904 may then display the display data via the display 2924.

[0503] Example computing device

[0504] Figure 30FIG. 3000 is a block diagram of an example computing device 3000 suitable for implementing some embodiments of the present disclosure. Computing device 3000 may include an interconnect system 3002 that directly or indirectly couples the following devices: a memory 3004, one or more central processing units (CPUs) 3006, one or more graphics processing units (GPUs) 3008, a communication interface 3010, input / output (I / O) ports 3012, input / output components 3014, a power supply 3016, one or more presentation components 3018 (e.g., one or more displays), and one or more logic units 3020. In at least one embodiment, one or more computing devices 3000 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). For non-limiting examples, one or more of the GPUs 3008 may include one or more vGPUs, one or more of the CPUs 3006 may include one or more vCPUs, and / or one or more of the logic units 3020 may include one or more virtual logic units. Accordingly, one or more computing devices 3000 may include discrete components (e.g., a full GPU dedicated to computing device 3000), virtual components (e.g., a portion of a GPU dedicated to computing device 3000), or a combination thereof.

[0505] Although Figure 30 each of the boxes of FIG. 3000 is shown as being connected to the lines via the interconnect system 3002, this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 3018 (such as a display device) may be considered an I / O component 3014 (e.g., if the display is a touch screen). As another example, the CPU 3006 and / or the GPU 3008 may include memory (e.g., in addition to the memory of the GPU 3008, the CPU 3006, and / or other components, the memory 3004 may represent a storage device). In other words, Figure 30 the computing devices of FIG. 3000 are merely illustrative. No distinction is made between categories such as "workstation", "server", "laptop computer", "desktop computer", "tablet computer", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", and / or other device or system types, as all are contemplated within the scope of Figure 30 the computing devices of FIG. 3000.

[0506] The interconnect system 3002 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 3002 may include one or more types of buses or links, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. For example, the CPU 3006 may be directly connected to the memory 3004. Further, the CPU 3006 may be directly connected to the GPU 3008. In cases where there are direct connections or point-to-point connections between components, the interconnect system 3002 may include a PCIe link to perform the connection. In these examples, the computing device 3000 need not include a PCI bus.

[0507] The memory 3004 may include any of a variety of computer-readable media. The computer-readable media may be any available media that can be accessed by the computing device 3000. The computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, the computer-readable media may include computer storage media and communication media.

[0508] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or o...

Claims

1. One or more processors, including processing circuitry, the processing circuitry being configured to: receiving, by one or more action servers handling one or more overlays of visual content that supplement one or more conversations with an interactive agent, one or more events representing one or more visual content actions classified using an interaction classification scheme and indicating one or more updates to the one or more overlays in one or more graphical user interfaces (GUIs); generating, by the one or more action servers, one or more visual layouts representing the one or more updates specified by the one or more events; as well as A rendering of the one or more visual layouts in a scene is caused to be presented.

2. One or more processors according to claim 1, wherein the processing circuit is further used to: select one or more template visual layouts based at least on one or more content blocks specified by the one or more events, and generate the one or more visual layouts based at least on filling one or more placeholders in the one or more template visual layouts with the content specified by the one or more events.

3. One or more processors according to claim 1, wherein the processing circuit is also used to: convert the one or more events into one or more modular graphical user interface configurations, and the one or more modular graphical user interface configurations specify one or more content blocks corresponding to one or more fields specified by the one or more events.

4. One or more processors according to claim 1, wherein the one or more events include one or more fields, and the one or more fields identify supported action types for classifying the one or more first visual content actions, the status of the one or more first visual content actions, and the representation of the indicated visual content.

5. One or more processors according to claim 1, wherein the one or more visual content actions include a visual information scene action that indicates a visualization of information about a topic associated with the one or more conversations.

6. One or more processors of claim 1, wherein the one or more visual content actions include a visual selection action that indicates a visualization of one or more selections associated with the one or more conversations.

7. One or more processors according to claim 1, wherein the one or more visual content actions include a visual form action, which indicates a visualization of one or more form fields, and the one or more form fields accept one or more inputs associated with the one or more dialogs.

8. The one or more processors of claim 1, wherein the one or more events include one or more fields specifying text generated using one or more large language models.

9. One or more processors according to claim 1, wherein the one or more events include one or more fields indicating that the one or more images were retrieved or generated based at least on one or more natural language descriptions of the one or more images.

10. The one or more processors of claim 1, wherein the processing circuit is further configured to direct inclusion of the one or more visual layouts in a stack of the one or more overlays.

11. The one or more processors of claim 1 , wherein the one or more processors are included in at least one of: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing simulation operations; Systems for performing digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; Systems for performing deep learning operations; Systems for performing remote operations; A system for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; A system implementing one or more language models; A system implementing one or more large language models (LLMs); A system implementing one or more visual language models (VLMs); A system implementing one or more multimodal language models; Systems for generating synthetic data; Systems for generating synthetic data using AI; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.

12. A system comprising one or more processors, the one or more processors being configured to: generate one or more visual layouts by one or more action servers, the one or more action servers handling one or more overlays of visual content that supplement one or more conversations with an interactive agent, the one or more visual layouts representing one or more updates to the one or more overlays specified by one or more events, the one or more events representing one or more visual content actions classified using an interaction classification scheme.

13. A system according to claim 12, wherein the one or more processors are further used to: select one or more template visual layouts based at least on one or more content blocks specified by the one or more events, and generate the one or more visual layouts based at least on filling one or more placeholders in the one or more template visual layouts with the content specified by the one or more events.

14. A system according to claim 12, wherein the one or more processors are further used to: convert the one or more events into one or more modular graphical user interface configurations, wherein the one or more modular graphical user interface configurations specify one or more content blocks corresponding to one or more fields specified by the one or more events.

15. A system according to claim 12, wherein the one or more events include one or more fields that identify supported action types for classifying the one or more visual content actions, the status of the one or more visual content actions, and a representation of the indicated visual content.

16. The system of claim 12, wherein the one or more visual content actions include a visual information scene action that indicates a visualization of information about a topic associated with the one or more conversations.

17. The system of claim 12, wherein the one or more visual content actions include a visual selection action that indicates a visualization of one or more selections associated with the one or more conversations.

18. The system of claim 12, wherein the system is included in at least one of the following: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing simulation operations; Systems for performing digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; Systems for performing deep learning operations; Systems for performing remote operations; A system for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; A system implementing one or more language models; A system implementing one or more large language models (LLMs); A system implementing one or more visual language models (VLMs); A system implementing one or more multimodal language models; Systems for generating synthetic data; Systems for generating synthetic data using AI; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.

19. A method comprising: receiving one or more events representing one or more visual content actions classified using an interaction classification scheme and indicating one or more updates to one or more overlays of visual content that supplement one or more conversations with an interactive agent; as well as The one or more events are converted into one or more visual layouts representing the one or more updates specified by the one or more events.

20. The method of claim 19, wherein the method is performed by at least one of: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing simulation operations; Systems for performing digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; Systems for performing deep learning operations; Systems for performing remote operations; A system for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; A system implementing one or more language models; A system implementing one or more large language models (LLMs); A system implementing one or more visual language models (VLMs); A system implementing one or more multimodal language models; Systems for generating synthetic data; Systems for generating synthetic data using AI; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.