Virtual ai representative
The system addresses challenges in NLP models by using a state machine and knowledge base to facilitate directed conversations, ensuring coherent and contextually accurate interactions with seamless transitions between conversation and presentation, and rapid response generation for a natural flow.
Patent Information
- Application Number
- PCT/IB2024/062070
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-02
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-05
AI Technical Summary
Standard natural language processing (NLP) models are not optimized for long, purposeful, real-time, interactive dialogues, leading to contextually inaccurate or incoherent responses. Additionally, maintaining a seamless transition between conversation and interactive visual presentation, especially when conditional on dialogue flow, is complex. Harmonizing multiple threads to monitor user engagement, presence, or intent, and ensuring rapid response generation to maintain a natural conversation, are further challenges.
A system and method utilizing a state machine to control directed conversations, integrating a knowledge base, and providing an interactive presentation. The system includes a controller unit that manages event queues, a state manager unit acting as a dynamic state machine, and a knowledge base unit that provides personalized and contextual information. This setup ensures coherent and contextually accurate interactions by synchronizing conversation flow with interactive presentation.
The solution enables effective facilitation of directed conversations that are contextually accurate and coherent, ensuring seamless transitions between conversation and interactive visual presentation. It maintains a natural conversation flow by generating responses within a fraction of a second, improving user engagement and overall interaction quality.
Smart Images

Figure IB2024062070_05062025_PF_FP_ABST
Abstract
Description
VIRTUAL Al REPRESENTATIVEBACKGROUND
[0001] The present invention relates artificial intelligent (Al) assistants that can be adapted to perform a directed function, and more specifically, to conversationally interact according to the directed function.SUMMARY
[0002] According to one embodiment of the invention, there is provided a method for facilitating a directed conversation according a criteria between an artificially-intelligent (Al) agent and an audience. A state machine is received by an Al system containing the Al agent. The state machine controls the directed conversation. A knowledge base for the directed conversation is ingested by the system. An interactive presentation for the directed conversation as indicated by the state machine is provided by the Al agent.
[0003] According to one embodiment of the invention, there is provided an information handling system that implements the steps of the method for facilitating a directed conversation according to a criteria between an artificially-intelligent (Al) agent and an audience.
[0004] According to one embodiment of the invention, there is provided a computer program product running program instructions executable on a processing circuit to cause the processing circuit to perform the steps facilitating a directed conversation according to a criteria between an artificially-intelligent (Al) agent and an audience.
[0005] The foregoing is a summary and thus contains, by necessity, simplifications, generalizations, and omissions of detail; consequently, those skilled in the art will appreciate that the summary is illustrative only and is not intended to be in any waylimiting. Other aspects, inventive features, and advantages of the present invention will be apparent in the non-limiting detailed description set forth below.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The accompanying drawings, which are included to provide a further understanding of the inventive concepts and are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the inventive concepts, and, together with the description, serve to explain the principles of the inventive concepts.
[0007] FIG. 1 illustrates virtual Al representative core architecture.
[0008] FIG. 2 illustrates the process flow for virtual Al representative core operative steps.
[0009] FIG. 3 illustrates the virtual Al representative architecture including user dashboard, data storage and virtual Al representative fleet manager and core.
[0010] FIG. 4 illustrates an exemplary website that employs a virtual Al representative as a sales agent to present the product to interested participants.
[0011] FIG. 5 illustrates a participant requesting for initiating a virtual Al presentation session.
[0012] FIG. 6 illustrates a participant joining a meeting session after requesting one.
[0013] FIG. 7 illustrates a virtual Al representative starting a meeting session.
[0014] FIG. 8 illustrates the user dashboard for a product owner to define the specifications of the virtual Al representative.
[0015] FIG. 9 illustrates an exemplary hardware architecture required to implement the present invention.DETAILED DESCRIPTION OF THE INVENTION
[0016] In the following description, for the purposes of explanation, numerous specific details are set forth to provide a thorough understanding of various exemplary embodiments. It is apparent, however, that various exemplary embodiments may be practiced without these specific details or with one or more equivalent embodiments.
[0017] In the accompanying figures, the size and relative sizes of elements may be exaggerated for clarity and descriptive purposes.
[0018] The terminology used herein is for the purpose of describing particular embodiments and is not intended to be limiting. As used herein, the singular forms,
[0019] “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Moreover, the terms “comprises,” “comprising,” “includes,” and / or “including,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, components, and / or groups thereof, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0020] Implementing a virtual Al representative may face a range of technical challenges that require sophisticated solutions. One important challenge is that standard natural language processing (NLP) models may not be optimized for long, purposeful, real-time, interactive dialogues and might produce responses that are not contextually accurate or coherent with the flow and purpose of the conversation. Another challenge is maintaining a seamless transition between the conversation and the interactive visual presentation, especially when the interactive presentation is conditional on the dialogue flow. Multiple threads are required to monitor various aspects of the conversation, such as user engagement, presence, or intent. Harmonizing these threads to produce a coherent interaction that follows the flow ofthe conversation is not straightforward. Another complexity is the response rate; to maintain a natural conversation, the system needs to generate responses within a fraction of a second.
[0021] FIG. 1 shows embodiments of the present invention that includes a system and method for an artificially intelligent virtual representative. Elements shown in FIG.1 are in the form of software. As shown in FIG. 1 , the system of the present invention includes the following components:
[0022] Controller unit 100 serves as the central processing and orchestration unit in the system. It is the brain behind the operations, ensuring synchronization between different threads and processes. Through a series of event queues, controller unit 100 communicates with various components, responding to and processing events such as user interactions, system updates, and audio inputs. An event queue is a data structure that operates based on the First-In-First-Out (FIFO) principle. The event queue is used to store and manage events or messages that need to be processed. In multithreaded applications such as the present invention, an event queue helps in achieving thread-safe communication between threads.
[0023] User input unit 102 is responsible for receiving and processing user voice inputs that come from the meeting application or medium. Transcriber unit 118 resides within user input unit 102. The primary role of transcriber unit 118 is to convert the captured audio data into textual format, essentially "transcribing" spoken words into readable text. Leveraging available advanced speech recognition algorithms, transcriber unit 118 analyzes the audio data. Controller unit 100 messages user input unit 102 at the beginning of the conversation to mark the start of the conversation.State manager unit 106 functions as a dynamic state machine, meticulously tracking and guiding the flow of conversation. The state manager utilizes a range ofpredefined states to facilitate a structured yet adaptable interaction, catering to a variety of conversational objectives. Each state within this system is defined by unique attributes including a unique identifier, directives on how to respond in each state, optional associated visual content, instructions for the next course of action (transiting to the next state and the conditions for the transit), and if the state is a “wait for response” state for user to provide a response or a “move forward” state that does not wait for the user’s input. When a message is received and transcribed by transcriber, transcriber assigns a unique number to it, so the message looks like this {identifier :2345, message: "how can your product help us?"}. This identifier is used throughout the life cycle of the message, for handling interruption or speeding up the response process.
[0024] State manager unit 106 includes two groups of states: user-defined states and system-defined states. System-defined states include “audio connection,” “first state,” "hold," "interrupt," and “tangent.” Any other states defined by the user to customize the virtual Al representative for their specific use and to ensure a fluid and intuitive interaction are called user-defined states. Controller unit 100 waits in “audio connection” state until it receives a message from the user at the beginning of the meeting to transit to the “first state.” All user-defined states can transit to the “interrupt” state if the user interrupts the virtual Al representative while presenting; reverting back post-interruption. Queries deviating from the meeting's flow trigger a transition to the “tangent” state, allowing the virtual Al representative to address off-topic inquiries. A user request for a pause shifts the state to “hold.” Each state associates with corresponding visual content on the meeting platform, which pauses when the state transitions and resumes when back in that state again. Transitions between states are guided by conditions that act as triggers, dictating the requirements for movement andidentifying the destination state. LLM interactor-conversation unit 108 decides if the transitions conditions are met and determines the state of conversation in each conversation cycle, the conversation cycle consists of a back and forth between the participant and the virtual Al representative.
[0025] State manager unit 106 can be adjusted to act as a persona with different set of states. For instance, the virtual Al representative presented in this patent can emulate a virtual Al sales agent when provided with suitable set of states and a product knowledge base to provide contextual information for knowledge base unit 126. States dictates how the agent navigates the presentation while demonstrating the product and the knowledge base provides the agent with prior information about the product. The states for this specific example are included in Table 1. Each state has a name, instruction, transition condition, the next state, and the action the agent must take after delivering the instruction.Table 1. States for the virtual Al representative to emulate a virtual sales agent
[0026] The user-defined states for this specific example are Agenda, Product, and Final. User-defined states provided in Table 1 can be more than the ones presented here to refine the conversation and to provide more instruction to the Al sales agent. System-defined states are hold, tangent, interruption, audio connection, and first state. At the beginning of the conversation, the Al agent is in state audio-connection. When the Al agent receives a participant's voice, the Al agent transits to the first-state inwhich it welcomes the participant. The agent transits to the agenda state in which it outlines the agenda for the meeting. When there is a message from the participants, controller unit 100 sends the message to LLM interactor-conversation unit 108 and LLM interactor-conversation unit 108 answers the message and determines the state in which the Al agent resides.
[0027] Arranging the set of states as in T able 2 can tailor the virtual Al representative to emulate an instructor. A course curriculum and related information on the topic of interest is provided to the virtual Al representative via knowledge base unit 126. User- defined states provided in Table 2 can be more than the ones presented here to refine the conversation and to provide more instruction to the virtual Al instructor.Table 2. States for the virtual Al representative to emulate a virtual instructor
[0028] Arranging the set of states as in T able 3 can tailor the virtual Al representative to emulate a healthcare provider. Related medical knowledge on the topic of specialty is provided to the virtual Al representative via knowledge base unit 126. User-defined states provided in Table 3 can be more than the ones presented here to refine the conversation and to provide more instruction to the virtual Al healthcare provider.Table 3. States for the virtual Al representative to emulate a virtual healthcare provider
[0029] The set of states in Table 4 can be used for the virtual Al representative to emulate a customer service representative. User-defined states provided in Table 4 can be more than the ones presented here to refine the conversation.Table 4. States for the virtual Al representative to emulate a virtual customer service representative
[0030] The set of states in Table 5 can be used for the virtual Al representative to emulate a virtual advisory service provider (i.e. a financial service advisor). User- defined states provided in Table 5 can be more than the ones presented here to refine the conversation.Table 5. States for the virtual representative to emulate a virtual advisory service provider
[0031] The set of states in Table 6 can be used for the virtual Al representative to emulate a virtual recruiter. User-defined states provided in Table 6 can be more than the ones presented here to refine the conversation.
[0032] The current state of the conversation is determined by LLM interactiveconversation unit 108. The progression of the states is not strictly sequential and can follow various paths depending on the input or other conditions. States with associated visual content can deliver relevant visual information or demonstrations throughout the conversation.
[0033] Action controller unit 110 is an integrated system that encompasses three primary components: action recorder unit 112, action player unit 114, and video recorder / player unit 116. Video recorder / player unit 116 records brief video snippets during the initialization of the virtual Al representative instance. These recorded snippets serve as a reservoir of content, ready for playback during presentations. Their deployment is contingent upon the presentation's context and state of the conversation passed by controller unit 100. Action recorder unit 112 meticulously records all events, including mouse clicks and keyboard strokes, capturing their precise timing when defining the virtual Al representative. Additionally, it embeds "merge tags" within these recordings. Such tags allow for real-time adaptability. For example, if a user originally searched for the weather in Vancouver, the embedded merge tag for "Vancouver" can be seamlessly replaced with another city during a later conversation. Action player unit 114 can mold screen activities during an interactive presentation based on the conversation's context, especially when the virtual Al representative is introducing a new product using the merge tags and the pre-recorded videos. In live presentations,action player unit 114 performs two critical roles. Firstly, it ensures that the timing of the playback mirrors the initial recording. Secondly, it actively monitors browser network activities, making real-time adjustments to the event timings. As an example, if a webpage originally took 2 seconds based on the data provided by action recorder unit 112 but requires 5 seconds during a live presentation, action player unit 114 recalibrates the timing of subsequent events.
[0034] Vocalizer unit 138 is an audio processing system, seamlessly integrating three specialized sub-units to deliver optimized voice outputs including audio generator unit 120, audio caching unit 122, audio player unit 124. Audio generator unit 120 generates voice snippets for individual sentences. While several available deep learning models can be employed for this purpose, fine-tuning of the model is required to ensure the fastest response in voice generation. Fine-tuning is done by providing the LLM by some sample conversation scenarios. Audio caching unit 122 serves as a repository, diligently maintaining a database of each vocalized sentence. The primary advantage of this cache is swift access when possible. By storing pre-vocalized sentences, the system dramatically reduces the time required to generate voice snippets for frequently used words or phrases, enhancing overall efficiency and speed. Audio player unit 124 is responsible for the actual playback of the voice snippets. The choice of both the voice format and the playback technology is rooted in their reliability and efficiency. However, the modular nature of vocalizer unit 138 ensures flexibility. If the need arises, alternative technologies and libraries can be integrated to replace the current voice format and playback mechanism.
[0035] Knowledge base unit 126 is a system designed to consolidate, process, and provide information tailored to both the product being presented and the user engaged in the conversation. The main objective of knowledge base unit 126 is to providepersonalization and context for a purposeful conversation. This unit amalgamates three pivotal components: knowledge base encoder unit 128, LLM interactor - user profiler unit 130, and knowledge base 132. Knowledge base 132 acts as a contextual hub. As discussions around the product evolve, knowledge base 132 dynamically provides relevant product-specific information and user-specific recommendations, ensuring that the conversation remains both informed and engaging.
[0036] Knowledge base encoder unit 128 is adept at transforming raw documents into structured, searchable formats. Knowledge base encoder unit 128 employs advanced vectorization techniques to convert documents into a format conducive to rapid searches and retrievals. Subsequent to vectorization, knowledge base encoder unit 128 establishes a database. This reservoir is primed with rich information about the product under discussion, ensuring that the Al virtual representative is equipped with comprehensive product knowledge.
[0037] LLM interactor - user profiler unit 130 gathers insights about the user throughout the presentation's duration, as interactions with the user progress, LLM interactor - user profiler unit 130 assiduously records and updates the background information acquired about the user. This includes preferences, past interactions, queries, feedback, and other pertinent details. This reservoir of insights not only ensures that every engagement with the user is rooted in historical context but also paves the way for more personalized and intuitive future interactions. Beyond cataloging user details, LLM interactor - user profiler unit 130 also holds the responsibility of strategizing and noting down future actions post the user interaction. For instance, if a discussion culminates in the decision to share a contract with the user, this action is duly noted and passed to controller unit 100, which eventually will be passed to LLM interactor - user conversation unit 108. Similarly, commitmentsmade during the conversation, like sharing case studies or further information, are systematically recorded. This proactive approach ensures that every commitment made during an interaction is passed to controller unit 100 for required actions after meetings.
[0038] User conversation encoder unit 134 acts as a reservoir that encodes users' questions and inputs into vectors across all meetings with different participants for a specific instance of virtual Al representative and then uses this reservoir to find similar question and answer sets. Controller unit 100 polls user conversation encoder unit 134 every time a new user message is received. If user conversation encoder unit 134 finds an existing suitable answer to the user message from before, controller unit 100 uses the existing message as a response to the user and skips sending the message to LLM interactor - user conversation unit 108. The main objective of the unit is to improve response time.
[0039] Interrupt and user monitoring unit 136 monitors user presence and interrupts to inform controller unit 100 if there is a need to change the state of the conversation. This unit maintains two event queues: “user_activity_event_queue” and “controller_event_queue.” “user_activity_event_queue” is used by controller unit 100 to inform the Interrupt and user monitoring unit 136 about other interactions using the following events: “final_state_timeout_triggered,” “long_inactivity_timeout_triggered,” “user_inactivity_timeout_triggered,” and “user_response_playback_triggered.” Controller unit 100 uses “user_inactivity_timeout_triggered” message to start a process of checking on the user every 20 seconds and uses “long_inactivity_timeout_triggered” message to end the conversation after 5 minutes if there is no answer. When in the final state, controller unit 100 uses “final_state_timeout_triggered” message to end the conversation aftera period of inactivity from the user to ensure the conversation has ended gracefully. Controller unit 100 uses “user_response_playback_triggered” message to inform interrupt and user monitoring unit 136 that the user is done talking and now we are waiting on the Al response from LLM interactor - user conversation unit 108.
[0040] Application Programming Interface (API) Server Unit 140, as embodied in the present invention, serves as an interface for the virtual Al representative, designed to handle synchronous communication events and audio data transmissions. The primary objective of this unit is to efficiently manage a series of events, such as participants joining or leaving a virtual meeting platform (meeting application unit 142), or any status changes within the meeting through its ' / webhook' endpoint. Depending on the nature of the event received, API server unit 140 triggers an appropriate function, placing the event details into an event queue for subsequent handling by controller unit 100. Another salient feature of API server unit 140 is its capability to handle raw audio data from virtual meetings. Through the ' / meeting-raw-audio' API endpoint, the unit accepts raw binary audio data and subsequently queues it into an “audio_output_queue” for controller unit 100 to pass it to transcriber unit 118. In sum, API server unit 140 in the present invention, effectively bridges the virtual Al representative with external systems, while ensuring seamless event and audio data management.
[0041] Meeting application unit 142 used in the virtual Al representative is to provide a bidirectional communication channel between the virtual Al representative and a potential participant. The modular design of the virtual Al representative makes it possible for any meeting application to be used as a component as long as it has the capability of passing the raw audio and autonomous screen share. For the present innovation, Zoom SDK is used as the meeting application.
[0042] Data flow within the virtual Al representative core is depicted in FIG. 2. At the start of each conversation cycle, the conversation cycle consists of a back and forth between the participant and the virtual Al representative, upon reception of user's verbal communication (step 200), user input unit 102 commences speech-to-text conversion (step 202), resulting in one or more transcribed interim messages. Each transcribed interim message is tagged with a unique integer identifier before being forwarded to controller unit 100. In step 212 of FIG. 2, controller unit 100 sends an inquiry to user conversation encoder unit 134 to check if there is any available Al response in the cache before making an inquiry. Controller unit 100 sends an inquiry to knowledge base unit 126 to find relevant information based on the user’s message (step 204); if the poll results in any related information or answer, controller unit 100 creates a system message based on the poll. Controller unit 100 sends user messages alongside the system message to LLM interactive-conversation unit 108.
[0043] Upon receipt of LLM interactive-conversation unit 108 response (Al response) in step 206, the state of the conversation is determined (step 208) and controller unit 100 prompts audio generator unit 120 to synthesize an audio file corresponding to the Al response (step 210). The audio file may be played (step 214). Any visuals may be rendered on the screen according to the state and Al response. Once the audio file is generated, it is sent back to controller unit 100, and then forwarded to vocalizer unit 138, setting it in standby mode.
[0044] If a new interim message from the participant is detected during this process, the existing audio file is discarded. The system reverts to the interim message handling stage, and the cycle repeats to generate a new response for the virtual sales agent.
[0045] When user input unit 102 receives the participant's final spoken message (step 218 final state yes), controller unit 100 checks its similarity against the last interimmessage. If they are similar, controller unit 100 prompts vocalizer unit 138 to play the already generated audio. Otherwise, the system returns to the interim message handling stage (step 200) to generate a new Al response corresponding to the user's final message. This new response is then vocalized and played. At step 220, the conversation is ended. At step 222, next steps to support CRM are sent to CRM.
[0046] FIG. 3 draws an overview of the platform software architecture. User dashboard frontend 300 is a stand-alone application including Virtual Al representative frontend module 302 and knowledge base frontend module 304 that provides user 358 with access to create or manage virtual Al representative instances to present a product. User dashboard backend 306 includes API module 308 via API calls 322 to communicate with database 314 accessing data storage 312, and via API calls 324 to communicate with virtual Al representative instances, and fleet manager 310. Fleet manager 310 uses API calls 326 to communicate with virtual Al representative core instance 318.
[0047] In FIG. 3, presenter docker 316 is created using a serverless compute engine (such as AWS Fargate or similar services). User dashboard backend 306 oversees the containers, handling tasks such as creation, stopping, and status querying using fleet manager 310. Subsequently, fleet manager 310 invokes presenter docker 316. A new presenter container is initialized for every meeting session (i.e. presenter docker 316 is a dedicated container for only one meeting). Presenter docker 316 comprises two components: Virtual Al representative core instance 318 and meeting application 320. API calls 328 are used to communicate between Virtual Al representative core instance 318 and meeting application 320. API calls 328 are used to communicate between Virtual Al representative core instance 318 and meeting application 320.
[0048] Upon the initiation of a presenter docker container, two main instances are activated to start and manage the meeting. The first is virtual Al representative coreinstance 318, which is responsible for overseeing meeting application instance 320 and ensuring seamless communication with the user dashboard backend 306. Its role is pivotal; if this process were to exit, the container would stop functioning, indicating its significance in the architecture.
[0049] Meeting application instance 320 is launched in conjunction with virtual Al representative core instance 318. This secondary instance is governed by virtual Al representative core instance 318 and operates under the directives of a representational state transfer (REST) API specific to the meeting application. Its primary function is to start a meeting session that allows for the display of presentations through window sharing. Moreover, it supports bidirectional audio streams, facilitating interactive communication channels during meetings.
[0050] FIG. 4 illustrates an exemplary website indicated by Wishpond 302 that employs a virtual Al representative to present the product to interested leads. Upon clicking on Get a Demo 300 button, participant 500 is asked to login 301 and after logging in the participant 500 is asked for his / her email address and the meeting link is sent to the email address. By clicking on the Uniform Resource Locator (URL) or the colloquially known as an address on the Web, the meeting starts and the virtual Al representative starts the presentation showing how to reach new customers and increase sales affordably 303.
[0051] FIG. 5 illustrates in detail the chain of events when a participant requests a meeting / presentation. To start a presentation, fleet manager 310 starts presenter docker 316 and injects environment variables. The environment variables are: “meeting id” and API credentials. Meeting id identifies a specific instance of a virtual Al representative (e.g. the same participant might have multiple meetings scheduled).API credentials are used by virtual Al representative core instance 318 to call into API module 145.
[0052] Virtual Al representative instance 318 makes API calls to user dashboard backend 306 to fetch the blueprint of states, lead information (participant name to use in the meeting etc.), and knowledge base information.
[0053] Virtual Al representative core instance 318 kicks off the process by first stopping all existing meeting application instance 320 processes within presenter docker 316, and then starting meeting application instance 320 via the command line. Meeting application instance 320 sends a meeting URL to virtual Al representative core instance 318 via webhooks to http : / / localhost: 4000. Virtual Al representative core instance 318 sends the meeting URL to user dashboard backend 308 using REST API POST. When meeting application instance 320 starts, virtual Al representative core instance 318 controls it using a REST API located at local host: 3000 with “start_meeting,” “stop_meeting,” “play_audio,” and “share_window” end points.
[0054] Webhooks sent by meeting application instance 320 to virtual Al representative core instance 318 includes “meeting_started,” “meeting_stopped,” “meeting_failed,” “meeting_connecting,” “meeting_disconnecting,” “userjoined,” “userjeft,” “sharing_status_changed.”
[0055] Meeting application instance 320 sends raw audio from the participant to virtual Al representative core instance 318.
[0056] To launch a meeting, virtual Al representative core instance 316 fetches information about the meeting from dashboard backend 306, then runs a worker job to start the meeting (FIG. 6). Upon receiving the meeting URL from virtual Al representative core instance 318, user dashboard backend 306 sends the meeting URL to participant 500. If user dashboard backend 306 does not receive the meetingURL after a period of time, it can decide to terminate presented docker 316 and start the container again if desired.
[0057] In FIG. 7, virtual Al representative core instance 318 starts the meeting with an API call to meeting application instance 320 and sends the welcome voice snippet. Meeting application instance 320 confirms receiving the voice snippet and relays it to meeting instance 512. Then virtual Al representative core instance 318 initiates screen share and waits for the response from meeting instance 512. Upon receiving the response, virtual Al representative core instance 318 follows the steps in FIG. 2 and continues the conversation.
[0058] FIG. 8 illustrates user dashboard frontend 141. User 358 uses the software tool available on user dashboard frontend 141 to create and manage virtual Al representatives and the flow of the conversation via defining states for state manger unit 102. The example user dashboard shown contains entries Sales closer by Wishpond 401 , Al Agents 400, Knowledge base 402, Analytics 404, and Recordings 406 as user selectable selections. Details include voice 410, product 412, and Knowledge Base 414.
[0059]
[0060] FIG. 9 illustrates the hardware architecture of the present invention. The present invention's platform architecture is outlined as follows: Users engage with system server 900 via client device 901. Client device 901 connects to server 902 through network 914 and can operate on any chosen computing platform. Server 902 interfaces with client devices over this network to provide a user or graphical interface (GUI) for system 900. This interface, accessible via web browsers or specific software applications, facilitates data display, entry, publication, and management, acting as a meeting interface. The term "network" here refers to a network collection appearing asone to users, including the Internet, which connects using Internet Protocol (IP) and similar protocols. The public network 914 depicted in FIG. 9 serves only as an example.
[0061] Server 902 may offer services relying on a database system accessible over a network and via server 936. The GUI or meeting interface, provided by server 902 on client device 901 via a web browser or app, allows for operation and utilization of service system 900. The components in system server 902 and 936 represent a combination necessary for providing the services and tools envisioned by the invention. These components, which may communicate over a WAN or LAN, include an application server or executing unit 904 comprising a web server 906 and 942 and a computer server 908 and 944. The web server responds to HTTP requests from remote browsers or software applications, providing the necessary user interface. The computer server may include a processor 910 and 946, RAM, and ROM, controlled by operating system software for resource allocation and task management.
[0062] The database tier, with at least one database server 903, interfaces with multiple databases 912, updated via private networks including the Internet. Although described as a single database, separate databases can store various user data and files.
[0063] Application server 940, custom-built for this invention, enables various tasks related to creating and customizing the virtual Al representative sits on an exemplary system server 938. "User dashboard" henceforth refers to the web browser interfaces for accessing application server 940 of this invention. Application server 940 communicates with application 905 via API calls through network 914. “Virtual Al representatives instance” henceforth refers to application 905. Users interact withmeeting application 907 via web server 906. “Meeting instance” henceforth refers to the web interface of meeting application 907.
[0064] Client devices 901 may include a range of electronic devices with various components. For instance, client device 901 may feature a display 918, processor 920, input device 922, transceiver 924, memory 928, app 930, local data store, and a data bus interconnecting these components. The term "transceiver" encompasses any known transmitter or receiver for communication. These components may vary, and alternative embodiments are considered within the invention's scope.
[0065] In an embodiment, communication begins when an audio message is sent by either the virtual representative or the user, triggering the communication. This audio is then translated into written text, each instance of which is assigned a distinct numerical identifier before being forwarded to controller unit 100. Controller unit 100, in turn, instructs user conversation encoder unit 134 to search knowledge base unit 126 for pertinent information. Utilizing this information, the system crafts messages from both the system's and the user's perspectives and directs them to LLM interactive-conversation unit 108. LLM interactive- user conversation unit 108 then produces a text-based reply, which is subsequently synthesized into an audio message for the user's consumption in vocalizer unit 138. Should there be an interruption with a new message from the user while this process is underway, the audio response is modified to reflect this latest communication. Only an audio file that is confirmed to be current and representative of the user's most recent message is played. With each round of dialogue, the unique numerical tag is advanced, readying the system for the next round of interaction.
[0066] In an embodiment, at each step controller unit 100 uses LLM interactive- user conversation unit 108 and state manager unit 106 to infer the state and parameters ofthe conversation that are passed to action controller unit 110 to create the suitable action to be presented on the screen alongside the vocalized response from LLM interactive-conversation unit 108. Synchronizing the visual part of the interactive presentation with the conversation is a challenge that this embodiment addresses via interaction between controller unit 100, action controller unit 110, and state manager unit 106.
[0067] The embodiment further includes the various states of the conversation comprising preparation, hold, wait, abandon, or finalized. There may be further states as well and this is flexible and may be provided to controller unit 100. For each different product that the Al virtual representative presents, the number of states can be adjusted accordingly.
[0068] Fine-tuning LLM interactive-conversation unit 108 for interactive conversation is essential because standard NLP models 200 may not be optimized for real-time, interactive dialogues, and they might produce responses that are not contextually accurate or coherent. Leveraging an LLM interactor as a knowledge base for context, combined with another LLM interactor for user profiling that provides related information as personalized context, can help fine-tune pre-trained language models on domain-specific data, thereby significantly enhancing performance and yielding more contextually accurate and coherent responses.
[0069] Synchronizing conversation flow and interactive presentation is an essential aspect in creating a seamless transition especially when the presentation is conditional on the dialogue flow. To solve this problem, in the present invention, event-driven architecture is implemented in controller unit 100 to trigger specific presentation steps based on a blueprint provided to state manager unit 106 at the time of the creation of the Al virtual representative code 154. State manager unit 106 is a robust dialoguemanagement system used by controller unit 100 alongside the LLM interactor-user conversation unit 108 that is capable of adaptively controlling the flow of the conversation. To create synchronization between the audio and video controller unit 100 infers the step and parameters of the conversation from the response of LLM interactor-conversation unit 108 and sends it to action controller unit 110 to be played alongside the vocalized response of LLM interactor-user conversation unit 108.
[0070] Harmonizing asynchronous threads is a complex task, especially when multiple threads are running to monitor various aspects of the conversation, including user engagement, sentiment, or intent. However, in the present invention, the use of message queues, shared state-management systems, flags, and events within the threads can be instrumental in synchronizing these various asynchronous tasks, ensuring a more coherent interaction.
[0071] Maintaining a natural conversation flow and minimizing response delay are crucial for user experience. To ensure a conversation feels natural, the system must generate responses within a fraction of a second, a challenge due to both the computational complexity of LLMs and the network response rate. One solution is to implement a stateful conversation model that remembers past interactions and context, helping preserve a seamless flow. When users pose a new inquiry, controller unit 100 polls user conversation encoder unit 134 to identify useful Al responses from the past. If a match is found, controller unit 100 quickly prompts vocalizer unit 138 to ensure a swift and relevant reply.
[0072] Systems such as traditional sales models that rely heavily on human agents to manage customer queries, presentations, and follow-ups often face scalability challenges. In contrast, the virtual Al representative can manage multiple interactionsat once and offers easy scalability. This capability enables businesses to cater to an expanding customer base without the need to proportionally increase their workforce.
[0073] Systems that rely heavily on human resources, such as those with a large sales team, can become expensive due to salaries, benefits, and training costs. In contrast, the virtual Al representative described in this invention offers a more cost- effective solution over time. The virtual Al representative not only eliminates the need for a sizable team but also ensures continuous 24 / 7 service.
[0074] Human representatives might sometimes lack immediate access to comprehensive customer data, hindering their ability to offer a truly personalized experience. In contrast, the Al virtual representative has the capability to swiftly analyze user’s data, enabling it to provide highly personalized recommendations and solutions. This not only enhances user engagement but also potentially boosts conversion rates.
[0075] Human representatives can occasionally experience off days, and their level of expertise might differ from one individual to another, which can result in varying presentation experiences. On the other hand, the virtual Al representative is designed to provide a consistent level of service, guaranteeing that each interaction aligns with the desired quality standards.
[0076] Unlike human representatives who aren't available 24 / 7, potentially posing challenges for businesses that operate across various time zones or for users who seek interactions beyond standard business hours, the virtual Al representatives have the advantage of being available continuously. This ensures constant support and engagement for users at any given time.
[0077] While human representatives typically manage just one interaction at a time and might exhibit slower response times during peak hours or while multitasking, thevirtual Al representatives excel in offering prompt feedback. This capability ensures that users receive answers or information with minimal delay, enhancing the overall user experience.
[0078] Decision-making during a course of a real-time interaction often hinges on intuition and experience rather than concrete data when done by human representatives. However, the virtual Al representative is equipped to amass and scrutinize extensive data, furnishing invaluable insights into user behaviors and predilections. Such insights can be pivotal for shaping future strategies and making informed decisions. This advantage is not just limited to sales; various other domains can also benefit from employing virtual Al representatives to harness data-driven insights.
[0079] When businesses or organizations venture into global markets, they often encounter language barriers, especially if they lack employees proficient in the target market's language at various locations. In contrast, virtual Al representatives can be endowed with capabilities to understand and communicate in multiple languages. This adaptability facilitates seamless engagement with a diverse and global user base.By addressing these challenges, the present invention provides a virtual Al representative offers a transformative solution for businesses and organizations, enabling them to improve customer engagement, drive sales, and operate more efficiently, improve customer care, and serve better.
Claims
1 . A computer program product for facilitating a directed conversation according to a criteria between an artificially-intelligent (Al) agent and an audience having program instructions embodied therewith, the program instructions executable on a processing circuit to cause the processing circuit to perform the steps comprising: receiving by an Al system including the Al agent, a state machine controlling the directed conversation; ingesting a knowledge base for the directed conversation by the Al system; and provide an interactive presentation for the directed conversation, by the Al agent, as indicated by the state machine.
2. The computer program product of claim 1 , wherein each entry in the received state machine further comprises: a name, an instruction to perform by the Al agent, a transition condition, a next state, and an action for the Al agent to take after performing the instruction.
3. The computer program product of claim 1 , wherein the criteria is defined by the received state machine.
4. The computer program product of claim 1 , wherein the directed conversation is emulating a virtual sales agent.
5. The computer program product of claim 1 , wherein the directed conversation is emulating a virtual customer service representative.
6. The computer program product of claim 1 , wherein the directed conversation is emulating a virtual healthcare provider.
7. The computer program product of claim 1 , further comprising:incorporating a set of predefined built-in system states in the state machine and wherein the set of predefined built-in system states include audio connection, a first state, a hold state, an interrupt state, and a tangent state.
8. The computer program product of claim 1 , wherein the directed conversation is emulating a virtual healthcare provider.
9. The computer program product of claim 1 , wherein the directed conversation is emulating a virtual instructor.
10. The computer program product of claim 1 , wherein the directed conversation is emulating a virtual advisory service provider.
11. The computer program product of claim 1 , wherein the directed conversation is emulating a virtual recruiter.
Citation Information
Patent Citations
Customizable expert agent
US20030028498A1
Conversation runtime
US20180131642A1
Artificial conversational entity methods and systems
US20190087707A1