Multi-party communication using large language model intermediary

By using a Large Language Model (LLM) as an intermediary, multi-party communication is managed based on scenario templates and role assignments, solving the problem of inefficiency in existing systems and enabling more efficient and accurate multi-party communication and direct communication options.

CN121753016APending Publication Date: 2026-03-27M3G TECHNOLOGY CO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-22
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Multi-party communication is difficult to manage efficiently in complex situations, leading to miscommunication, delays, and privacy issues. Existing systems are unable to provide direct communication options as needed.

Method used

Using a Large Language Model (LLM) as an intermediary, output messages in multi-party communication are generated and managed through scene templates and role assignments. The LLM processes input messages and generates response outputs according to a predefined grammatical structure and prompt words.

Benefits of technology

It improves the efficiency and accuracy of multi-party communication, reduces miscommunication and delays, provides direct communication options as needed, and enhances communication management and problem-solving capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121753016A_ABST
    Figure CN121753016A_ABST
Patent Text Reader

Abstract

A method for providing multi-party communication using a large language model (LLM) is provided. The method begins with receiving an input message from a user device associated with a plurality of communicating parties participating in a scenario. The method continues with the following operations: generating text input cues for LLM based on the scene template. The method further includes sending the text input cue word to the LLM. The LLM analyzes the text input cues to compile a response output of the LLM. The method continues by analyzing the response output of the LLM to generate an output message to be sent to one or more of the plurality of communicators. The method also includes sending an output message from the communication endpoint representing the LLM to a user device associated with one or more of the plurality of communicators.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 461,068, filed April 21, 2023, entitled “MULTIPARTY COMMUNICATION USINGA LARGE LANGUAGE MODEL INTERMEDIARY,” which is incorporated herein by reference in its entirety for all purposes. Technical Field

[0003] This disclosure generally relates to communication systems, and more specifically, to an architecture for multi-party communication and problem-solving using a large-scale language model (also known as a large language model (LLM)) as an intermediary between communicating parties, the architecture having the ability to selectively enable direct communication between communicating parties at the discretion of the licensee. Background Technology

[0004] In various scenarios, such as customer support or service collaboration, multi-party communication can become complex and difficult to manage. Traditional methods that facilitate communication, such as teleconferencing or group messaging, may be inefficient and can lead to miscommunication, delays, privacy issues, and / or loss of control. There is a need for improved systems and methods to provide, manage, and coordinate multi-party communication, while also offering options for direct communication between parties as needed. Summary of the Invention

[0005] This summary is provided to introduce, in a simplified form, some concepts further described in the following detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0006] According to an example embodiment of this disclosure, a system for providing multi-party communication using an LLM is provided. The system may include a processor and a memory storing instructions to be executed by the processor. The processor may be configured to receive input messages from user devices associated with multiple communicating parties participating in a scenario. The LLM may be configured to understand the scenario. The processor may also be configured to generate text input prompts for the LLM based on a scenario template. The scenario template may include rules for organizing the input messages based on the scenario. The multiple communicating parties and the LLM may be assigned roles as participants in the scenario. The scenario template may include rules for instructing the LLM to analyze the text input prompts and compute a response output of the LLM defined by the LLM's role in the scenario. The processor may also be configured to send the text input prompts to the LLM. The LLM may analyze the text input prompts to compose the LLM's response output. The processor may also be configured to analyze the LLM's response output based on a syntactic structure provided for the LLM in the scenario template to generate an output message to be sent to one or more of the multiple communicating parties. The one or more communicating parties may be identified in the LLM's response output by the roles of those one or more communicating parties. The processor can also be configured to send output messages from the communication endpoint representing the LLM to a user device associated with one or more of a plurality of communication parties, based on the role of the LLM in the scene.

[0007] According to an example embodiment of this disclosure, a method for providing multi-party communication using an LLM is provided. The method may begin by receiving an input message from a user device associated with multiple communicating parties participating in a scenario. The LLM may be configured to understand the scenario. The method may continue by generating text input prompts for the LLM based on a scenario template. The scenario template may include rules for organizing the input message based on the scenario. The multiple communicating parties and the LLM may be assigned roles as participants in the scenario. The scenario template may include rules for instructing the LLM to analyze the text input prompts and compute a response output defined by the LLM's role in the scenario. The method may further include sending the text input prompts to the LLM. The LLM may analyze the text input prompts to compose the LLM's response output. The method may continue by analyzing the LLM's response output based on a syntactic structure provided for the LLM in the scenario template to generate an output message to be sent to one or more of the multiple communicating parties. The one or more communicating parties may be identified in the LLM's response output by the roles of those one or more communicating parties. The method may also include sending output messages from a communication endpoint representing the LLM to a user device associated with one or more of a plurality of communication parties, based on the role of the LLM in the scenario.

[0008] According to yet another example embodiment of this disclosure, the operation of the above method is stored on a non-transitory computer-readable storage medium including instructions that, when implemented by one or more processors, execute the described operation.

[0009] Other exemplary embodiments and aspects of this disclosure will become apparent from the following description taken in conjunction with the accompanying drawings. Attached Figure Description

[0010] In the figures of the accompanying drawings, exemplary embodiments are shown by way of example rather than limitation, and the same reference numerals in the drawings indicate similar elements.

[0011] Figure 1 An environment within which systems and methods for providing multi-party communication using LLM can be implemented is shown.

[0012] Figure 2 This is a schematic diagram illustrating the interaction provided between communicating parties by a system for providing multi-party communication using LLM, according to an example implementation.

[0013] Figure 3 This is a schematic diagram illustrating the interaction provided between communicating parties by a system for providing multi-party communication using LLM, according to an example implementation.

[0014] Figure 4 This is a schematic diagram illustrating the architecture of a system for providing multi-party communication using LLM, according to an example implementation.

[0015] Figure 5 This is a schematic diagram illustrating how a first participant, a second participant, and a third participant access a multi-party session within the context of a scenario, according to an example implementation.

[0016] Figure 6 This is a schematic diagram illustrating a system for providing multi-party communication using an LLM as a robot, according to an example embodiment.

[0017] Figure 7 This is a schematic diagram illustrating components of a system for providing multi-party communication using an LLM, according to an example embodiment.

[0018] Figure 8 This is a block diagram illustrating a sequence of steps for enabling a system acting as a robot to access multi-party communication using LLM, according to an example embodiment.

[0019] Figure 9 This is a schematic diagram illustrating key components of prompts written for and provided to an LLM according to an example implementation.

[0020] Figure 10 This is a schematic diagram illustrating the various components of a prompt word according to an example implementation.

[0021] Figure 11 An example dialogue supported by a system for providing multi-party communication using LLM is shown according to an example implementation.

[0022] Figure 12 This is a flowchart of a method for providing multi-party communication using an LLM, according to an example implementation.

[0023] Figure 13 A schematic diagram illustrating the various components of a labeled prompt word according to an example implementation is shown.

[0024] Figure 14 This is a flowchart of a method for providing multi-party communication using an LLM, according to an example implementation.

[0025] Figure 15 This is a high-level block diagram of an example computer system in which a set of instructions can be executed to cause the machine to perform any or more of the methods discussed herein. Detailed Implementation

[0026] The following detailed description includes references to the accompanying drawings, which form part of the detailed description. The methods described in this section do not constitute prior art to the claims and are not considered prior art by virtue of their inclusion in this section. The drawings illustrate exemplary embodiments. These exemplary embodiments (which may also be referred to herein as "examples") are described in sufficient detail to enable those skilled in the art to practice the subject matter. Embodiments may be combined, other embodiments may be utilized, or structural, logical, and operational changes may be made without departing from the scope of the claims. Therefore, the following detailed description should not be considered limiting, and its scope is defined by the appended claims and their equivalents.

[0027] In general, embodiments of this disclosure relate to systems and methods for providing multi-party communication facilitated by generative artificial intelligence (AI) and LLM. The LLM can be capable of processing text input and producing text output. Example LLMs can include GPT-4®, a transformer-based model pre-trained to predict the next lexical unit in a document (see “GPT-4 Technical Report,” OpenAI (2023) (https: / / cdn.openai.com / papers / gpt-4.pdf - 27 Mar 2023)). LLMs have demonstrated human-level performance on various professional and academic benchmarks. However, the system and method also anticipate future advancements in language models that can modify or replace transformer architectures. It should not be construed as being limited to this architecture.

[0028] LLM is a machine learning model that can recognize, predict, and generate human language based on very large text-based datasets. The system disclosed herein can form a request to an LLM in the form of text input prompts, receive a response output from the LLM to the request, and process the response output from the LLM to generate an output message to be sent to one or more communicating parties.

[0029] Traditional systems used for providing multi-party communication typically employ bots (such as chatbots), which are designed to communicate back to a single person, similar to how most people use ChatGPT®. When these systems connect to customer support, the customer sends a message, which is initially intercepted by the bot that connects to the customer's two-way conversation. When the bot can no longer handle the conversation or the customer requests to communicate with a human agent, the bot transfers the conversation to the human agent. The conversation then becomes a separate two-way conversation between the customer and the human agent, without bot interaction. At this stage, the bot may offer suggestions to the human agent on what to say or to summarize the conversation, but the bot no longer communicates with the customer.

[0030] Replacing two separate two-way conversations with multi-party conversations involving customers, human support agents, and chatbots can improve customer experience and make human support agents more effective. This is an important special case of multi-party communication.

[0031] In an example implementation of the system disclosed herein, when a customer sends a message, a bot provided by the system joins the conversation along with a human agent. According to one example implementation, the conversation is a multi-party conversation with two human participants (the customer and the human agent) and a bot coordinating the conversation. The bot can handle all initial responses from the customer, but it can also bring in the human agent as needed for more information or instructions. The bot never leaves the conversation. This approach has key advantages. First, the bot doesn't have to "think" independently. Specifically, the bot knows the agent can help answer questions, so the bot doesn't have illusions. Second, the bot helps the agent be more productive, and the bot determines when and whether the agent needs to join in. When the agent does join in, the agent can speak directly to the customer, and the agent can also communicate through the bot. Third, the bot can independently determine how to handle the issue, including updating external systems. However, even in this implementation, the bot calculates the determinism of the update and confirms it with the agent before execution, if the update determinism does not exceed a threshold set by the organization.

[0032] There exist multiple parties who need to communicate over a period of time, with participants having diverse business and personal backgrounds and being geographically distant from each other. When communication is synchronous or "real-time" (such as in a phone call or video conference), participants are assumed to be fully engaged and focused, and can use cues derived from the session to know when they are being asked to respond. In multi-party communication where participants are geographically distant and the communication is asynchronous, each participant is less aware of whether they need to pay attention to the session or when they are being asked to respond. Participants are also referred to as the communicating parties in a multi-party communication.

[0033] When two parties are communicating, if the first party sends a message, the second party knows that the message is addressed to the first party without having to read it. The second party may need to read the message to know if it is urgent or requires a response, but there is no doubt that the message was intended for the second party. In multi-party communication, if the first party sends a message, the other parties do not know whether they need to pay attention to it or whether they should respond unless they all read the message.

[0034] In an example three-way conversation involving a delivery to a customer's home, involving a customer, dispatcher, and delivery driver, if the customer asks, "When is the ETA?" (where "ETA" refers to the estimated time of arrival), both the driver and the dispatcher need to review the customer's question to determine who should respond. In this case, the driver is the logical participant who should respond. Therefore, there are two questions facing the participants in a three-way conversation: which participant should pay attention to the message and who should provide a response to it.

[0035] As the number of simultaneous sessions increases, as does the number of participants in any session, these problems become a more complex combination, thus making a universal solution necessary.

[0036] The system disclosed herein is based on a script conception in which several characters play different roles to advance the plot. Given a plot and the first few lines of dialogue in the script, the LLM can accurately predict the optimal next line in the script. The system disclosed herein maps a multi-party session to the form of a script, where the LLM is one of the characters and the writer of their own lines. The communicating parties in the multi-party session are other characters with specific roles and backstories. The context of the session is the occasion. The goals of the participants in the scene are the driving forces of the plot. The LLM's role is as a facilitator, promoter, or assistant to help the parties communicate to resolve the core issues of the story. When each communicating party submits a message, the system composes a cue word that includes all the lines, characters, plot, and occasion of the script, and asks the LLM to write the next line for the character they are playing to say to the other characters. The system accepts these lines and sends them as messages to the other communicating parties. Thus, each participant in the multi-party session participates in a script that is written as they speak.

[0037] The system disclosed herein facilitates communication between multiple communicating parties, where an LLM acts as an intermediary. The system can be configured to receive input messages from each communicating party and algorithmically format the input messages to conform to a specified syntax structure that can be used by the LLM. After formatting the input messages, the system can send the input messages to the LLM. The LLM can process the input messages and generate its own response output based on the input messages. The system can also be configured to analyze the content of the response output received from the LLM to draft output messages to the appropriate communicating party in response to the input messages using the specified syntax structure. The system can algorithmically format the output messages generated based on the LLM's response output and route these output messages to the appropriate communicating party.

[0038] The system disclosed herein can be associated with prompts specifically designed for LLM to facilitate communication between multiple communicating parties. The prompts can be constructed using natural language and are also referred to herein as text input prompts. Prompts may include settings defining the activities and roles of the LLM in a session, a definition of the number of communicating parties, and a definition of the persona of each communicating party (potentially explicitly defining which information each communicating party accesses). Prompts may also include specified actions to be taken in the event of missing information to keep communicating parties informed and / or alleviate illusions. Prompts may also include a defined syntactic structure of input messages and response outputs for interfacing with the computation unit and may include potential examples from previous sessions. The computation unit may also switch between different prompts for LLM whenever a specific condition is met in a session.

[0039] Various embodiments are now described with reference to the accompanying drawings, wherein the same reference numerals denote the same parts and components throughout the various views. It should be noted that reference to various embodiments does not limit the scope of the appended claims. Furthermore, any examples outlined in this specification are not intended to be limiting, but merely illustrate some of the many possible embodiments with respect to the appended claims.

[0040] Now refer to the attached diagram, Figure 1 An environment 100 is shown within which systems and methods for providing multi-party communication using LLM can be implemented. As used herein, the term "multi-party communication" refers to communication between two or more communicating parties. Environment 100 may include a communicating party 102 associated with user device 104, a communicating party 106 associated with user device 108, a communicating party 110 associated with user device 112, a system 114 (also referred to herein as system 114) for providing multi-party communication using LLM, and a data network shown as network 116. System 114 may communicate with user device 104, user device 108, and user device 112 (also referred to herein as communication endpoints) via network 116.

[0041] Each of user device 104, user device 108, and user device 112 may include, but is not limited to, smartphones, laptops, personal computers, desktop computers, tablets, phablets, personal digital assistants, mobile phones, smart TVs, personal computing devices, etc.

[0042] System 114 may include a computing unit 118, a memory 120 storing instructions to be executed by the computing unit 118, and an AI unit 122. System 114 may also include or communicate with an LLM 124. LLM 124 may include a machine learning model configured to recognize, predict, and generate human language based on a very large text-based dataset. System 114 may communicate with LLM 124 via network 116.

[0043] In one example implementation, AI unit 122 may use generative AI. Generative AI can be configured (typically in response to prompts) to use a generative model to generate text, images, videos, or other data. LLM 124 is an example of a generative model. LLM 124 may be a generative AI model configured to learn patterns and structures from input training data and generate new data with similar characteristics. LLMs can also be configured to use retrieval-enhanced generation (RAG) or similar techniques to provide higher-quality responses by leveraging data sources that may contain proprietary data, including vector databases.

[0044] In one example implementation, the operation of computing unit 118 and AI unit 122 may be performed by one or more processors communicating with memory 120. In some example implementations, the operation of LLM 124 may be performed by one or more processors.

[0045] Network 116 can refer to any wired, wireless, or optical network, including, for example, the Internet, intranet, local area network (LAN), personal area network, wide area network (WAN), virtual private network, Wi-Fi® network, cellular telephone network (e.g., GSM network, packet-switched communication network, circuit-switched communication network), Bluetooth™ radio, Ethernet, IEEE 802.11-based radio frequency network, frame relay network, Internet Protocol (IP) communication network, or any other data communication network that utilizes physical layer, link layer, or network layer capabilities to carry data packets, or any combination of the data networks listed above. In some implementations, network 116 may include an enterprise network, data center network, service provider network, mobile operator network, or any combination thereof.

[0046] Figure 2 This is a schematic diagram 200 illustrating the interaction provided between communicating parties by a system 114 for providing multi-party communication using LLM, according to an example implementation. Figure 2In this scenario, a first participant (shown as communicator 102), a second participant (shown as communicator 106), and a third participant (shown as communicator 110) access multi-party communication using a first communication medium (shown as user device 104), a second communication medium (shown as user device 108), and a third communication medium (shown as user device 112), respectively. Communicator 102 sends an incoming message 204 (e.g., "When is the ETA?"). For typical asynchronous communication, communicator 106 and communicator 110 receive incoming message notifications 206 and 208, respectively. When communicator 106 and communicator 110 receive incoming message notifications 206 and 208, respectively, they need to decide whether they should pay attention to the incoming message notifications 206 and 208, open the incoming message 204 to view it, and further decide whether they should respond. These problems, which are resolved by system 114, are shown as multi-party communication questions 202.

[0047] For arbitrary messages without context, it may be unclear whether any given participant in a multi-party communication should pay attention and respond. However, many multi-party communications take place within the context of a specific scenario, where each participant is playing a role, and within that role, actions are expected from the participants, and information about the scenario of the multi-party communication develops as the communication unfolds over time.

[0048] Figure 3 This is a schematic diagram 302 illustrating the interaction provided between communicating parties by a system 114 for providing multi-party communication using LLM, according to an example embodiment. System 114 can receive requests in the form of incoming messages 204 from communicating parties such as communicating party 102 (e.g., a customer), communicating party 106 (e.g., a dispatcher from a service provider, such as a trailer dispatcher), and communicating party 110 (e.g., a driver providing service on behalf of a service provider). In one example embodiment, system 114 may act as a robot providing and facilitating interaction between communicating parties. In one example embodiment, system 114 can provide multi-party communication between two or more communicating parties, which may include human communicating parties and robots provided by system 114. Although... Figure 3 Three human communication parties (customer, dispatcher, and driver) are shown, but system 114 can be configured to provide multi-party communication between two or more human communication parties.

[0049] Upon receiving incoming message 204, system 114 can access data related to scenario 310. Data related to scenario 310 may include data related to the following: the circumstances associated with multi-party communication (e.g., the type of service to be provided), the goal of multi-party communication, the participants in multi-party communication, etc.

[0050] In one example implementation, system 114 may communicate with an operations management system associated with a service provider. The communicating party 106 may be a dispatcher of the service provider. Examples of operations management systems 306 include transportation management systems (TMS) for managing delivery (e.g., third-party systems from Onfleet®, Oracle®, Project44®, or BlueYonder®, and company-owned systems such as those operated by Ryder® or DoorDash®) and customer relationship management (CRM) systems for managing field services (e.g., third-party systems from Salesforce® and Microsoft®).

[0051] System 114 can also be configured to receive context 304 associated with multi-party communication. In one example implementation, context 304 may be received from an operations management system 306. Context 304 may include previous messages associated with previous sessions between any of these communicating parties, business rules of the service provider or a party providing services on behalf of the service provider, order details, the time of request for service, etc.

[0052] System 114 can combine data associated with incoming message 204 and scenario 310 with context 304 and provide the combined data to LLM 124 communicating with system 114. System 114 also provides one or more multi-party communication queries to LLM 124. The one or more multi-party communication queries may include questions about who to respond to, how to respond, how to update, and who needs to follow up. LLM 124 can be configured to provide responses to one or more multi-party communication queries based on the data provided by system 114 to LLM 124.

[0053] LLM 124 can receive and process incoming message 204, scenario 310, context 304, and one or more multi-party communication queries, and provide response output to system 114. The response output may include one or more responses 308 to be sent to the communicating parties, and multi-party communication instructions. Multi-party communication instructions may include responses to one or more multi-party communication queries. System 114 can parse the information contained in the response output to determine one or more responses 308 and responses to one or more multi-party communication queries. Responses to one or more multi-party communication queries may include responses to queries that include: which communicating party to respond to, how to respond, how to update data associated with the multi-party communication, and who needs to follow up in the multi-party communication.

[0054] Based on the information contained in the multi-party communication instructions, system 114 can send one or more responses 308 to communication parties 102, 106, and 110 as needed. Furthermore, based on the multi-party communication instructions, system 114 can update information associated with the multi-party communication in the operation management system 306. In a response 308, system 114 can tag a communication party only if that communication party needs to be aware of the response 308 (i.e., add a tag associated with that communication party to the response 308).

[0055] Upon receiving response 308, any of the communicating parties 102, 106, and / or 110 can re-establish multi-party communication by sending an additional incoming message. System 114 triggers a repetition upon receiving an additional incoming message. Figure 3 The process is shown in the figure.

[0056] Figure 4 This is a schematic diagram illustrating the architecture 402 of a system 114 for providing multi-party communication using LLM, according to an example implementation. The architecture 402 of system 114 can be hosted in the cloud, such as Google Cloud® or Amazon Web Services (AWS)®.

[0057] System 114 may include an AI unit 122, a computing unit 118, a memory unit 120, and an application programming interface (API) 406. System 114 can communicate with and be embedded in a third-party operation management system 408. System 114 can communicate with and be embedded in a communication system 410. System 114 may also include or communicate with an LLM 124. The LLM 124 may be hosted by a cloud computing platform such as Microsoft® and Azure®. The LLM 124 used by system 114 may include GPT-4®, Gemini®, Llama®, Falcon®, and other language models.

[0058] System 114 can use LLM 124 to facilitate at least five key AI functions of AI unit 122. First, AI unit 122 can be configured to use LLM 124 to manage and coordinate chat in a session between multiple communicating parties. Second, AI unit 122 can be configured to translate incoming messages received by system 114 from communicating parties and output messages sent by system 114 to communicating parties between multiple languages. Third, AI unit 122 can be configured to read the entire session, determine which communicating party needs to pay attention to which message, and tag the appropriate communicating party when the message is sent to it to alert that party to the message. Fourth, AI unit 122 can be configured to formulate and update data in the operation management system communicating with system 114 by sending update requests based on interactions occurring in multi-party communication. Fifth, AI unit 122 can be configured to provide deterministic factors. These deterministic factors can be calculated by AI unit 122 by comparing the content of update requests with the content of the actual session. Then, AI unit 122 determines whether the deterministic factor exceeds a predetermined deterministic threshold in order to determine whether an automatic update to the operation management system is required (if the predetermined deterministic threshold is not exceeded), or whether the communication party (e.g., the scheduler) needs to pay attention to the session before the update is performed (if the predetermined deterministic threshold is exceeded).

[0059] System 114's API 406 can use the communication API to connect to multi-party communication system 410 to send and receive messages from multiple communicating parties. Examples of communication systems include third-party services such as TalkDesk®, Simple Messaging Service (SMS), WhatsApp®, Slack®, and Riptide®, as well as embedded communication systems (such as those embedded in TMS and CRM systems).

[0060] API 406 can also be configured to provide updates. In one example implementation, to provide an update, API 406 can trigger a workflow within the operation management system 408 to cause the correct update to occur.

[0061] API 406 can also be configured to collect context. In one example implementation, to collect context, API 406 can query the operations management system 408 to provide order information, business rules, or data about the communicating parties.

[0062] Figure 5 This is a schematic diagram 500 illustrating, according to an example implementation, enabling a first participant (shown as communicator 102), a second participant (shown as communicator 106), and a third participant (shown as communicator 110) to access a multi-party session in the context of scenario 310. Example scenario 310 involves a driver delivering a package to a customer, which is assigned and managed by a dispatcher. The first participant (communicator 102) plays a first role 502 (e.g., customer). The second participant (communicator 106) plays a second role 504 (e.g., dispatcher). The third participant (communicator 110) plays a third role 506 (e.g., driver). Given this scenario 310, the roles, and the content of an incoming message 204 from the first participant (communicator 102) playing the first role 502, system 114 can act as a smart session agent in the multi-party session. System 114 can be configured to infer that the second participant (communicator 106) playing the second role 504 can ignore the incoming message notification (i.e., Figure 2 The incoming message notification 206 shown is illustrated in box 508. Furthermore, system 114 can be configured to infer that the third participant (communication party 110) playing the third role 506 should pay attention to the incoming message notification (i.e., Figure 2 The incoming message notification 208 shown (as shown in box 510) should be responded to and answered in the query in incoming message 204 (as shown in box 512).

[0063] Therefore, given an understanding of scenario 310, roles, and session input (i.e., incoming message 204), system 114 can act as an intelligent session broker and facilitate multi-party sessions by converting the multi-party session into several bidirectional sessions between each participant and the intelligent session broker. Alternatively, system 114 acting as an intelligent session broker can facilitate multi-party sessions by acting as an active participant in the multi-party session, accessing it when other participants are slow to respond or unable to respond, or prompting other participants to access the multi-party session. In some example implementations, system 114 acting as an intelligent session broker can facilitate multi-party sessions by acting as an active participant in the multi-party session while simultaneously accessing a bidirectional or multi-party session about that multi-party session individually with selected participants. Many other example implementations may exist for system 114 acting as an intelligent session broker to facilitate multi-party sessions.

[0064] Figure 6 This is a schematic diagram 600 illustrating a system 114 acting as a robot 602 according to an example embodiment. As used herein, robot 602 may include an autonomous program that operates in a network and can interact with systems and / or users in the network.

[0065] Robot 602 can be configured to understand scenario 310, roles 502, 504, 506, and incoming message 204 (session input). Robot 602 can act as an intermediary to conduct bidirectional conversations with each of the first participant (shown as communicator 102), the second participant (shown as communicator 106), and the third participant (shown as communicator 110).

[0066] Given scenario 310, roles, and the content of incoming message 204 from the first participant (communication party 102) playing the first role 502, system 114 can be configured to infer incoming message notification to the second participant (communication party 106) playing the second role 504 (i.e., Figure 2 The incoming message notification 206 shown can be suppressed or downgraded, as shown in block 604. Furthermore, system 114 can be configured to infer incoming message notifications (i.e., notifications to the third participant (communication party 110) playing the third role 506) to the third participant (i.e., the communication party 110). Figure 2 The incoming message notification 208 shown should be given priority (as shown in box 606), and the system 114 can be configured to infer that a response to the message is specifically requested from the third participant (communication party 110) playing the third role 506 (as shown in box 608).

[0067] In the second example implementation, when robot 602 is instilled with an understanding of scenario 310, roles 502, 504, 506 and incoming message 204 (session input), it can facilitate the multi-party conversation by acting as an active participant in the multi-party conversation, joining when other participants are slow to respond or unable to respond, or prompting other participants to join the multi-party conversation.

[0068] In the third example implementation, when robot 602 is instilled with an understanding of scenario 310, roles 502, 504, 506, and incoming message 204 (session input), it can simultaneously access a multi-party conversation by acting as an active participant in the multi-party conversation, while simultaneously accessing a two- or multi-party conversation about the multi-party conversation separately with selected participants, thereby coordinating multiple dialogues of one or more subsets of multiple communicating parties as part of the same conversation associated with the multi-party communication.

[0069] This disclosure addresses the aforementioned needs by providing an architecture for multi-party communication and problem-solving using an LLM as an intermediary between communicating parties. The system of this disclosure is configured to manage communication more efficiently, share information among communicating parties as needed, and facilitate problem-solving. Furthermore, the system provides the licensee with the ability to selectively enable direct communication between two or more communicating parties based on their own discretion.

[0070] The system disclosed herein may include a robot 602 acting as an intelligent conversation agent configured to participate in multi-party conversations by acting as an intermediary between participants who communicate asynchronously and play roles within the context of a human scene. The robot 602 may communicate with an LLM configured to understand scene 310, roles 502, 504, 506, and incoming message 204 based on pre-training of the LLM, fine-tuning of the LLM using additional training data, or by incorporating additional data into the cue words using RAG or similar methods.

[0071] Figure 7 This is a schematic diagram 700 illustrating components of a system according to an example embodiment of the present disclosure that act as a robot 602. In such... Figure 7 In the example implementation shown, the system of this disclosure may include two main components. The first component is an LLM 124, which can be configured to receive specially constructed prompt words and generate the LLM's response output using a predetermined syntax structure that indicates to which messages in the response output should be routed. The second component is a computing unit 118, which can process the response output of the LLM 124 and format the input messages being provided to the LLM 124 using predetermined rules, both based on the predetermined syntax structure.

[0072] Robot 602 can be configured to communicate with multiple participants (such as communicator 102, communicator 106, and communicator 110) by sending messages between a communication endpoint 702 associated with robot 602 and user devices 104, 108, and 112 respectively associated with communicators 102, 106, and 110. Robot 602 can also be configured to communicate with an additional communicator 710 (shown as participant N) using a user device 712 associated with an additional communicator 710.

[0073] The computing unit 118 can receive input in the form of input messages from various external data sources 708, and calculate and write the input in the form of prompts (i.e. text input prompts) for LLM 124.

[0074] LLM 124 may include OpenAI®'s GPT-4® or other LLMs that are sufficiently capable of correctly understanding and responding to text input prompts. LLM 124 may be configured to receive input in the form of text input prompts (as shown in step 714) and compute output in the form of a response output. Text input prompts may include natural language prompts, text in a data exchange format (such as JSON (JavaScript Object Notation), XML (Extensible Markup Language), or other data exchange formats), or a combination of both. Response output may include natural language responses, text in a data exchange format (such as JSON, XML, or other data exchange formats), or a combination of both.

[0075] In other implementations, the input prompt may have a combination of modalities, such as text, image, audio, video, location, and any other data type or format. A modality can refer to a defined data category based on how the data is received, represented, and understood. In other implementations, the response output may have a combination of modalities, such as text, image, audio, video, location, and any other data type or format.

[0076] LLM 124 can compute response output based on text input prompts. LLM 124 can compute response output based on pre-training of LLM 124 across multiple scenarios, fine-tuning of LLM using additional training data, or incorporating additional data into the prompts using RAG or similar methods. The response output of LLM 124 can be structured and formatted based on predetermined rules, including the grammatical structure provided in the text input prompts. The response output can include a set of messages, including message content and instructions on which participant (communicator) each message should be sent to.

[0077] The computing unit 118 can receive response output from the LLM 124 (as shown in step 716) and interpret the instructions present in the response output to take action to distribute the message to the participants.

[0078] In one example implementation, the input received by the computing unit 118 may further include a robot definition 704. The robot definition 704 may include a scenario selected from multiple scenarios, wherein the LLM 124 is configured to understand the scenario based on pre-training of the LLM 124, fine-tuning of the LLM 124 using additional training data, or incorporating additional data into cue words using RAG or similar methods. The robot definition 704 may also include the roles of potential participants (communicators) in the scenario and triggers including conditions such as responses received from one or more participants that allow the robot 602 to access.

[0079] The input received by the computing unit 118 may also include robot instantiation 706 for a specific session that robot 602 can facilitate. Robot instantiation 706 may include the identity of the initial participant, the communication endpoint associated with the participant (note that one or more participants may share a communication endpoint, as long as the messages sent or received can be associated with the participant's identity), the assignment of the initial participant to roles in the scene, and initial conditions for the scene, which may be received from the participants or external data sources, or inferred from the session, and may define information known about the scene.

[0080] The computing unit 118 can be configured to send and receive input from one or more of the following: the initial participant of the session using a communication endpoint assigned to the initial participant, additional participants identified by the robot 602 that can be added to the session or can join the session independently, and an external data source 708 that can be queried by the robot 602 or can publish scene-related information to the robot 602.

[0081] In one example implementation, participants in multi-party communication may include additional robots that can manage other scenarios and multiple groups of participants. Robots can send information to each other about internal sessions they are arbitrating. For example, one robot may be managing a session in the finance department, while another robot may be managing a session in the engineering department. If the finance department robot notices information that might be useful to the engineering department, it can access the session with the engineering department robot to share that information with the relevant participants.

[0082] Figure 8 This illustrates a method for enabling, according to an example embodiment, to act as Figure 6 and Figure 7The block diagram 800 illustrates an example sequence of steps for system access to multi-party communication for robot 602. In one example embodiment, robot 602 may include the system disclosed herein for providing multi-party communication using LLM. Figure 8 The sequence of steps shown is not intended to be limiting, as other sequences that enable robot 602 to join can be used to achieve a similar effect. For example, the scenario can be predefined or selected by the participants or an external data source, or the scenario can be inferred based on the participants' identities and / or initial sessions. Robot 602 can join a multi-party communication from the outset, or it can join later by invitation or triggered.

[0083] In the scenario selection step shown in box 802, the scenario can be pre-configured. Alternatively, the scenario can be selected later by participants in multi-party communication or by an external data source, or inferred based on the participant's identity and / or initial session.

[0084] In the robot configuration step shown in box 804, the robot can be configured using a robot definition. The robot definition can include a scene, a role, and a trigger.

[0085] In the session initiation step shown in box 806, a participant or external data source can initiate a session with one or more other initial participants within the context of the scenario.

[0086] In the robot instantiation step shown in box 808, the robot can join a session or be invited by a participant or an external data source. The robot can act as an intermediary for one, some, or all participants in a multi-party communication.

[0087] In the role assignment step shown in box 810, the bot can assign roles to participants based on their identity or relevant metadata. In one example implementation, participants can choose their roles, or roles can be assigned to participants by an external data source.

[0088] In the initial condition receiving step shown in box 812, the robot may receive initial conditions that define information known about the scene.

[0089] In the triggering robot step shown in box 814, the robot can be triggered, i.e., the robot can be brought in to participate in multi-party communication based on a trigger. The robot can be brought in to start a session with the participants. The robot can then wait until the participants respond. Alternatively, the robot can listen for messages and look for specific content within the messages as a trigger. Alternatively, the robot can wait until too much time has passed and the communicating party has still not responded. The robot can determine the amount of waiting time based on the scenario, or the amount of waiting time can be predetermined.

[0090] In the "Write Prompts and Send to LLM" step shown in box 816, the computational unit associated with the robot can use the following to write prompts: scene, roles, participant-to-role assignments, initial conditions, and any dialogue that has already occurred between the participants, including the robot, or input from external data sources. The computational unit can send the prompts to the LLM in the form of text input prompts.

[0091] In the step of writing a response and sending it to the computation unit, shown in block 818, the LLM can receive a text input prompt and return a response output with instructions for the computation unit. The LLM uses its scene- and world-related knowledge, typically acquired during pre-training and / or fine-tuning, to process the text input prompt and computes a series of lexical units as indicated by the text input prompt. This series of lexical units can represent the response output returned by the LLM to the computation unit. Specifically, the LLM can receive the prompt (as shown in block 820) and analyze the prompt (as shown in block 822). Based on the analysis of the prompt, the LLM can write a response output (as shown in block 824) and send the response output to the computation unit (as shown in block 826). The prompt and the flow of operations performed by the LLM are key elements of the system of this disclosure.

[0092] In the instruction-based action step shown in block 828, the computing unit may follow instructions provided in the response output received from the LLM. Based on these instructions, the computing unit may generate an output message and send it to the appropriate participant, or take other actions as instructed by the LLM.

[0093] This process is repeated as follows: The added content of the session is continuously reviewed against the triggering conditions until the session is terminated by one of the participants, an external data source, or a robot based on its understanding of the scene and dialogue, as shown in decision box 830. If the session ends in decision box 830, the process stops in box 832. If the session has not ended in decision box 830, the process returns to box 814.

[0094] Figure 9 This is a schematic diagram 900 illustrating key components of a prompt 902 written for and provided to an LLM according to an example implementation. The "Activity" component 904 of the prompt 902 may include settings defining activities that the LLM may be participating in. Activities may include sessions, workflows, projects, or any human activity that provides the LLM with additional context regarding expected content and response methods.

[0095] The “Number of Parties” component 906 can define the number of communicating parties in a multi-party communication. The number of communicating parties can be a fixed number, or it can indicate that communicating parties are expected to join or leave arbitrarily at different points in time or at turning points in the activity. The number of communicating parties can be implied by the “Scene and Roles”.

[0096] The “Scenario and Role” component 908 may include defining the roles or characters of the participants, including knowledge categories related to the scenario in which the participants can be expected to have information about them. For example, in a roadside service scenario, the driver is one of the roles. The cue word 902 may include a statement that the driver is expected to know their ETA. These additional knowledge categories that the participant can know may be explicit in the cue word 902 or may be implied by the scenario. For example, GPT-4® LLM, based on pre-training, is likely to infer that the driver in the roadside service scenario can know their ETA.

[0097] Component 910, "Robot Role in the Scenario (Robot Role)," may include defining the robot's role in a session occurring within a multi-party communication. For example, it may be anticipated that the robot engages only in bidirectional conversations between participants. In some example implementations, the robot may be anticipated as a facilitator that integrates into the multi-party conversation to help maintain the flow of the conversation. In one example implementation, it may be anticipated that the robot both facilitates multi-party conversations and engages in individual bidirectional or multi-party conversations with other participants. Additionally, it may be anticipated that the robot enables other participants, other robots, or other external data sources to integrate into the multi-party conversation as needed.

[0098] Component 912, “Restrictions to Prevent Illusions (Illusion Restrictions),” may include defining restrictions to prevent illusions. “Illusion” refers to an LLM providing one or more inaccurate or unreliable lexical units when it predicts a series of lexical units based on cue word 902 as a response output (see OpenAI (2023)). While the LLM’s response output may appear convincingly correct, the response is a trick of choosing the possible next word and is not based on any “real-world” information. To account for the potential for LLM to produce illusions, illusion restrictions include explicit guidance on providing alternative responses when information is unavailable. For example, in a roadside service scenario without any restrictions to prevent illusions, if the ETA is not already shared in the session, the LLM might respond with “45 minutes” when a customer asks “When is the ETA?”, since 45 minutes could be a typical arrival time based on pre-training or other training inputs. However, by instructing the LLM in cue word 902 to “take a specific action if asked for information that is currently unavailable in the session”, the LLM is guided to take that action instead of guessing.

[0099] Component 914, “Actions to be taken in restricted situations (illusionary actions),” can include defining specific actions to be taken in situations where information is missing, in order to keep the communicating parties informed and mitigate the illusion. Continuing with the roadside service scenario example above, cue word 902 could include the following text: “If asked for information that is not currently available in the session, draft a message to whichever communicating party you believe may have that information. If you believe that no party has the information, honestly admit that you do not have a response.” In this case, the illusionary action is to ask one of the other participants in the session for a response to the question. Illusionary actions, combined with knowledge categories explicitly identified or implied by the scenario and roles, enable the LLM to create messages directed to participants with specific roles to provide information that might otherwise be unreliably generated (creating an illusion).

[0100] Component 916, “Syntax Structure of Incoming Messages from Computing Units (Incoming Syntax Structure),” may include defining the syntax structure of incoming messages. Instructions in prompt word 902 may allow the LLM to precisely interpret previous messages in the session, including the role of the participant who spoke the message in the session. For example, instructions regarding the syntax structure of incoming messages may include: “Messages from customers begin with ‘Customer:’, messages from drivers begin with ‘Driver:’, and messages from employees begin with ‘Employee:’.”

[0101] In another implementation, the “syntax structure of incoming messages from computing units (incoming syntax structure)” component 916 can be defined more precisely using JSON, XML, or other data exchange formats.

[0102] Component 918, “Syntax structure of the response to be sent to the computing unit (response syntax structure),” may include a syntax structure defining the output to be read by the computing unit (i.e., the LLM’s response output). These instructions in prompt 902 can guide the LLM to precisely format the response output, enabling the computing unit to deterministically generate the output message and take action. For example, prompt 902 regarding the syntax structure of the output message may include “When you draft a message, make sure it begins with ‘I to the driver:’, ‘I to the customer:’, or ‘I to the employee:’, depending on who you wish to send the message to.”

[0103] In another implementation, component 918, “syntax structure of the response to be sent to the computing unit (response syntax structure)”, can be defined more precisely using JSON, XML, or other data exchange formats.

[0104] The "Previous Messages" component 922 may include messages sent up to this point in time, including messages from the robot. The "Previous Messages" component 922 may also include additional information about the scenario, which may be known to one or more of the participants or may be known based on access to an external data source. By recreating the cue word 902 with all previous messages in the session, the robot constrains the LLM to only provide the next set of messages in the session and defines what is known and unknown about the session at this point in time. Previous messages may be provided following the syntax structure defined in the "Syntax Structure of Incoming Messages from the Computation Unit (Incoming Syntax Structure)" component 916.

[0105] Some or all of these components may be included in cue word 902 to ensure correct output (i.e., the LLM's response output) within the context of the scenario. Additional elements may be included in cue word 902 to guide the LLM in further activities or to provide additional focus on the response output provided by the LLM. Additional elements may include input confirmation component 920, which may include expectations and (potentially) requirements for the input message that confirms the response output provided by the LLM. Additional elements may also include instructional examples for the LLM (i.e., “few-sample” cue words), which include potential examples from previous sessions to further guide the LLM's behavior.

[0106] Figure 10 A schematic diagram 1000 illustrates the various components of a prompt 902 according to an example implementation. The components of prompt 902 can be provided using natural language as text input prompts. Example prompt 1002 for the “Activity” component 904 could include “Today, you will help facilitate the conversation.” Example prompt 1004 for the “Number of Parties” component 906 could include “Among three communicating parties.” Example prompt 1006 for the “Scene and Roles” component 908 could be “Trailer dispatcher, trailer driver, and customer requiring assistance.” Example prompt 1008 for the “Robot’s Role in the Scene (Robot Role)” component 910 could be “They will each communicate with you individually, but not with each other. That is, you are the ‘middleman’ in the conversation. Your job is to deliver any information that each party needs as it appears in the conversation.” Example prompt 1010 for the “Prevention of Illusion Restrictions (Illusion Restrictions)” component 912 could include “Or, if a communicating party requests information that is currently unavailable in the conversation.”

[0107] Furthermore, example prompt 1012 for component 914, “Action to be taken under restricted conditions (illusory action)”, could include “Draft a message to whichever communication party you believe will have the information. If you believe no communication party has the information, then honestly admit that you do not have a reply.” Example prompt 1014 for component 916, “Syntax structure of incoming messages from the computing unit (incoming syntax structure)”, could be: “Messages from the customer begin with ‘Customer:’, messages from the driver begin with ‘Driver:’, and messages from the employee begin with ‘Employee:’.”

[0108] Example prompt 1016 for component 918, “Syntax structure of the response to be sent to the computing unit (response syntax structure),” could include “When you draft a message, make sure it begins with ‘I to the driver:’, ‘I to the customer’, or ‘I to the employee:’, depending on who you want to send the message to. If you don’t think you need to draft a message to anyone, just use ‘%%%’ for the response.” Example prompt 1018 for component 920, “Respond to more than one party at a time, with messages separated by newlines. Be sure to do this whenever a customer provides you with new information or asks a question.” Example prompt 1020 for component 922, “Previous messages,” could include “Note that you may have already responded in the session. Here is the session so far:…”

[0109] Depending on the input message, the LLM may respond, or it may not respond and suggest an alternative action. In one example implementation, if a participant responds with "Thank you!" when a response is not required, the LLM can be trained or instructed to notice that a response is not necessary.

[0110] In one example implementation, when it is a response that does not require interaction, if a participant asks a question in the prompt that an answer is available, the LLM can respond directly to the participant without involving any of the other participants.

[0111] In an example implementation of response and information delivery, if a participant provides information that can be related to other participants (i.e., other participants do not have the information), the LLM can use the information provided by that participant to write communications to those other participants.

[0112] In an example implementation of responses and queries, if a participant seeks information that is unavailable in the prompt, the LLM can identify which other participant is most likely to be able to respond to that participant and compose a communication to that other participant to seek the information. The LLM can also notify the participant that the LLM has asked the question and periodically update the participant on the status of the question. For example, if the driver has not responded after a set amount of time, the LLM can notify the participant. In an example implementation of responses where no one might know the information at this time, if the participant seeks information that none of the current participants might be able to provide, the LLM can respond by stating that no participant in the session might know the information.

[0113] In one example implementation, the system can include timestamps on previous messages sent by multiple communicating parties and by the LLM, and can include the current timestamp in the input prompt. The LLM can use this information to reason about questions or issues raised by one or more of the communicating parties. For example, if a driver sends a message at 2:00 PM with an ETA of 20 minutes, and later at 2:15 PM, a customer asks, “Where is my package? When will it arrive?”, the LLM can respond, “According to the driver’s ETA, the package should arrive within five minutes.” However, if the customer asks the same question at 3:00 PM, the LLM can respond, “It looks like the driver is late. Let me ask the driver when they expect to deliver your package.”

[0114] The systems and methods disclosed herein also consider additional features and enhancements. Specifically, the LLM can also be trained / tuned using typical multi-party session data from specific scenarios to better understand how and when to respond. In some example implementations, cue words may include new information published to the computation unit from an external data source. For example, the external data source may periodically update ETAs or role assignments. Cue words can also be enhanced using RAG or similar techniques, leveraging examples from particularly good previous sessions. The computation unit can be configured to treat this new information as a trigger for the LLM to write messages to the appropriate participants.

[0115] In some example implementations, the system can enable the automatic detection of escalating problems or conflicts (utilizing the ability to route such problems to designated human oversight bodies or advanced resolution mechanisms). In one example implementation, the system may also include an implementation that allows for customizable privacy settings for each communicating party to control the extent of information sharing. The system can be integrated with various communication platforms, including voice, text, and video.

[0116] In some example implementations, the authorized party may be able to establish direct communication lines between any or all communicating parties in a session, including role-based access control to define each participant’s permissions and control levels for multi-party communication.

[0117] In one example implementation, the response output from the LLM may be an original message drafted by the LLM that incorporates information from the session, or a direct quote from one or more communicating parties (if direct communication is enabled by the authorized party).

[0118] In some example implementations, the system can enable the management of multiple parallel communication threads for simultaneously discussing multiple issues or topics with all or a subset of multiple communicating parties.

[0119] In some example implementations, the system can enable the dynamic creation, modification, or disbanding of communication groups based on the current context, question, or requirement. For instance, the LLM can determine, based on prompt words, that a new participant should be added to the session and instruct the computing unit to contact that new participant using their communication medium. If an answer to a question requires input from someone else not yet included in the scenario, the system can compose a message to that other person or system.

[0120] In some example implementations, the system can be adapted to scales that manage the entire organization as a “multi-party session.” For example, each department or team can be viewed as a communication group, where the system manages the flow of information within the organization as a whole.

[0121] As the volume of communication for any given participant that can be involved in multiple simultaneous sessions increases, prioritizing communications to that participant can be useful. The system can be configured to manage all sessions for a given participant, respond to sessions on behalf of a participant, and prioritize requests based on the urgency or importance of those requests to that participant. In one example implementation, the system can represent multiple participants sharing a single role in multiple simultaneous sessions, and thereby distribute requests to be responded to among these participants based on the availability and expertise of individual participants.

[0122] In some example implementations, the system's computing unit can be configured to perform real-time progress tracking to monitor the status of the ongoing problem and the steps taken.

[0123] At the role assignment step, when participants are identified and assigned roles, the system can be informed of their language preferences, or the system can analyze the content of previous messages from participants to infer their language preferences. Cue words for the LLM can include language preferences, allowing the LLM to analyze and respond to individuals in their preferred languages. For example, the system could support conversations between Mandarin-speaking customers, English-speaking dispatchers, and Spanish-speaking drivers.

[0124] In many multi-party asynchronous sessions, there can be multiple communicating parties capable of responding to queries raised by one of the participants. In this case, the first participant can join the session, but if the first participant does not respond in a timely manner, a second participant can join the session. The prompt can be configured to require the LLM to return a request for information from several participants, where all participants are simultaneously queried for the information; or each participant's requests for information can be sequential, with a time delay instruction so that the computation unit sends subsequent requests only if the first request is not responded to within the time limit; and other participants can be notified that the request is being sent to them.

[0125] In another implementation, the prompt can be configured to instruct the LLM to raise questions with multiple parties and synthesize the "best" response based on all collected responses. Furthermore, the prompt can be configured to instruct the LLM to gather information from external sources, including the internet or internally available data sources, and to provide accurate citations.

[0126] Figure 11 An example dialogue 1102 supported by a system for providing multi-party communication using an LLM is shown according to an example implementation. The example dialogue 1102 may include an input message 1104, a text input prompt 1106 to be sent to the LLM, and a response output 1108 received from the LLM.

[0127] In example dialogue 1102, the computing unit generates an initial message (shown as first input message 1110) based on a template and using information about the service available from an external data source to begin the session and generate the initial conditions for the setup. First input message 1110 can be sent by the system acting as the robot. First input message 1110 is identified as having an "employee" role. The first part of first input message 1110 can be, for example, as follows: "Employee: Hello Deanna, this is the dispatch office. It's a pleasure to assist with your service request today. Our driver will meet you at 1801, 30th Street, Boulder 80301 to provide battery service. Please reply to this text if you have any questions or need assistance. If you do not wish to receive texts, please reply STOP. Thank you!"

[0128] Input from the customer is what enables the robot to connect and... Figure 10 The complete prompt word shown, along with any messages received up to that point in the session, is passed to the LLM.

[0129] The customer can respond using two messages. The second part of the first input message 1110 can include, for example, the following two messages received from the customer: "Customer: When can I expect them to arrive?" and "Customer: I am at 1821 30th Street".

[0130] The calculation unit can pause before processing input from the customer to check if the customer has any additional input. Upon receiving input from the customer, the calculation unit can generate a first text input prompt 1112 to be sent to the LLM. The first text input prompt 1112 may include, for example, the following text: Today, you will help facilitate a conversation between three parties: a trailer dispatcher, a trailer driver, and a customer who needs assistance. They will each communicate with you, but not with each other. That is, you are the "middleman" in the conversation. Your job is to deliver any information that each party needs to the conversation as it becomes available. Alternatively, if a party requests information that is not currently available in the conversation, draft a message to any party you believe will have that information. If you believe no party has the information, honestly admit that you do not have a response. Messages from the customer begin with "Customer:", messages from the driver begin with "Driver:", and messages from the staff begin with "Staff:". When drafting a message, ensure it begins with "I am writing to the driver..." Start with ":", "I'm to the customer:", or "I'm to the employee:", depending on who you want to send the message to. If you don't think you need to draft a message to anyone, just use "%%%" to respond. You can respond to more than one party at a time, with messages separated by newlines. Always do this whenever a customer provides you with new information or asks a question. For example, if a customer asks about the driver's ETA, you might want to draft a message to the driver and a message to the customer informing them that you've asked the driver how long it will take. That is, a message starts with "I'm to the driver:", followed by a newline, followed by a message starting with "I'm to the customer:". Note that you may already be responding in the conversation. Here is the conversation so far: Staff: Hello Deanna, this is the dispatch office. It's a pleasure to assist with your service request today. Our driver will meet you at 1801, 30th Street, Boulder 80301 to provide battery service. Please reply to this text if you have any questions or need further assistance. If you do not wish to receive texts, please reply STOP. Thank you!

[0131] Customer: When can I expect them to arrive?

[0132] Customer: I am at 1821 30th Street.

[0133] The LLM can receive and analyze the first text input prompt 1112, and can generate the first response output 1114 based on the following rules: - First, using the incoming syntax structure, LLM can analyze example dialogue 1102 with messages received before this moment; Next, the illusionary constraint in the first text input prompt 1112 instructs the LLM not to simply fabricate a response and react to the customer. Without the illusionary constraint, the LLM would likely directly reply to the customer with "45 minutes," as this is a commonly predicted response for an LLM trained based on historical conversations; - Conversely, following the restricted actions, the LLM, based on its understanding of the scenario and the roles and based on the training data, determines that the driver is a participant in the session who can have the information requested by the customer, and writes instructions to the computing unit to send messages to the driver. - Following the response syntax structure, the first response output 1114 begins with the syntactic instruction "I to the driver:" to instruct the computing unit to send a message to the participant playing the role of the driver and request an update to the ETA. Based on the scenario and role, the message to be sent to the driver is written in the dispatcher's tone based on training data. The first part of the first response output 1114 may include a message to the driver and may be as follows: "I to the driver: Hello, this is the dispatch office. The customer is at 1821 30th Street. Please update your ETA."; - If the first text input prompt 1112 has already instructed the driver to speak a different language, then the message to the driver can be in the driver's language. The LLM can also interpret the driver's response in the driver's own language and then respond to those participants in the languages ​​chosen by other participants; - Following the input confirmation instruction, the LLM also generates a confirmation for the customer, causing its response to begin with "I to Customer:". The second part of the first response output 1114 can include a message for the customer, and can be as follows: "I to Customer: Hi Deanna, I have inquired with the driver about the ETA. We will update you once we have more information." The computing unit receives a first response output 1114, which includes messages for the driver and messages for the customer. The computing unit can use the communication medium of the participants identified as the driver and customer in the session to send messages to each participant.

[0134] The participant identified as the driver responds using messages. The message from the driver is the second input message 1116, and it acts as a trigger for the computational unit to compose a second text input prompt 1118 to be sent to the LLM. An example of the second input message 1116 could be: “Driver: I’ll be there in about 30 minutes.” The second text input prompt 1118 can fully include the first text input prompt 1112, the previous message previously composed by the LLM in the first response output 1114, and the second input message 1116 received from the driver.

[0135] LLM analyzes the second text input prompt word 1118 and generates a second response output 1120 based on the following rules: - Using the incoming syntax structure, LLM analyzes the conversation up to this point, including the first response output 1114 generated by LLM itself; - The LLM follows the robot's role and writes a second response output 1120, which includes a response directed to the customer with new information provided by the driver. The second response output 1120 may include the following text: "I am writing to the customer: Hello Deanna, the driver estimates they will arrive in approximately 30 minutes." - LLM is based on training data, uses the tone of dispatcher-customer communication, and translates the text "Get there in about 30 minutes" into a more accurate response; - Following the response syntax structure, the second response output 1120 begins with the syntax instruction "I to the client:" to instruct the computing unit to send a message with the updated ETA to the participant playing the role of the client.

[0136] The computing unit receives the second response output 1120 and uses the client's communication medium to send messages to participants in the session identified as clients.

[0137] Participants identified as customers in the session respond using messages. The customer's message is the third input message 1122 and acts as a trigger for the computing unit to compose a third text input prompt 1124 to be sent to the LLM. The third input message 1122 could be, for example, "Customer: Okay, you can tell him it's actually the same address, except for 31st Street." The third text input prompt 1124 generated by the computing unit for the LLM fully includes the second text input prompt 1118, the second response output 1120 previously generated by the LLM, and the third input message 1122 received from the customer.

[0138] The LLM can receive and analyze the third text input prompt 1124, and generate a third response output 1126 based on the following rules: - Using the incoming syntax structure, LLM analyzes the conversation up to this point in the session, including the response output generated by LLM; - The LLM follows the robot's role and writes the first part of the third response output 1126, which includes a message directed to the driver with new information provided by the customer. The first part of the third response output 1126 could be, for example, as follows: "To the driver: Hello, the customer has updated their address to 1821 31st Street. Please update your ETA." Based on its training in natural language and its interpretation of illusionary limitations, the LLM identifies that it has sufficient information to interpret "same address, except for 31st Street" as "1821 31st Street". If the LLM does not have prior information about the address in the second text input prompt 1118, the LLM will ask one of the participants for clarification. -Based on training in this scenario, the LLM infers that a change in address may require a change in ETA and requests this additional information from the driver; - Following the response syntax structure, the third response output 1126 begins with the syntax instruction "I to the driver:" to instruct the computing unit to send a message with the updated address to the participant playing the role of the driver; Following the input confirmation instruction, the LLM also generates a second part of the third response output 1126, which includes confirmation to the customer, causing its response to begin with "I to Customer:". The second part of the third response output 1126 may include, for example, the following text: "I to Customer: Hello Dear Deanna, I have asked the driver to update their ETA based on the new address. We will update you once we have more information." The computing unit can receive the third response output 1126 and send messages to each participant via the communication medium of the participants identified as the driver and the customer in the session.

[0139] The participant identified as an employee during the session responds using a fourth input message 1128, which acts as a trigger for composing a fourth text input prompt 1130 to be sent to the LLM for the computing unit. The fourth input message 1128 may include, for example, the following text: “Employee: The trailer company informs Deanna that there is now a new driver, Doug.”

[0140] The fourth text input prompt 1130 generated by the computing unit 118 fully includes the third text input prompt 1124, the third response output 1126 previously generated by the LLM, and the fourth input message 1128 received from the employee.

[0141] The LLM receives and analyzes the fourth text input prompt 1130, and generates the fourth response output 1132 based on the following rules: - Using the incoming syntax, LLM analyzes the conversation in the session that has received all messages so far, including the response output generated by LLM; - The LLM follows the robot's role and writes a fourth response output 1132, which includes a response directed to the customer with new information provided by the employee. The fourth response output 1132 could include, for example, the following text: "I am writing to the customer: Hello Deanna, we have a new driver, Doug, who will assist you. We will update you as soon as we have more information." -Based on training data, LLM can identify that information about new drivers should be shared with customers and that the driver's role is to "assist you"; - Following the response syntax structure, the fourth response output 1132 begins with the syntax instruction "I to the client:" to instruct the computing unit to send a message with updated information to the participant playing the role of the client.

[0142] The computing unit receives the fourth response output 1132 and uses the client's communication medium to send messages to participants in the session who are identified as clients.

[0143] A participant identified as a customer in the session can respond using a fifth input message 1134, which acts as a trigger for composing a fifth text input prompt 1136 to be sent to the LLM for the computing unit. The fifth input message 1134 may include, for example, the following text: “Customer: Ugh… how long will it take now?>:(”. The fifth text input prompt 1136 may fully include the fourth text input prompt 1130, the fourth response output 1132 generated by the LLM, and the fifth input message 1134 received from the customer.

[0144] The LLM receives and analyzes the fifth text input prompt 1136, and generates the fifth response output 1138 based on the following rules: - Using the incoming syntax structure, LLM analyzes the conversation in the session with the messages received so far, including the response output generated by LLM; - The LLM follows the robot's role and writes the first part of the fifth response output 1138, which includes a response directed to the driver with a request for the ETA from the customer. The first part of the fifth response output 1138 may include, for example, the following text: "I am to the driver: Hello Doug, the customer has inquired about the ETA. Please update your ETA with us."; - Based on training data, the LLM will rephrase the fifth input message 1134, which includes emojis, received from the customer, into a request for the ETA. In this case, the request protects the driver from the customer's frustration, which can help maintain the morale of the participants, who might otherwise face abusive language from other participants; - Following the response syntax structure, the first part of the fifth response output 1138 begins with the syntax instruction "I to the driver:" to instruct the computing unit to send a message with updated information to the participant playing the role of the driver; Following the input confirmation instruction, the LLM also generates a second part of the fifth response output 1138, which includes confirmation to the customer, causing its response to begin with "I to Customer:". The second part of the fifth response output 1138 may include, for example, the following text: "I to Customer: Hello Dear Deanna, I have inquired with the driver about the ETA. We will update you once we have more information." The computing unit can receive the fifth response output 1138 and use the communication medium of the participants identified as the driver and customer in the session to send messages to each participant.

[0145] The participant identified as the driver in the session can respond using a sixth input message 1140, which acts as a trigger for the computational unit to compose a sixth text input prompt 1142 to be sent to the LLM. The sixth input message 1140 may include, for example, the following text: “Driver: I am about to arrive.” The sixth text input prompt 1142 may fully include the fifth text input prompt 1136, the fifth response output 1138 previously composed by the LLM, and the sixth input message 1140 received from the driver.

[0146] The LLM can receive and analyze the sixth text input prompt 1142 to generate a sixth response output 1144 based on the following rules: - Using the incoming syntax structure, LLM analyzes the conversation in the session with all messages received so far, including the responses generated by LLM; - The LLM follows the robot's role and writes a sixth response output 1144, which includes a response directed to the customer with new information provided by the driver. The sixth response output 1144 could include, for example, the following text: "Hello Deanna, the driver will be arriving now."; - Following the response syntax structure, the sixth response output 1144 begins with the syntax instruction "I to the client:" to instruct the computing unit to send a message with the updated ETA to the participant playing the role of the client.

[0147] The computing unit receives the sixth response output 1144 and uses the client's communication medium to send messages to participants in the session who are identified as clients.

[0148] Figure 12 This is a flowchart of a method 1200 for providing multi-party communication using an LLM, according to an example implementation. In some implementations, operations may be combined, performed in parallel, or performed in a different order. Method 1200 may also include additional or fewer operations compared to those shown. Method 1200 may be executed by processing logic, which may include hardware (e.g., decision logic, dedicated logic, programmable logic, and microcode), software (such as software running on a general-purpose computer system or a dedicated machine), or a combination of both. In an example implementation, the operation of method 1200 may be performed by... Figure 1 The computation unit 118 and / or AI unit 122 shown are executed.

[0149] Method 1200 may begin in block 1202 by receiving input messages from user devices associated with multiple communicating parties of an instance participating in the scenario. The LLM may be configured to understand the scenario. In one example implementation, the input messages may originate from multiple sources. Multiple sources in Figure 7 The data source is shown as external data source 708. Multiple sources may include one or more of the following: data events, machines, sensors, additional robots, and other sources (such as a communication party among multiple communication parties and a database or system accessible to an enterprise associated with multiple communication parties).

[0150] In box 1204, method 1200 may continue with generating text input prompts for the LLM based on a scenario template. The scenario template may include rules for organizing input messages based on the scenario, as shown in box 1206. Multiple communicating parties and the LLM may be assigned roles as participants in the scenario. The scenario template may also include rules for instructing the LLM, as shown in box 1208, to analyze the text input prompts and compute the LLM's response output, which is defined by the LLM's role in the scenario.

[0151] In one example implementation, the scenario template may include an LLM and activities to be participated in by multiple communicating parties. The scenario template may also include the number of communicating parties among the multiple communicating parties. In some example implementations, the scenario template may also include references to a scenario. The scenario can provide context for the activities. The LLM may be configured to understand the scenario based on at least one of the following: pre-training the LLM, fine-tuning the LLM using additional training data, and incorporating additional data into the prompt words using RAG or a similar method. The scenario template may also include roles for each of the multiple communicating parties, which are assigned based on the scenario. The scenario template may also include roles for the LLM, which are assigned based on the scenario and the activities. In some example implementations, the scenario template may also include a first syntactic structure and a second syntactic structure. The first syntactic structure can be used to construct input messages from the multiple communicating parties and to assign input messages based on roles in the scenario. The second syntactic structure can be used by the LLM to construct response outputs to identify the content of the output message to be sent on behalf of the LLM and the intended recipient.

[0152] In some example implementations, the scenario template may include additional rules regarding the actions to be taken by the LLM when it is requested to provide information not included in the text input prompt. These additional rules may include generating a message to request information from one of the multiple communicating parties. The information may be expected to be available from that communicating party based on its role or rules within the scenario template.

[0153] In some example implementations, additional rules may include generating a message to request the information from an external data source. External data sources may include one or more of the following: additional machines, sensors, robots, and other sources (such as a database or system accessible to a communicator among multiple communicators or an enterprise associated with that communicator among multiple communicators). The communicator among multiple communicators may be expected to possess the information based on rules in the communicator's role or scenario template.

[0154] In one example implementation, the scenario template may include examples from previous sessions to guide the LLM on how to respond.

[0155] In box 1210, method 1200 may proceed by sending a text input prompt to the LLM. The LLM can analyze the text input prompt to calculate the LLM's response output.

[0156] In block 1212, method 1200 may include analyzing the response output of the LLM. This analysis may be performed based on the syntax structure specified for the LLM in the scenario template. Based on the analysis of the LLM's response output, output messages to be sent to one or more of a plurality of communicating parties may be generated. The one or more communicating parties may be identified in the LLM's response output by the roles of those one or more communicating parties.

[0157] In block 1214, method 1200 may continue to perform the following operation: sending output messages from the communication endpoint representing the LLM to the user device associated with one or more of the multiple communication parties, based on the role of the LLM in the scenario.

[0158] Sending an output message to a user device associated with one or more of a plurality of communicating parties can result in an automatic notification to that party regarding the receipt of a new message. Depending on the message content, context, and the role of the receiving party, it can be expected that the receiving party will either follow up with a response or not. For example, if the output response to the receiving party is "Thank you," no follow-up is needed. If the output response to the receiving party is "Please provide the updated ETA for delivery," then follow-up is required. To fully understand whether follow-up is needed, all previous messages must be reviewed to identify any unresolved expectations from the previous message loop.

[0159] In one example implementation, after block 1214, the method may continue to deliver previous messages (including messages from the LLM) to a separate LLM configured with tagged prompts to determine when follow-up is expected from one or more of the multiple communicating parties.

[0160] Figure 13 A schematic diagram 1302 illustrates the various components of a tagged prompt word 1304 according to an example embodiment. The tagged prompt word is used... Figure 9The prompt 902 shown contains many of the same components. Component 1306 for the “Activity” component 1308 in prompt 1304 could include “Today, you will help with a conversation about delivery.” An example prompt 1304 for the “Scenario and Roles” component 1312 could be “The people in the conversation are the customer receiving the delivery, the delivery dispatcher, the delivery driver, and the assistant mediating between them,” where the “Number of Parties” component 1310 is implied by the number of parties listed. Note that “assistant mediating between them” in this list refers to the robot’s role in the conversation to distinguish messages sent by the robot in the conversation. In this context, example prompt 1304 for component 1314, "Robot's Role in the Scene," refers to the role of the tagged robot, which could be, "You are Angie, a customer service expert with 30 years of experience. Your job is to use your expertise to determine if anyone is waiting for a response from anyone else." Example prompt 1304 for component 1318, "Syntax Structure of the Response to be Sent to the Computing Unit (Response Syntax Structure)," could be, "Is anyone currently waiting for a response from anyone else? Please reply using a JSON dictionary containing the keywords 'customer,' 'driver,' and 'dispatcher,' where each keyword corresponds to a Boolean value that is true if anyone is waiting for a response from that party, and false otherwise. Do not reply using any text other than a dictionary."

[0161] Method 1200 may include repeating the operation in an additional loop each time an additional input message with a previous message is received from any of the multiple communicating parties. The previous message may include messages from the LLM, which are included as input messages in additional text input prompts following a scene template. Repeating the operation in an additional loop may be triggered based on one or more of the following: an algorithm, an external data source, LLM analysis of the input message to determine whether the LLM should enter an additional loop, etc.

[0162] In one example implementation, the LLM can be configured to acknowledge receipt of a message from one of a plurality of communicating parties.

[0163] In some implementations, an LLM can be configured to send instructions and messages to one or more external systems. Additionally, an LLM can be configured to perform actions or trigger workflows in one or more external systems, or modify or switch its own prompts. Furthermore, an LLM can be configured to determine whether it has collected the necessary input to perform an action or trigger a workflow. Based on analysis of sessions associated with multi-party communication, an LLM can be configured to estimate the certainty of the action to be taken and evaluate that certainty against a threshold set by the organization associated with one or more external systems. When the certainty is above the threshold, the LLM can perform an action or trigger a workflow based on the necessary input collected from the multi-party communication. When the certainty is below the threshold, the LLM can send a message to the appropriate party among the multiple communicating parties for confirmation or instructions on how to proceed.

[0164] In one example implementation, the LLM can be configured to generate code in a computer language associated with one or more external systems. The LLM can use this code to send instructions to one or more external systems, send messages to one or more external systems, perform actions on one or more external systems, trigger workflows on one or more external systems, and perform other operations associated with one or more external systems.

[0165] In one example implementation, the LLM can be configured to understand the language of one or more of a plurality of communicating parties. The LLM can communicate with one or more communicating parties in a language selected by one or more of the communicating parties in a multi-party communication. Additionally, the LLM can automatically translate messages from a first communicating party (one of the one or more communicating parties) from a first language selected by the first communicating party to a second language selected by a second communicating party (one of the one or more communicating parties).

[0166] In one example implementation, the LLM can be configured to identify another communication party among one or more communication parties that speaks a language different from the language in question, and to suggest on behalf of the third communication party that the other communication party translate the other communication party's message into the other language.

[0167] In one example implementation, an LLM can be configured to coordinate multiple conversations of one or more subsets of multiple communicating parties as part of the same session.

[0168] In one example implementation, the input message can be associated with certain modalities, such as text, images, audio, video, location, and any other data type or format. A modality can refer to a data category defined by how the data is received, represented, and understood. For example, a customer might take a picture of a damaged package at the front door and send it to the system along with the message "This is unacceptable!". The LLM can receive the message and the picture, combine the text and image of the message to understand that it is a picture of a damaged package, and based on the analysis of the input, provide assistance to the customer or agent to resolve the delivery problem, such as by issuing a refund.

[0169] In one example implementation, an LLM can be configured to write response outputs with modalities corresponding to the modalities of the input messages.

[0170] Furthermore, the LLM can be configured to send messages to the system that control the notification of additional messages from one of the multiple communication parties to another of the multiple communication parties. Additionally, the LLM can be configured to prevent notification of additional incoming messages from one or more of the other communication parties when their attention is not required. In one example implementation, the LLM can be configured to send notification of additional incoming messages to one or more of the multiple communication parties. The notification may optionally include a measure of the urgency of the expected response from the one or more other communication parties.

[0171] In one example implementation, an LLM can be configured to identify a subset of multiple communicating parties that are communicating directly with each other. Based on this identification, the LLM can decide not to interfere with the session. Specifically, the LLM can avoid interfering with multi-party communication between subsets of multiple communicating parties.

[0172] Method 1200 may also include instructing the LLM to add a communicator to multiple communicators in a text input prompt. A communicator may be added based on one or more of the following: a turning point in the scenario, an external event, a request from another communicator, a request for information not available in the text input prompt or among the communicators, a requirement provided in the scenario and inferred by the LLM, etc. Based on the text input prompt, a message can be sent to the system controlling the multiple communicators, inviting the communicator to join the multiple communicators.

[0173] Method 1200 may further include instructing the LLM to remove a communicator from a plurality of communicators via a text input prompt. Removal of a communicator may be performed based on one or more of the following: a turning point in the scenario, an external event, a request from another communicator, a request from the communicator to be removed, a requirement provided in the scenario and inferred by the LLM, etc. Based on the text input prompt, a message may be sent to the system controlling the plurality of communicators, causing the communicator to be removed from the plurality of communicators.

[0174] Figure 14 This is a flowchart of a method 1400 for providing multi-party communication using an LLM, according to an example implementation. In some implementations, operations may be combined, performed in parallel, or performed in a different order. Method 1400 may also include additional or fewer operations compared to those shown. Method 1400 may be executed by processing logic, which may include hardware (e.g., decision logic, dedicated logic, programmable logic, and microcode), software (such as software running on a general-purpose computer system or a dedicated machine), or a combination of both. In an example implementation, the operation of method 1400 may be performed by... Figure 1 The computation unit 118 and / or AI unit 122 shown are executed.

[0175] Method 1400 may begin in block 1402 by receiving input messages from user devices associated with multiple communicating parties involved in the scenario. In block 1404, method 1400 may include receiving scenario-related data. Scenario-related data may include data relating to: the circumstances associated with the multi-party communication (e.g., the type of service to be provided), the goal of the multi-party communication, the participants in the multi-party communication, examples of previous sessions of that type, etc.

[0176] In block 1406, method 1400 continues by receiving context associated with the multi-party communication. In one example implementation, the context may be received from a TMS / CRM system associated with one of the multiple communicating parties. The context may include previous messages associated with previous sessions between any of the communicating parties, business rules of the service provider or the communicating party providing services on behalf of the service provider, order details, the time of request for service, etc.

[0177] In block 1408, method 1400 may proceed by generating text input prompts for the LLM based on the input message, scene-related data, and context. The text input prompts may be generated based on a scene template. The scene template may include rules for organizing the input message based on the scene. Multiple communicating parties and the LLM may be assigned roles as participants in the scene. The scene template may include rules instructing the LLM to analyze the text input prompts and calculate its response output, defined by the LLM's role in the scene. In one example implementation, the rules instructing the LLM to analyze the text input prompts and calculate its response output may include multi-party communication questions, such as to whom to respond, how to respond, how to update, and who needs to follow up.

[0178] In block 1410, method 1400 may proceed by sending a text input prompt to the LLM. The LLM may analyze the text input prompt to compose a response output for the LLM. The response output may include one or more messages to be sent to one or more of a plurality of communicating parties. The response output may also include multi-party communication instructions. Multi-party communication instructions may include responses to one or more multi-party communication queries.

[0179] In box 1412, method 1400 can perform the following operations: analyze the LLM's response output based on the syntax structure specified for the LLM in the scenario template. Based on the analysis of the response output and one or more messages provided in the response output, output messages to be sent to one or more communication parties can be generated. The one or more communication parties can be identified in the LLM's response output by the roles of those one or more communication parties.

[0180] In block 1414, method 1400 may perform the following operation: adding a tag to at least one output message in the output messages based on a multi-party communication instruction. In block 1416, method 1400 may further perform the following operation: sending the output message from the communication endpoint representing the LLM to a user device associated with one or more of the multiple communication parties, based on the multi-party communication instruction and the LLM's role in the scene. The tag added to at least one output message in the output messages may notify at least one of the multiple communication parties of the receipt of the at least one output message.

[0181] In block 1418, method 1400 may proceed to update the context associated with the multi-party communication based on multi-party communication instructions. In one example implementation, updating the context may include updating information associated with the multi-party communication in a TMS / CRM system associated with one of the multiple communicating parties.

[0182] Figure 15 This is a high-level block diagram illustrating an example computer system 1500 within which a set of instructions for causing a machine to perform any or more of the methods discussed herein can be executed. Computer system 1500 may include, refers to, or is a component of one or more devices of various types, such as general-purpose computers, desktop computers, laptop computers, tablet computers, netbooks, mobile phones, smartphones, personal digital assistants, smart television devices, and servers. In some embodiments, computer system 1500 is an example of a system for providing multi-party communication using an LLM. It is noteworthy that... Figure 15 Only one example of computer system 1500 is shown, and in some embodiments, computer system 1500 may have the same... Figure 15 The number of components / modules shown is less than that of the previous ones. Figure 15 The diagram shows a greater number of components / modules compared to the previous one.

[0183] Computer system 1500 may include one or more processors 1502, memory 1504, one or more mass storage devices 1506, one or more input devices 1508, one or more output devices 1510, and a network interface 1512. In some examples, processor 1502 is configured to perform functions and / or process instructions for execution within computer system 1500. For example, processor 1502 may process instructions stored in memory 1504 and / or instructions stored on mass storage device 1506. Such instructions may include components of operating system 1515 or software applications 1516. Computer system 1500 may also include components not in memory 1504. Figure 15 One or more additional components are shown in the diagram.

[0184] According to one example, memory 1504 is configured to store information within computer system 1500 during operation. In some example implementations, memory 1504 may refer to a non-transitory computer-readable storage medium or a computer-readable storage device. In some examples, memory 1504 is transient memory, meaning that the primary purpose of memory 1504 may not be long-term storage. Memory 1504 may also refer to volatile memory, meaning that memory 1504 does not retain its stored contents when memory 1504 is not receiving power. Examples of volatile memory include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), and other forms of volatile memory known in the art. In some examples, memory 1504 is used to store program instructions for execution by processor 1502. In one example, memory 1504 is used by software (e.g., operating system 1514 or software application 1516). Typically, software application 1516 refers to a software application suitable for implementing at least some of the operations of the method for providing multi-party communication using LLM as described herein.

[0185] Mass storage device 1506 may include one or more transient or non-transitory computer-readable storage media and / or computer-readable storage devices. In some embodiments, mass storage device 1506 may be configured to store a larger amount of information than memory 1504. Mass storage device 1506 may also be configured for long-term storage of information. In some examples, mass storage device 1506 includes non-volatile storage elements. Examples of such non-volatile storage elements include magnetic hard disks, optical disks, solid-state drives, flash memory, various forms of electrically programmable memory (EPROM) or electrically erasable programmable memory, and other forms of non-volatile memory known in the art.

[0186] In some examples, input device 1508 may be configured to receive input from a user via tactile, audio, video, or biometric channels. Examples of input device 1508 may include a keyboard, keypad, mouse, trackball, touchscreen, touchpad, microphone, one or more cameras, image sensor, fingerprint sensor, or any other device capable of detecting input from a user or other source and relaying that input to computer system 1500 or its components.

[0187] In some examples, output device 1510 may be configured to provide output to a user via a visual or auditory channel. Output device 1510 may include a video graphics adapter card, a liquid crystal display monitor, a light-emitting diode (LED) monitor, an organic LED monitor, a sound card, a speaker, a lighting device, an LED, a projector, or any other device capable of generating output that the user can understand. Output device 1510 may also include a touchscreen, a presence-sensitive display, and other displays known in the art with input / output capabilities.

[0188] In some example implementations, the network interface 1512 of the computer system 1500 can be used to communicate with external devices via one or more data networks, such as one or more wired networks, wireless networks, or optical networks, including, for example, the Internet, intranets, LANs, WANs, cellular telephone networks, Bluetooth radios, IEEE 902.11-based radio frequency networks, and Wi-Fi networks®. The network interface 1512 can be a network interface card, such as an Ethernet card, an optical transceiver, a radio frequency transceiver, or any other type of device capable of sending and receiving information.

[0189] Operating system 1514 can control one or more functions of computer system 1500 and / or its components. For example, operating system 1514 can interact with software application 1516 and can facilitate one or more interactions between software application 1516 and components of computer system 1500. Figure 15 As shown, the operating system 1514 can interact with or otherwise couple to the software application 1516 and its components. In some embodiments, the software application 1516 may be included within the operating system 1514. In these and other examples, a virtual module, firmware, or software may be part of the software application 1516.

[0190] Therefore, systems and methods for providing multi-party communication using LLM have been described. Although embodiments have been described with reference to specific example embodiments, it will be apparent that various modifications and changes can be made to these example embodiments without departing from the broader spirit and scope of this application. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.

Claims

1. A method for providing multi-party communication using a Large Language Model (LLM), the method being implemented by a computer system, the method comprising: The LLM receives input messages from user devices associated with multiple communicating parties involved in the scenario, and is configured to understand the scenario. Text input prompts for the LLM are generated based on a scenario template, wherein the scenario template includes rules for the following: The input messages are organized based on the scenario, wherein the plurality of communicating parties and the LLM are assigned roles as participants in the scenario; and The LLM is instructed to analyze text input prompts and calculate its response output, which is defined by the role of the LLM in the scene. The text input prompt is sent to the LLM, which analyzes the text input prompt to write the LLM's response output; The response output of the LLM is analyzed based on the syntax structure provided for the LLM in the scenario template to generate an output message to be sent to one or more of the plurality of communication parties, wherein the one or more communication parties are identified by their roles in the response output of the LLM; and Based on the role of the LLM in the scenario, the output message is sent from the communication endpoint representing the LLM to the user device associated with one or more of the plurality of communication parties.

2. The method according to claim 1, wherein, The scene template includes: The activities that the LLM and the multiple communicating parties need to participate in; The number of communication parties among the plurality of communication parties; The reference to the scenario provides context for the activity, wherein the LLM is configured to understand the scenario based on one of the following: pre-training the LLM, fine-tuning the LLM with additional training data, and incorporating additional data into the text input prompts by using retrieval enhancement generation; The role of each of the plurality of communicating parties is assigned based on the scenario; The LLM roles are assigned based on the scenario and the activity; A first syntactic structure for constructing the input messages from the plurality of communicating parties and assigning the input messages based on roles in the scenario; and The LLM is used to construct the response output to identify the content of the output message to be sent on behalf of the LLM and the second syntax structure of the intended recipient.

3. The method according to claim 1, further comprising: Each time an additional input message with a previous message is received from one or more of the plurality of communicating parties, the operation is repeated in an additional loop, the previous message including a message from the LLM that is included as the input message in additional text input prompts following the scene template.

4. The method according to claim 3, wherein, The repeating of the operation in the additional loop is triggered based on one or more of the following: an algorithm, an external data source, and the LLM analyzing the input message to determine whether the LLM should access the additional loop.

5. The method according to claim 1, wherein, Other elements of the input message and the text input prompt are derived from sources, which include one or more of the following: data events, machines, sensors, additional robots, and other sources accessible to the communicators among the plurality of communicators and the enterprises associated with the plurality of communicators.

6. The method according to claim 1, wherein, The input message is associated with one or more modalities, which include one of the following: text, image, audio, video, and location; The LLM is configured to write the response output having a modality associated with the input message.

7. The method according to claim 1, wherein, The scenario template includes additional rules regarding the actions to be taken by the LLM when it is requested to provide information not included in the text input prompt; The additional rules include one or more of the following: A message is generated to request the information from one of the plurality of communicating parties, wherein the communicating party is expected to possess the information based on its role or the rules in the scenario template; and Generate additional messages to request the information from external data sources, which include one or more of the following: additional machines; sensors; robots; and additional sources including systems, databases, and other sources accessible to the communicators among the plurality of communicators or to enterprises associated with the communicators among the plurality of communicators, wherein the communicators among the plurality of communicators are expected to have the information based on the role of the communicators among the plurality of communicators or the rules in the scenario template.

8. The method according to claim 1, wherein, The LLM is configured to acknowledge receipt of a message from one of the plurality of communicating parties.

9. The method according to claim 1, wherein, The LLM is configured as follows: Analyze previous messages, including messages from the LLM; Determine when one or more of the plurality of communicating parties are expected to follow up; and This ensures that the notification is sent to the appropriate communication party.

10. The method according to claim 1, wherein, The LLM is configured as follows: It was identified that a subset of the plurality of communicating parties were communicating directly with each other; and In response to identification, interference with multi-party communication between subsets of the multiple communicating parties is avoided.

11. The method according to claim 1, wherein, The LLM is configured to perform one or more of the following: Send instructions and messages to one or more external systems; Perform actions in the one or more external systems; Trigger workflows in one or more external systems; Verify that the LLM has collected the necessary inputs to perform the action or trigger the workflow; Based on the analysis of the sessions associated with the multi-party communication, the certainty of taking the action is estimated; The certainty is assessed against a threshold set by the organization associated with the one or more external systems; The action is performed or the workflow is triggered based on the assessment that the certainty is higher than the threshold, and based on the necessary input collected from the multi-party communication. as well as When the certainty is below the threshold, a message is sent to the appropriate party among the plurality of communicating parties for confirmation or instruction on how to proceed.

12. The method according to claim 1, wherein, The LLM is configured as follows: To generate code using a computer language associated with one or more external systems; and Based on the code, perform one or more of the following: Send instructions to the one or more external systems; Send the message to the one or more external systems; Perform actions in the one or more external systems; as well as Trigger workflows in one or more external systems.

13. The method according to claim 1, wherein, The LLM is configured as follows: Understand the language of one or more of the plurality of communicating parties; In the multi-party communication, the communication is conducted in the language selected by the one or more communicating parties. and Automatically translate messages from the first communication party among the one or more communication parties from a first language selected by the first communication party into a second language selected by the second communication party among the one or more communication parties.

14. The method according to claim 1, wherein, The LLM is configured to coordinate multiple conversations of one or more subsets of the multiple communicating parties as part of the same session associated with the multi-party communication.

15. The method of claim 1, further comprising instructing the LLM to perform one or more of the following in the text input prompt: A communication party may be added to the plurality of communication parties based on one or more of the following: a turning point in the scenario, an external event, a request from another communication party, a request for information that is not available in the text input prompt or among the plurality of communication parties, and a requirement provided in the scenario and inferred by the LLM. Based on the text input prompt, a message is sent to the system that controls the multiple communication parties, so that the communication party is invited to join the multiple communication parties; The communication party among the plurality of communication parties is removed based on one or more of the following: a turning point in the scenario, an external event, a request from another communication party, a request from the communication party to be removed, and a requirement provided in the scenario and inferred by the LLM. as well as Based on the text input prompt, another message is sent to the system controlling the multiple communication parties, causing the communication party to be removed from the multiple communication parties.

16. A system for providing multi-party communication using a Large Language Model (LLM), the system comprising: processor; as well as A memory for storing instructions, which, when executed by the processor, configure the processor to: The LLM receives input messages from user devices associated with multiple communicating parties involved in the scenario, and is configured to understand the scenario. Text input prompts for the LLM are generated based on a scenario template, wherein the scenario template includes rules for the following: The input messages are organized based on the scenario, wherein the plurality of communicating parties and the LLM are assigned roles as participants in the scenario; and The LLM is instructed to analyze text input prompts and calculate its response output, which is defined by the role of the LLM in the scene. The text input prompt is sent to the LLM, which analyzes the text input prompt to write the LLM's response output; The response output of the LLM is analyzed based on the syntax structure provided for the LLM in the scenario template to generate an output message to be sent to one or more of the plurality of communication parties, wherein the one or more communication parties are identified by their roles in the response output of the LLM; and Based on the role of the LLM in the scenario, the output message is sent from the communication endpoint representing the LLM to the user device associated with one or more of the plurality of communication parties.

17. The system according to claim 16, wherein, The scene template includes: The activities that the LLM and the multiple communicating parties need to participate in; The number of communication parties among the plurality of communication parties; The reference to the scenario provides context for the activity, wherein the LLM is configured to understand the scenario based on one of the following: pre-training the LLM, fine-tuning the LLM with additional training data, and incorporating additional data into the text input prompts by using retrieval enhancement generation; The role of each of the plurality of communicating parties is assigned based on the scenario; The LLM roles are assigned based on the scenario and the activity; A first syntactic structure for constructing the input messages from the plurality of communicating parties and assigning the input messages based on roles in the scenario; and The LLM is used to construct the response output to identify the content of the output message to be sent on behalf of the LLM and the second syntax structure of the intended recipient.

18. The system according to claim 16, wherein, The processor is further configured to repeat the operation in a separate loop each time an additional input message with a previous message is received from any of the plurality of communicating parties, the previous message including a message from the LLM that is included as the input message in additional text input prompts following the scene template.

19. The system according to claim 18, wherein, The repeating of the operation in the additional loop is triggered based on one or more of the following: an algorithm, an external data source, and the LLM analyzing the input message to determine whether the LLM should access the additional loop.

20. The system according to claim 16, wherein, Other elements of the input message and the text input prompt are derived from sources, which include one or more of the following: data events, machines, sensors, additional robots, and other sources accessible to the communicators among the plurality of communicators and the enterprises associated with the plurality of communicators.

21. The system according to claim 16, wherein, The input message is associated with one or more modalities, which include one of the following: text, image, audio, video, and location; The LLM is configured to write the response output having a modality associated with the input message.

22. The system according to claim 16, wherein, The scenario template includes additional rules regarding the actions to be taken by the LLM when it is requested to provide information not included in the text input prompt, and wherein the additional rules include one or more of the following: A message is generated to request the information from one of the plurality of communicating parties, wherein the communicating party is expected to possess the information based on its role or the rules in the scenario template; and Generate additional messages to request the information from external data sources, which include one or more of the following: additional machines; sensors; robots; and additional sources including systems, databases, and other sources accessible to the communicators among the plurality of communicators or to enterprises associated with the communicators among the plurality of communicators, wherein the communicators among the plurality of communicators are expected to have the information based on the role of the communicators among the plurality of communicators or the rules in the scenario template.

23. The system according to claim 16, wherein, The LLM is configured as follows: Analyze previous messages, including messages from the LLM; Determine when one or more of the plurality of communicating parties are expected to follow up; and This ensures that the notification is sent to the appropriate communication party.

24. The system according to claim 16, wherein, The LLM is configured as follows: It was identified that a subset of the plurality of communicating parties were communicating directly with each other; and In response to identification, interference with multi-party communication between subsets of the multiple communicating parties is avoided.

25. The system according to claim 16, wherein, The LLM is configured to perform one or more of the following: Send additional instructions and messages to one or more external systems; Perform actions in the one or more external systems; Trigger workflows in one or more external systems; Verify that the LLM has collected the necessary inputs to perform the action or trigger the workflow; Based on the analysis of the sessions associated with the multi-party communication, the certainty of taking the action is estimated; The certainty is assessed against a threshold set by the organization associated with the one or more external systems; The action is performed or the workflow is triggered based on the assessment that the certainty is higher than the threshold, and based on the necessary input collected from the multi-party communication. as well as When the certainty is below the threshold, a message is sent to the appropriate party among the plurality of communicating parties for confirmation or instruction on how to proceed.

26. The system according to claim 16, wherein, The LLM is configured as follows: To generate code using a computer language associated with one or more external systems; and Based on the code, perform one or more of the following: Send additional instructions to the one or more external systems; Send the message to the one or more external systems; Perform actions in the one or more external systems; as well as Trigger workflows in one or more external systems.

27. The system according to claim 16, wherein, The LLM is configured as follows: Understand the language of one or more of the plurality of communicating parties; In the multi-party communication, the communication is conducted in the language selected by the one or more communicating parties. and Automatically translate messages from the first communication party among the one or more communication parties from a first language selected by the first communication party into a second language selected by the second communication party among the one or more communication parties.

28. The system according to claim 16, wherein, The LLM is configured to coordinate multiple conversations of one or more subsets of the multiple communicating parties as part of the same session associated with the multi-party communication.

29. The system according to claim 19, wherein, The processor is also configured to instruct the LLM to perform one or more of the following in the text input prompt: A communication party may be added to the plurality of communication parties based on one or more of the following: a turning point in the scenario, an external event, a request from another communication party, a request for information that is not available in the text input prompt or among the plurality of communication parties, and a requirement provided in the scenario and inferred by the LLM. Based on the text input prompt, a message is sent to the system that controls the multiple communication parties, so that the communication party is invited to join the multiple communication parties; The communication party among the plurality of communication parties is removed based on one or more of the following: a turning point in the scenario, an external event, a request from another communication party, a request from the communication party to be removed, and a requirement provided in the scenario and inferred by the LLM. as well as Based on the text input prompt, another message is sent to the system controlling the plurality of communication parties, causing the communication party to be removed from the plurality of communication parties.

30. A method for providing multi-party communication using a Large Language Model (LLM), the method being implemented by a computer system, the method comprising: The LLM receives input messages from user devices associated with multiple communicating parties involved in the scenario, and is configured to understand the scenario. Generate text input prompts for the LLM, the text input prompts being generated based on a scene template, wherein the scene template includes rules for the following: The input messages are organized based on the scenario, wherein the plurality of communicating parties and the LLM are assigned roles as participants in the scenario; and The LLM is instructed to analyze text input prompts and calculate its response output, which is defined by the role of the LLM in the scene. The scene template includes: The activities that the LLM and the multiple communicating parties need to participate in; The number of communication parties among the plurality of communication parties; The reference to the scenario provides context for the activity, wherein the LLM is configured to understand the scenario based on one of the following: pre-training the LLM, fine-tuning the LLM with additional training data, and incorporating additional data into the text input prompts by using retrieval enhancement generation; The role of each of the plurality of communicating parties is assigned based on the scenario; The LLM roles are assigned based on the scenario and the activity; A first syntactic structure for constructing the input messages from the plurality of communicating parties and assigning the input messages based on the roles in the scenario; The LLM is used to construct the response output to identify the content of the output message to be sent on behalf of the LLM and the second syntax structure of the intended receiver; and Regarding additional rules for the action to be taken by the LLM when it is requested to provide information not included in the text input prompt, the additional rules include generating a message to request the information from one of the plurality of communication parties, wherein the communication party among the plurality of communication parties is expected to have the information based on the role of the communication party among the plurality of communication parties or the rules in the scene template; The text input prompt is sent to the LLM, which analyzes the text input prompt to write its response output. The response output includes one or more messages to be sent to one or more of the plurality of communication parties, as well as multi-party communication instructions. The LLM is configured to: It was identified that a subset of the multiple communicating parties were communicating directly with each other; Based on identification, interference with multi-party communication between subsets of the multiple communicating parties is avoided; and Coordinate multiple conversations of one or more subsets of the multiple communicating parties as part of the same session associated with the multi-party communication; The response output of the LLM is analyzed based on the syntax structure provided for the LLM in the scenario template, so as to generate an output message to be sent to one or more of the plurality of communication parties based on the one or more messages, wherein the one or more communication parties are identified by the roles of the one or more communication parties in the response output of the LLM; The tag is added to at least one output message based on the multi-party communication instructions; Based on the multi-party communication instructions, the output message is sent from the communication endpoint representing the LLM to the user device associated with one or more of the multiple communication parties, according to the role of the LLM in the scenario. Analyzing previous messages, including messages from the LLM, to determine when the attention of one or more of the plurality of communication parties is needed and to cause a notification to be sent to one or more of the plurality of communication parties; and Each time an additional input message with a previous message is received from one or more of the plurality of communicating parties, the operation is repeated in an additional loop, the previous message including a message from the LLM that is included as the input message in additional text input prompts following the scene template.