Computer-implemented multi-user messaging application

By integrating generative models into computer-implemented messaging applications, providing group conversation prompts, and leveraging a search system, we address the problem of generative models maintaining context in group conversations, enabling effective participation and collaboration in a multi-user environment.

CN120677689APending Publication Date: 2025-09-19MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480011488.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-29
Filing Date
2024-03-06
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing generative models struggle to maintain conversational context in group conversations and are unable to effectively handle ambiguity between multiple users, making it impossible to effectively participate in group conversations in computer-implemented messaging applications.

Method used

By integrating the generative model into a computer-implemented messaging application, prompts related to group conversations are provided, the generative model is trained to recognize multi-user environments, and messages from multiple users are processed in combination with images. The generative model generates output based on these prompts, and uses the search system to obtain relevant data to support group conversations.

Benefits of technology

The generative model has been enabled to effectively participate in group conversations, and can summarize conversations, identify user input, assist in scheduling events, generate graphics, translate content, answer questions, etc., thus improving the collaborative efficiency and user experience of group conversations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120677689A_ABST
    Figure CN120677689A_ABST
Patent Text Reader

Abstract

A computing system includes a processor and a memory storing instructions that, when executed by the processor, cause the processor to perform a number of actions. The actions include receiving a plurality of messages from a plurality of users in a messaging application supporting a group conversation, wherein the plurality of messages are included in the group conversation. The actions also include providing a cue to the generative model, wherein the cue includes a plurality of messages. The actions also include receiving, from the generative model, an output generated by the generative model based on the cue, and including the output as a turn in the group conversation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 451,547, filed on March 10, 2023, entitled “COMPUTER-IMPLEMENTED MULTI-USER MESSAGING APPLICATION WITH AI BOT.” The entire contents of that application are incorporated herein by reference. Background Art

[0003] There are various types of computer-implemented messaging applications that allow multiple users of multiple client computing devices to participate in group conversations with each other (synchronously and / or asynchronously). Examples of computer-implemented messaging applications include, but are not limited to, text messaging applications, instant messaging applications, unified communications (UC) applications, and the like.

[0004] Relatively recently, generative models have been developed, including generative language models (GLMs) (also known as large language models (LMSs)), models that generate images based on input (where the input can be text, speech, images, etc.), models that generate videos based on input, and the like. An example of a GLM is the Generative Pretrained Transformer 4 (GPT-4) model. Another example of a GLM is the Large Open Science Open Access Multilingual Model (BLOOM) model, which is also a transformer-based model. In short, a generative model is configured to generate an output (such as text in a human-readable language, source code, music, video, etc.) based on a prompt provided as an input to the generative model, wherein the generative model generates the output in near real time (e.g., within seconds of receiving the prompt).

[0005] Computer-implemented applications are being developed to include generative models. For example, generative models have been incorporated into chat applications, where a single user can interact with the generative model through the chat application. Accordingly, the (single) user provides input to the chat application and the generative model generates output based on this input and presents the output to the user. However, conventional generative models are designed to interact with a single user, and the design of such generative models has prevented them from being incorporated into computer-implemented messaging applications that support group conversations. For example, conventional generative models always maintain the conversation context between the user and the generative model during a conversation between the user and the generative model. Conventional generative models cannot maintain the conversation context in group conversations because (due to the design of the generative model) the generative model cannot disambiguate between the different users participating in the conversation. Summary of the Invention

[0006] The following is a brief summary of subject matter that is described in greater detail herein. This summary is not intended to limit the scope of the claims.

[0007] Various techniques are described herein that involve integrating generative models into computer-implemented messaging applications that support group messaging (messaging between at least two participants). Computer-implemented messaging applications include text messaging applications, instant messaging applications, unified communications applications, and other applications that support synchronous messaging between participants in a group conversation (computer-implemented messaging applications may also support asynchronous messaging).

[0008] In an example, a server computing system receives a plurality of messages from a plurality of participants in a group conversation. The messages in the plurality of messages may include text, images, audio (e.g., music, voice input, etc.), video, or other computer-readable input that can be processed by a generative model. The generative model generates an output based on the plurality of messages from the plurality of participants. The output may include text, images, video, audio, etc.

[0009] Various methods can be employed to allow a generative model to generate output based on multiple messages. In one example, a computer-implemented messaging application (through which a group conversation is conducted) constructs a prompt and provides the prompt to the generative model, wherein the prompt, for example, identifies that the conversation is a group conversation including multiple users, the identities of the users in the group conversation, and also includes multiple messages from the group conversation (with each message identifying the user who generated the message). The generative model then generates an output based on the prompt. In another example, the generative model is trained to inherently recognize multi-user conversations.

[0010] Integrating generative models into computer-implemented messaging applications that support group messaging enables a variety of use cases that have heretofore been unavailable. For example, generative models can summarize group conversations, identify input from specific users about a topic, propose follow-up actions based on the group conversation, assist in scheduling events for participants in the group conversation, assist in coordinating the schedules of users in the group conversation, generate graphics (such as images, avatars, emoticons, etc.) that are specific to the group conversation, assist in rewriting and reviewing the content of the group conversation (e.g., modifying the conversation to make the messages therein more professional, more humorous, correcting type and grammatical inconsistencies, translating the conversation into a different language, etc.), answer questions about the content of a uniform resource locator (URL) included in the group conversation (e.g., summarizing a news article, identifying facts mentioned in the article, reading structured data (such as a company earnings document), answering specific questions, etc.), assist in brainstorming sessions occurring in the group conversation, act as a translation engine for translating portions of the group conversation into a language understandable by other members of the group conversation, start and run games for entertaining members of the group conversation, etc.

[0011] The above summary presents a simplified summary of the invention in order to provide a basic understanding of some aspects of the systems and / or methods discussed herein. This summary is not a comprehensive overview of the systems and / or methods discussed herein. It is not intended to identify key / critical elements or to delineate the scope of such systems and / or methods. Its sole purpose is to present some concepts in a simplified form as a prelude to the detailed description presented later. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a functional block diagram of a computing system in which a generative model is integrated into a messaging application that supports group conversations.

[0013] Figure 2 is a functional block diagram of a client computing device displaying a group conversation including entries generated by a bot that uses a generative model to create output.

[0014] Figure 3 is a schematic diagram including a functional block diagram of a computer-implemented messaging application that supports group conversations.

[0015] Figure 4 Prompts that may be provided by a computer-implemented messaging application to a robot that includes a generative model are depicted.

[0016] Figure 5 Prompts that may be provided by a computer-implemented messaging application to a robot that includes a generative model are depicted.

[0017] Figure 6 is a flow chart depicting a method related to a robot participating in a group conversation via a computer-implemented messaging application.

[0018] Figure 7 is a flow chart depicting a method for constructing prompts to be provided to a robot including a generative model.

[0019] Figure 8 is a schematic diagram of a computing device. DETAILED DESCRIPTION

[0020] Various techniques related to integrating the functionality of a robot into a computer-implemented messaging application that supports group conversations will now be described with reference to the accompanying drawings, wherein the same reference numerals are used to refer to the same elements throughout. In the description below, for the purpose of explanation, a number of specific details are set forth in order to provide a thorough understanding of one or more aspects. However, it is apparent that these aspects can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form to facilitate describing one or more aspects. Further, it should be understood that functions described as being performed by certain system components can be performed by multiple components. Similarly, for example, a component can be configured to perform functions described as being performed by multiple components.

[0021] Furthermore, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless specified otherwise, or clear from the context, the phrase "X employs A or B" is intended to mean any of the natural inclusive permutations. That is, the phrase "X employs A or B" satisfies any of the following: X employs A; X employs B; or X employs both A and B. Furthermore, the articles "a" and "an" used in this application and the appended claims should generally be construed to mean "one or more" unless specified otherwise or the context clearly dictates a singular form.

[0022] Further, as used herein, the terms "component," "module," and "system" are intended to encompass a computer-readable data storage device configured with computer-executable instructions that, when executed by a processor, cause certain functions to be performed. Computer-executable instructions may include routines, functions, and the like. It should also be understood that a component or system may be located on a single device or distributed across multiple devices. Further, as used herein, the term "exemplary" is intended to mean serving as an illustration or example of something and is not intended to indicate a preference.

[0023] Various techniques related to integrating a robot including a generative model into a computer-implemented messaging application that supports group conversations are described herein. As described above, conventional generative models, when integrated into a chat application, are limited to conversations with only a single user. The techniques described herein are directed to enabling computer-implemented messaging applications that support group conversations to integrate generative models, thereby enabling several use cases that have been impossible to date. As will be described in more detail herein, a generative model can be integrated into a messaging application, wherein this integration is implemented based on a messaging application that provides a prompt to the generative model, wherein the prompt identifies that the conversation is a group conversation and also includes information related to the group conversation (such as the identification of the participants, messages previously transmitted in the group conversation, and other information). In another example, integrating a generative model into a messaging application is implemented by training the generative model to recognize that it participates in group conversations that include multiple users (rather than a single user).

[0024] Now refer to Figure 1 , which illustrates a functional block diagram of a computing architecture 100 that facilitates integrating a robot including a generative model into a computer-implemented messaging application that supports group messaging. Architecture 100 includes a server computing system 102 and a plurality of client computing devices 104 to 106. Server computing system 102 communicates with client computing devices 104 to 106 via a network 108, such as the Internet. Although not shown, each of client computing devices 104 to 106 has a client messaging application installed thereon. The computer-implemented messaging application supports synchronous messaging and, therefore, can be an instant messaging application, a unified communications application, or the like. In another example, the messaging application supports asynchronous communication and, therefore, can be a text messaging application. Other computer-implemented messaging applications that support group conversations are also contemplated. The client messaging application installed on client computing devices 104 to 106 can be a standalone application, or can be integrated into a browser, or the like.

[0025] The server computing system 102 includes a processor 110 and a memory 112, wherein the memory 112 includes instructions executed by the processor 110. Figure 1 As illustrated, the memory 112 includes a computer-implemented messaging application 114 through which users of the client computing devices 104 to 106 can participate in a group conversation. Figure 1 In the depicted embodiment, the messaging application 114 has a client-server architecture. However, it should be understood that the functionality described herein can be used in a peer-to-peer messaging application, where the bot acts as a peer in such an architecture.

[0026] Memory 112 also includes a robot 116, which includes or can access a generative model 118. In an example, the generative model 118 is a transformer-based model. The generative model 118 can be a GLM, so that the generative model 118 is configured to receive text input and generate text output. In this case, the robot 116 can be a chatbot. In other examples, the generative model 118 can generate audio, images, videos (with audio) or other appropriate outputs based on various inputs (such as voice, text, video, images, audio, etc.). Although the robot 116 is illustrated as being outside the messaging application 114, it is contemplated that the robot 116 can be included in the messaging application 114. As will be described in more detail herein, the messaging application 114 receives messages from the client computing devices 104 to 106, which are part of a group conversation between users of the client computing devices 104 to 106. The bot 116 generates output based on the received message by using the generative model 118 and provides the output as part of the group conversation to the client computing devices 104 to 106. Thus, the bot 116 can participate in the group conversation.

[0027] The server computing system 102 may also optionally include a search system 120. More specifically, the server computing system 102 may include several data sources 122 to 124, and the search system 120 may search the data sources 122 to 124 based on messages included in the conversation and / or queries generated by the robot 116, wherein the robot 116 generates the query based on at least one message in the group conversation. In an example, the data sources 122 to 124 may be or include a search engine index, an instant answer index, an image index, a knowledge graph, a data source that stores historical information (such as previous group conversations conducted between members of the group), a data source that includes calendar information of participants in the group conversation, and other information. In addition, although not shown, the search system 120 may access external data sources, such as a web server of a host website, a document repository, etc.

[0028] Now, an example operation of the server computing system 102 will be described. Users of the client computing devices 104 to 106 participate in a group conversation via the messaging application 114, such that, for example, the messaging application 114 receives a first message from a first client computing device 114 and transmits the first message to a computing device operated by a user participating in the group conversation. Subsequently, the messaging application 114 receives an nth message from an nth client computing device 106 and transmits the nth message to a computing device operated by a user participating in the group conversation.

[0029] According to an example, messaging application 114 receives an indication that bot 116 will be invoked. Messaging application 114 provides bot 116 with a first message and an nth message, as well as information related to the requested invocation of bot 116. Bot 116 generates output based on the information provided to bot 116 by messaging application 114, and bot 116 provides the generated output to messaging application 114. Messaging application 114 transmits the output to client computing devices 104-106. For example, the output may be transmitted to client computing devices 104-106 so that it appears as if bot 116 is a member of a group and is participating in a group conversation with users of client computing devices 104-106. Additionally, bot 116 may generate output based on data obtained by search system 120, where search system 120 obtains the data by searching data in one or more of data sources 122-124 and / or from external data sources.

[0030] Now refer to Figure 2 , which presents a functional block diagram of a first client computing device 104. The first client computing device 104 includes a processor 202 and a memory 204, wherein the memory 204 includes a client messaging application 206. When executed by the processor 202, the client messaging application 206 establishes communication with the messaging application 114 of the server computing system 102 so that data can be transferred between the client messaging application 206 and the messaging application 114. For example, the client messaging application 206 can receive a message to be included in a group conversation from a user of the first client computing device 104 and can transmit such a message to the messaging application 114. Upon receiving the message, the messaging application 114 can transmit the message to computing devices operated by other users participating in the group conversation. Similarly, when the messaging application 114 receives a message in the group conversation from the nth client computing device 106, the messaging application 114 can transmit the message to computing devices operated by other users participating in the group conversation (including the first client computing device 104).

[0031] The first client computing device 104 also includes a display 208, where the display 208 can depict a group conversation 210 implemented by the client messaging application 206 and the messaging application 114. The group conversation 210 includes a plurality of messages elaborated by several users participating in the group conversation 210. In some examples, each message is referred to as a "turn." Figure 2In the depicted example group conversation 210, a first user initially composes a first message in the group conversation 210, followed by a second user composing a second message, after which a third user composing a third message. The first user then contributes a fourth message to the group conversation 210, followed by a third user composing a fifth message to the group conversation 210. The group conversation 210 also includes a sixth message composing by the second user, wherein the sixth message invokes the bot 116 and includes a request for output from the bot 116. For example, the second user can indicate to the messaging application 114 that the bot 116 is to be invoked by using an "@" symbol followed by an identifier for the bot 116, and the request for output can follow the identifier for the bot 116.

[0032] The generative model 118 of the bot 116 generates output based on the request included in the sixth message and further based on the previous messages in the group conversation 210. As illustrated in the group conversation 210, the bot 116 generates output and provides the output to the messaging application 114, which in turn transmits the output to the first client computing device 104. It should also be noted that other participants in the group conversation can also invoke the bot 116. For example, later in the conversation, the first user indicates to the messaging application 114 that they want to invoke the bot 116, and the generative model 118 of the bot 116 generates output based on several messages in the group conversation 210.

[0033] Now refer to Figure 3 , which illustrates a functional block diagram of the messaging application 114. The messaging application 114 includes an invocation detector module 302, a prompt generator module 304, and conversation information 306. As previously indicated, a user in a group conversation can invoke the bot 116 (e.g., request that the bot 116 participate in the group conversation). The invocation detector module 302 can monitor the group conversation and detect a request by a user participating in the group conversation to invoke the bot 116. For example, as identified above, a user can employ the "@" symbol and invoke the bot 116. In another example, a user can formulate a voice command requesting that the bot 116 be invoked.

[0034] When the call detector module 302 detects a request to call the bot 116, the call detector module 302 can notify the prompt generator module 304 to generate a prompt for the bot 116, where the prompt is an input provided to the generative model 118, and further, the input indicates the output to be generated by the generative model 118. In some examples, the prompt includes text indicating the output to be produced by the generative model. The prompt generator module 304 generates the prompt 308 based on the conversation information 306. The conversation information 306 can include information such as the time when the group conversation started, the identity of the users in the group conversation, an indication that the conversation is a group conversation, the time when the messaging application 114 obtained the message, the content of the message in the group conversation, and other information.

[0035] refer to Figure 4 , which depicts an example prompt 400 output by the prompt generator module 304. Historically, prompts provided to the generative model assume that the bot 116 is communicating directly with a single user, with no other participants in the conversation. Furthermore, such prompts include examples that the generative model 118 would employ when responding to a single user's query. Instead, the prompt generator module 304 appends new instructions to the typical prompt (below the context conventionally provided to the generative model 118 for conversing with a user on a one-on-one basis) to guide the generative model 118 in understanding the context of a group conversation. The prompt 400 generated by the prompt generator module 304 includes the bot 116's interpretation of the group conversation. Furthermore, the prompt 400 generated by the prompt generator module 304 includes the identities of the users participating in the group conversation. This information is important because, without the identities of the users in the group conversation, the generative model 118, when reading messages in the group conversation, may be confused about who the participants are and which participants are providing what messages. To obtain the list of participants' identities, prompt generator module 304 may obtain the identities from a client application executed by client computing devices 104 - 106 or by making backend calls to a roster application programming interface (API) maintained by messaging application 114 .

[0036] The prompt generator module 304 then includes a transcript of the messages in the group conversation in a prompt, where the transcript includes the content of the message, a timestamp indicating when the message was obtained by the messaging application 114, and the identity of the user who submitted the message. The prompt generator module 304 also identifies the user requesting to invoke the bot 116 in the prompt 400 and optionally includes the user's location in the prompt 400. The prompt generator module 304 may also include instructions in the prompt 400 to be provided to the bot 116 to constrain the bot 116's output. For example, the instructions may indicate that the user requesting to invoke the bot 116 is interested in responses from the context of the provided group conversation (e.g., from the transcript of the group conversation). Further, the prompt 400 may include a request for the bot 116 not to make inferences, not to provide personal information about people participating in the group conversation, not to make inferences about users participating in the group conversation, etc. These instructions may help reduce the illusion of output from the bot 116 because the bot 116 is constrained to respond using information within the context.

[0037] In the example, when the bot 116 is not restricted to this context in prompt 400, the user may formulate an invocation request including a sentence such as "Do you remember when John traveled to Australia?" to the bot 116. Without the restriction instructions referenced above, the bot 116 may generate a completely fictitious output describing John's travel to Australia, as the group conversation may not include information about John's travel to Australia.

[0038] Optionally, the prompt generator model 304 can include other information about the users in the group conversation in the prompt 400, such as information obtained from the users' profiles, including topics of interest to the users, keywords and facts mentioned by the users in other conversations, demographic information of the users (such as age), etc., allowing the robot 116 to customize the response.

[0039] Return to Figure 3 , a prompt 308 generated by the prompt generator module 304 is provided to the bot 116, which in turn provides the prompt 308 to the generative model 118. The generative model 118 generates an output 310 based on the prompt, and the bot 116 provides the output 310 to the messaging application 114. The messaging application 114 includes the output 310 in the group conversation, making it appear as if the bot 116 is participating in the group conversation with other users.

[0040] Optionally, the bot 116 may interact with the search system 120 in conjunction with creating updated prompts to be used by the bot 116 to generate output. In a non-limiting example, a user in a group conversation may request to invoke the bot 116 and formulate the input "We all want to see a movie this afternoon. What time is convenient for everyone, and what movies are playing at that time?" The prompt generator module 304 may generate a prompt for the bot 116 based on the invocation request. In this case, the prompt 308 generated by the prompt generator module 304 does not include a restriction that limits the response to the context included in the prompt 308. The generation model 118 receives the prompt 308 and generates a query based on the prompt 308 and provides the query to the search system 120. The search system 120 generates a query based on the query in one or more data sources 122 to 124 ( Figure 1 ) and provides at least some of the search results as part of a prompt to be used by the generation model 118 to respond to the user request. For example, the first data source 122 may include movie times and locations, while the mth data source 124 may include calendar information for participants in the group message. The generation model 118 generates output 310 based on this prompt and provides the output 310 to the messaging application 114, which includes the output 310 in the ongoing group conversation between the users of the client computing devices 104 to 106.

[0041] Prompt 400 is illustrated as an example prompt associated with a conventional generative model that has been modified to allow the generative model 118 to be integrated into the messaging application 114. In other examples, the prompt generator module 304 can generate a prompt that correctly describes to the bot 116 from the outset that the bot 116 is in a multi-user group conversation.

[0042] Brief reference Figure 5, which depicts another example prompt 500 that may be generated by the prompt generator module 304. Prompt 500 is similar to prompt 400, except that the dialogue information 306 included in prompt 500 includes output from the robot 116. It has been observed that when a robot 116 generates output based on its own previous output, the subsequently generated output may become unpredictable, erratic, and generally undesirable. To address this issue, the prompt generator module 304 may optionally hide one or more of the robot 116's outputs included in the dialogue information 306. Various methods may be used to determine how much of the robot 116's output to include in prompt 500 and / or how much of the robot 116's output to hide. For example, a sliding window may be employed, where only a threshold number of the robot 116's recent outputs are included in prompt 500 (while the other outputs are hidden). In another example, a rolling window of the robot 116's output may be retained, while the others are hidden. An example of such a rolling window is described below.

[0043] In this example, the threshold number of outputs for the bot 116 is 5. This means that when one of the users in the group calls the bot 116 for the sixth time, the bot 116 is provided with a prompt containing its past five responses. When the user calls the bot for the seventh time, the prompt generator module 304 hides all of the bot's output in the prompt. This is represented algorithmically as follows:

[0044] ●So resetConversationInTurn=5

[0045] ●Then botTurnsToKeep=previousTurnCount%(resetConversationInTurn+1)

[0046] o If this is the 6th call of the bot, previousTurnCount=6-1=5, and botTurnsToKeep=5%(5+1)=5 -> the last 5 bot replies are included in the prompt, and other earlier replies are replaced with [hidden].

[0047] If this is the 7th call of the bot, then previousTurnCount = 7-1 = 6, and botTurnsToKeep = 6%(5+1) = 0 -> bot output is not included in the prompt; all bot output is replaced with [hidden]

[0048] ○ If this is the 8th call of the bot, then previousTurnCount = 8 - 1 = 7, botTurnsToKeep = 1, so the record will look like:

[0049] Human A: XXXXXXX

[0050] Human B: YYYYY

[0051] ■Robot:[Hidden]

[0052] Human A: @Robot Problem

[0053] ■Robot:Reply

[0054] Human B: @Robot's new question

[0055] This can be illustrated in the table below:

[0056]

[0057] again, Figure 5 The prompt 500 is illustrated as including concealment of the output of the robot 116 .

[0058] The bot 116 is limited to receiving prompts of a certain size (e.g., a threshold number of tokens). Accordingly, the prompt generator module 304 can perform one or more actions to ensure that the number of tokens in the prompt 308 provided to the bot 116 is equal to or less than the threshold, so that the bot 116 can process the prompt 308. As previously indicated, the prompt generator module 304 can include various types of information in the prompt and / or the search system 120 can include information in the prompt 308, where such information can include general instructions for how the bot 116 should generate output, including hypothetical examples of questions and answers between the bot 116 and users in the group conversation. The prompt 308 can also include a JSON document and a web snippet of search results initiated by the bot 116 and retrieved by the search system 120. The prompt 308 also includes at least some of the conversation information in the conversation information 306, including a transcript of the group conversation prior to receiving the call request. Finally, the prompt 308 can include a request to the bot associated with the call.

[0059] All of this information consumes space in the prompt 308; in order to keep the prompt 308 within an allowed size (e.g., so that the number of tokens is less than a threshold), the prompt generator module 304 can truncate the transcript of the group conversation when the conversation becomes too long. The prompt generator module 304 can employ a variety of different strategies when truncating the conversation transcript. In an example, the prompt generator module 304 can obtain the entire conversation transcript sorted by time (from most recent to oldest). The prompt generator module 304 then fills a buffer of a specific number of tokens, line by line, from most recent to oldest (it is important to note that the term "token" is not necessarily equivalent to a word or character, but rather refers to an entity for which the robot 116 has a semantic understanding, such that a token in one language may represent four characters, while a token in another (more semantically dense) language may represent one character).

[0060] In this example, the prompt generator module 304 does not truncate the content of any individual message elaborated by the user. When the prompt generator module 304 determines that adding the next message to the record would cause the token count in the buffer to exceed a certain number, the message is not added to the context. Thereafter, the prompt generator module 304 reverses the order in the buffer so that the messages are arranged from oldest to newest, and the prompt generator module 304 appends the buffer to the context.

[0061] In another embodiment, because messages from the bot 116 may be relatively long, the prompt generator module 304 may request that the bot 116 truncate (summarize) its own output, thereby reducing the number of tokens required to represent the bot 116's output. In another example, when the number of tokens representing a conversation transcript exceeds a certain number, the bot 116 may be provided as input and requested to truncate (summarize) the transcript (e.g., reducing the transcript from X characters to Y characters, where Y is less than X). In yet another example, the prompt generator module 304 may utilize multiple different models to truncate a portion of the conversation transcript. For example, when the conversation transcript is relatively long, the prompt generator module 304 may select the next few lines, call another model to summarize the discussion, and then perform the same operation on the next batch until all lines have been processed.

[0062] In addition, to reduce computing resource usage, the messaging application 114 can be configured to identify certain robot calls as requiring less computation and other calls as requiring more computation. For example, the messaging application 114 communicates with multiple different robots, where different robots utilize different amounts of computing resources when generating output. For example, one robot can be trained to perform only summarization and therefore require a small amount of computing resources when generating output. Conversely, a second robot can be configured to obtain information generated by the search system 120 and generate an image based on that information, thereby requiring a much greater amount of computing resources than the first robot. The messaging application 114 can identify which robot should answer the call request and can provide a prompt to the appropriate robot, thereby conserving computing resources.

[0063] Further, as described above, when generating certain types of output (such as images, music, videos, etc.), the generation model 118 can utilize a large amount of computing resources. In some instances, users participating in a group conversation may have accounts associated with the messaging application 114, where the accounts include value units (such as reward points). When a user who calls the robot 116 submits a request for the robot 116 to utilize a large amount of computing resources to generate an output, the user's account can be charged for such output. For example, a certain number of reward points can be deducted from the account of the user who called the robot 116. In another example, participants in the group conversation are randomly selected and value units are deducted from the randomly selected user's account. In yet another example, when the robot 116 generates output based on a user request, a polling method is used to deduct value units from the user's account.

[0064] As expressed above, the technology described herein allows for use cases that have heretofore been impossible in messaging applications that allow group conversations. For example, the robot 116 can answer questions related to the context of the group conversation as well as general questions. In some implementations, the robot 116 generates output only when a user participating in the group conversation calls it. In another example, the robot 116 is provided with input for each turn of the group conversation, and the robot 116 decides when it is useful to join the group conversation by generating output. In yet another example, the messaging application 114 is associated with a UX canvas displayed on the displays of the client computing devices 104 to 106, wherein the UX canvas is continuously updated by the output of the robot 116 (wherein the output is a suggestion for joining the group conversation). When a user clicks on a suggestion, such suggestion is entered into the group conversation as a turn (and identified as generated by the robot 116).

[0065] The bot 116 can assist users in group conversations with various tasks. Examples of tasks that the bot 116 can assist with include, but are not limited to: 1) summarizing the current conversation; 2) answering questions about a topic someone is talking about in the conversation; 3) suggesting follow-up actions based on information in the group conversation; 4) helping the group plan events, such as vacations; 5) generating images, avatars, memes, and the like related to messages included in the group conversation; 6) helping rewrite and review text that the group is working on—making the text more professional, making the text more humorous, correcting typos in the text, translating the text into different languages, and the like; 7) answering questions about webpage content pointed to by URLs in the group conversation, such as summarizing news articles, identifying facts mentioned in news articles, reading structured data such as company earnings documents, answering specific questions about webpage content, and the like; 8) brainstorming, such as creating a roadmap for product development; 9) translating text from one language to another, making the language of the participants in the group irrelevant because the bot 116 can translate and convey information between participants in the language requested by the participants; and 10) starting and running text- and image-based games to entertain users in the group conversation.

[0066] Furthermore, the bot 116 can generate text in output, such as answering questions in plain text, following the flow of conversation, and understanding pasted structured text, such as tables. In another example, the bot 116 accepts audio as input, allowing users to provide audio annotations in group conversations, and the bot 116 can use speech-to-text technology to understand the user's speech. The bot 116 can reply to the input text with text or using text-to-speech. By receiving audio, the bot 116 can listen to and understand audio / video calls, and can take URLs pasted into group conversations as input, allowing the bot to follow URLs, download HTML, and answer questions about the page content. In another example, the bot 116 accesses a link to a document in a shared storage space. Regarding output, the bot 116 can generate output in any of the modes mentioned above (text, audio, image, etc.). The bot 116 can generate output within an image it has already generated, and can populate emoticon templates, generate graphics, and so on. The bot 116 can generate output in the tone selected by the users in the group chat, and the bot's 116 output can be shared to other applications.

[0067] Figure 6 and Figure 7The present invention relates to a method for integrating a generative model into a computer-implemented messaging application according to one or more embodiments described herein. Although these methods are shown and described as a series of actions performed in sequence, it should be understood and appreciated that these methods are not limited to the order of the sequence. For example, some actions may occur in an order different from that described herein. In addition, an action may occur simultaneously with another action. Further, in some instances, not all actions are required to implement the methods described herein.

[0068] Furthermore, the actions described herein may be computer-executable instructions that can be implemented by one or more processors and / or stored on one or more computer-readable media. Computer-executable instructions may include routines, subroutines, programs, execution threads, etc. Furthermore, the results of the actions of these methods may be stored on computer-readable media, displayed on a display device, etc.

[0069] Now only reference Figure 6 , which illustrates a flow chart of a method 600 that facilitates integrating a bot into a computer-implemented messaging application that is illustrated as supporting group conversations. Method 600 begins at 602, and at 604, in a messaging application that supports group conversations, a first message is received from a first client computing device operated by a first user participating in a group conversation that includes several other users. At 606, the first message is transmitted to a client computing device operated by the user participating in the group conversation.

[0070] At 608, in the messaging application, a second message is received from a second client computing device operated by a second user participating in the group conversation. At 610, the second message is transmitted to the client computing device operated by the user participating in the group conversation. Thus, it can be determined that the group conversation includes messages generated by multiple users.

[0071] At 612, a third message is generated by a robot including a generative model, wherein the third message is generated based on the first message and the second message in the group conversation (i.e., the previous message in the group conversation). In the example, as previously described, the robot generates the third message in response to receiving a call request from a user participating in the group conversation. At 614, the third message is transmitted to a computing device operated by the user participating in the group conversation. Method 600 is completed at 616.

[0072] Now go to Figure 7, which illustrates a flow diagram of a method 700 for including a robot's output in a group conversation. Method 700 begins at 702, and at 704, an indication is received, via a computer-implemented messaging application, that a user participating in the group conversation has invoked a robot. At 706, in response to receiving the indication, a prompt is constructed for the robot, wherein the prompt includes two previously received messages from two different users in the group conversation. The two previously received messages can be represented by markers in the prompt. At 708, the robot receives the prompt, wherein the robot generates output based on the prompt. At 710, the output generated by the robot is transmitted to the client computing devices of the users participating in the group conversation. Method 700 completes at 710.

[0073] Now refer to Figure 8 , which illustrates a high-level diagram of an exemplary computing device 800 that can be used in accordance with the systems and methods disclosed herein. For example, the computing device 800 can be a client computing device having a client messaging application executing thereon. As another example, the computing device 800 can be a server computing system that executes a server-side messaging application and / or generates a model. The computing device 800 includes at least one processor 802 that executes instructions stored in a memory 804. The instructions can be, for example, instructions for implementing functions described as being performed by one or more of the components described above, or instructions for implementing one or more of the methods described above. The processor 802 can access the memory 804 via a system bus 806. In addition to storing executable instructions, the memory 804 can also store content, graphical icons, profile information, and the like.

[0074] The computing device 800 also includes a data store 808 that can be accessed by the processor 802 via the system bus 806. The data store 808 may include executable instructions, graphical icons, profile information, content, etc. The computing device 800 also includes an input interface 810 that allows external devices to communicate with the computing device 800. For example, the input interface 810 can be used to receive instructions from an external computer device, from a user, etc. The computing device 800 also includes an output interface 812 that interfaces the computing device 800 with one or more external devices. For example, the computing device 800 can display text, images, etc. via the output interface 812.

[0075] It is contemplated that external devices that communicate with the computing device 800 via the input interface 810 and the output interface 812 can be included in an environment that provides substantially any type of user interface with which a user can interact. Examples of user interface types include graphical user interfaces, natural user interfaces, and the like. For example, a graphical user interface can accept input from a user employing (multiple) input devices such as a keyboard, mouse, remote control, and provide output on an output device such as a display. Further, a natural user interface can enable a user to interact with the computing device 800 in a manner that is not constrained by input devices such as a keyboard, mouse, remote control, and the like. In contrast, a natural user interface can rely on voice recognition, touch and stylus recognition, gesture recognition on and adjacent to the screen, mid-air gestures, head and eye tracking, sound and voice, vision, touch, gestures, machine intelligence, and the like.

[0076] Additionally, although illustrated as a single system, it should be understood that computing device 800 may be a distributed system. Thus, for example, several devices may communicate via a network connection and may collectively perform the tasks described as being performed by computing device 800.

[0077] The various functions described herein can be implemented with hardware, software, or any combination thereof. If implemented with software, these functions can be stored on a computer-readable medium or transmitted on a computer-readable medium as one or more instructions or codes. Computer-readable media include computer-readable storage media. Computer-readable storage media can be any available storage medium that can be accessed by a computer. As an example and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to carry or store desired program codes in the form of instructions or data structures and that can be accessed by a computer. The disks and optical disks used herein include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks, and Blu-ray disks (BDs), wherein disks typically reproduce data magnetically and optical disks typically reproduce data optically by lasers. Further, propagation signals are not included within the scope of computer-readable storage media. Computer-readable media also include communication media, which include any media that facilitates the transfer of a computer program from one place to another. For example, a connection can be a communication medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of communications media. Combinations of the above may also be included within the scope of computer-readable media.

[0078] Alternatively or additionally, the functions described herein may be performed, at least in part, by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0079]

[0014] Disclosed herein are techniques for using generative models in a multi-user messaging context, according to at least the following examples.

[0080] (A1) In one aspect, a method disclosed herein includes receiving, in a messaging application supporting group conversations, a plurality of messages from a plurality of client computing devices operated by a plurality of users, wherein the plurality of messages are included in a group conversation. The method also includes providing a prompt to a generative model, wherein the prompt includes the plurality of messages. The method also includes receiving, from the generative model, output generated by the generative model, wherein the generative model generates the output based on the prompt. The method also includes including the output as a turn in the group conversation.

[0081] (A2) In some embodiments of the method of (A1), the plurality of users includes a first user. The method further includes receiving a message from the first user, wherein the message includes a request to invoke a chatbot in a group conversation and a request for information from the chatbot. The method additionally includes detecting the request to invoke the chatbot in the message. The method further includes constructing a prompt in response to detecting the request to invoke the chatbot, wherein the request for information is included in the prompt.

[0082] (A3) In some embodiments of the method of at least one of (A1)-(A2), the messaging application is an instant messaging application.

[0083] (A4) In some embodiments of the method of at least one of (A1)-(A3), the method further includes constructing a record of messages in the group conversation. The record includes: 1) a first message received from a first client computing device among the plurality of client computing devices; 2) a first identifier for a first user of the first client computing device, wherein the first identifier is assigned to the first message; 3) a second message received from a second client computing device among the plurality of client computing devices; and 4) a second identifier for a second user of the second client computing device, wherein the second identifier is assigned to the second message. The method additionally includes constructing a prompt based on the record of messages in the group conversation.

[0084] (A5) In some embodiments of the method of (A4), the record also includes: 5) a third message previously generated by the generative model; and 6) a third identifier for the generative model, wherein the third identifier is assigned to the third message, wherein the prompt provided to the generative model includes the first message, the first identifier assigned to the first message, the second message, the second identifier assigned to the second message, the third message, and the third identifier assigned to the third message.

[0085] (A6) In some embodiments of the method of (A4), the transcript of the messages in the group conversation further includes: 1) a third message previously generated by the generative model; and 2) a third identifier for the generative model, wherein the third identifier is assigned to the third message. The method further includes hiding the third message generated by the generative model from the transcript to create an updated transcript, wherein the third message is hidden based on the third identifier assigned to the third message, and wherein the prompt also includes the updated transcript.

[0086] (A7) In some embodiments of the method of at least one of (A1)-(A6), the method further includes receiving a command to invoke the generative model from the first client computing device, wherein the command to invoke the generative model includes a request for output from the generative model, wherein the command is received before providing the prompt to the generative model. The method additionally includes constructing a record of the group conversation, wherein the record includes a number of messages in the group conversation previously generated by the generative model. The method further includes comparing the number of messages in the number of messages to a predetermined threshold. The method further includes determining that the number of messages in the number of messages is equal to the predetermined threshold. The method additionally includes hiding the number of messages in the record previously generated by the generative model to create an updated record, wherein the prompt includes the updated record, wherein the number of messages is hidden when it is determined that the number of messages in the number of messages is equal to the predetermined threshold.

[0087] (A8) In some embodiments of the method of at least one of (A1)-(A7), the prompting includes generating an indication that the model is participating in a group conversation.

[0088] (A9) In some embodiments of the method of at least one of (A1)-(A8), the method further includes obtaining identifiers of the plurality of users from user profiles of the plurality of users, wherein the identifiers are obtained before providing the prompt to the generative model.

[0089] (A10) In some embodiments of the method of at least one of (A1)-(A9), the method further includes receiving a command from the first client computing device to invoke the generative model, wherein the command to invoke the generative model includes a request for output from the generative model, and further wherein the command is received before providing a prompt to the generative model. The method further includes constructing a prompt based on the request for output, wherein the generative model generates a query based on the prompt and provides the query to a search engine, wherein the generative model receives at least a portion of search results identified by the search engine based on the query, and further wherein the generative model generates the output based on at least a portion of the search results identified by the search engine.

[0090] (B1) On the other hand, a method performed by a computing system executing a messaging application includes: receiving a first message in a group conversation from a first client computing device, the first client computing device communicating with the computing system via the messaging application, wherein the first client computing device is operated by a first user. The method also includes receiving a second message in the group conversation from a second client computing device, the second client computing device communicating with the computing system via the messaging application, wherein the second client computing device is operated by a second user. The method additionally includes constructing a prompt based on the first message and the second message. The method also includes providing a prompt to a generative model, wherein a third message is generated by the generative model based on the prompt. The method also includes transmitting a third message to the first client computing device and the second client computing device for display as part of the group conversation, wherein the third message is identified in the group conversation as being generated by the generative model.

[0091] (B2) In some embodiments of the method of (B1), the method further includes receiving a command from the first client computing device to invoke the generative model, wherein the command to invoke the generative model includes a request for a fourth message from the generative model, wherein the command is received after the third message is transmitted to the first client computing device and the second client computing device. The method additionally includes constructing a second prompt based on the command, wherein the second prompt includes the third message. The method further includes providing a second prompt to the generative model, wherein the generative model generates a fourth message based on the second prompt. The method additionally includes transmitting the fourth message to the first client computing device and the second client computing device for presentation as part of the group conversation, wherein the fourth message is identified in the group conversation as being generated by the generative model.

[0092] (B3) In some embodiments of the method of at least one of (B1)-(B2), the prompt includes an identifier assigned to the first user of the first message and an identifier assigned to the second user of the second message.

[0093] (B4) In some embodiments of the method of at least one of (B1)-(B3), constructing the prompt includes: 1) constructing a transcript of the group conversation; and 2) truncating the transcript of the group conversation so that the number of tags in the truncated transcript is less than a predetermined threshold, wherein the prompt includes the truncated transcript.

[0094] (B5) In some embodiments of the methods of at least one of (B1)-(B4), the method includes receiving a command from a first client computing device to invoke the generative model, wherein the command to invoke the generative model includes a request for a fourth message from the generative model, and further wherein the command is received after the third message is transmitted to the first client computing device and the second client computing device. The method additionally includes constructing a second prompt based on the command to invoke the generative model, wherein the second prompt includes the first message and the second message but does not include the third message. The method also includes providing a second prompt to the generative model, wherein the generative model generates the fourth message based on the second prompt. The method also includes transmitting the fourth message to the first client computing device and the second client computing device for presentation as part of the group conversation, wherein the fourth message is identified in the group conversation as generated by the generative model.

[0095] (B6) In some embodiments of the method of at least one of (B1)-(B5), the method further includes obtaining a first identifier of the first user and a second identifier of the second user from the first user profile and the second user profile, respectively, and wherein constructing the prompt includes including the first user identifier and the second user identifier in the prompt, wherein the generation model generates the third message based on at least one of the first user identifier or the second user identifier.

[0096] (B7) In some embodiments of the method of at least one of (B1)-(B6), the generative model generates an overview summarizing the first message and the second message based on the first message and the second message, and further wherein constructing the prompt includes including the overview in the prompt.

[0097] (B8) In some embodiments of the method of at least one of (A1)-(A7), the method further includes obtaining a topic identified in a user profile of the first user, wherein constructing the prompt includes including the topic in the prompt.

[0098] (C1) In another aspect, a computing system includes a processor and a memory, wherein the memory stores instructions that, when executed by the processor, cause the processor to perform at least one of the methods disclosed herein (e.g., any of the methods (A1)-(10) or (B1)-(B8)).

[0099] (D1) On the other hand, a computer-readable storage medium includes instructions that, when executed by a processor, cause the processor to perform at least one of the methods disclosed herein (e.g., any one of methods (A1)-(10) or (B1)-(B8)).

[0100] What has been described above includes examples of one or more embodiments. Of course, it is not possible to describe every conceivable modification and alteration of the apparatus or method above for the purposes of describing the aforementioned aspects, but those skilled in the art will recognize that many further modifications and permutations of the various aspects are possible. Accordingly, the various aspects described are intended to cover all such alterations, modifications, and variations that fall within the spirit and scope of the appended claims. Furthermore, to the extent that the term "includes" is used in the detailed description or the claims, such term is intended to be inclusive, similar to how the term "comprising" is understood when used as a transition word in a claim.

Claims

1. A computing system comprising: processor; as well as a memory storing instructions that, when executed by the processor, cause the processor to perform actions, the actions comprising: receiving, in a messaging application supporting group conversations, a plurality of messages from a plurality of client computing devices operated by a plurality of users, wherein the plurality of messages are included in a group conversation; providing a prompt to a generative model, wherein the prompt comprises the plurality of messages; receiving, from the generative model, an output generated by the generative model, wherein the generative model generates the output based on the prompt; and The output is included as a turn in the group conversation.

2. The computing system of claim 1 , wherein the plurality of users includes a first user, the actions further comprising: receiving a message from the first user, wherein the message includes a request to invoke a chatbot in the group conversation and a request for information from the chatbot; detecting, in the message, the request to invoke the chatbot; as well as The prompt is constructed in response to detecting the request to invoke the chatbot, wherein the request for information is included in the prompt. 3 . The computing system of claim 1 , wherein the messaging application is an instant messaging application.

4. The computing system of claim 1 , wherein the actions further comprise: Constructing a record of messages in the group conversation, wherein the record includes: a first message received from a first client computing device among the plurality of client computing devices; a first identifier for a first user of the first client computing device, wherein the first identifier is assigned to the first message; a second message received from a second client computing device of the plurality of client computing devices; a second identifier for a second user of the second client computing device, wherein the second identifier is assigned to the second message; and The prompt is constructed based on the record of the messages in the group conversation.

5. The computing system of claim 4, wherein the record of the messages in the group conversation further comprises: a third message previously generated by the generative model; as well as a third identifier for the generative model, wherein the third identifier is assigned to the third message, wherein the prompt provided to the generative model includes the first message, the first identifier assigned to the first message, the second message, the second identifier assigned to the second message, the third message, and the third identifier assigned to the third message.

6. The computing system of claim 4, wherein the record of the messages in the group conversation further comprises: a third message previously generated by the generative model; as well as a third identifier for the generated model, wherein the third identifier is assigned to the third message, wherein the actions further comprise: The third message generated by the generative model from the record is suppressed to create an updated record, wherein the third message is suppressed based on the third identifier assigned to the third message, and wherein the prompt also includes the updated record.

7. The computing system of claim 1 , wherein the actions further comprise: prior to providing the hint to the generative model, receiving a command from the first client computing device to invoke the generative model, wherein the command to invoke the generative model includes a request for the output from the generative model; constructing a record of the group conversation, wherein the record includes a plurality of messages in the group conversation that were previously generated by the generative model; comparing a number of messages in the plurality of messages with a predetermined threshold; determining that the number of messages in the plurality of messages is equal to the predetermined threshold; and Upon determining that the number of messages in the plurality of messages is equal to the predetermined threshold, hiding the plurality of messages in the record previously generated by the generative model to create an updated record, wherein the prompt includes the updated record.

8. The computing system of claim 1, wherein the prompt comprises an indication that the generative model is participating in the group conversation.

9. The computing system of claim 1 , wherein the actions further comprise: prior to providing the prompt to the generative model, obtaining identifiers for the plurality of users from user profiles of the plurality of users; as well as The prompt is constructed to include the identifiers for the plurality of users.

10. The computing system of claim 1, wherein the actions further comprise: prior to providing the hint to the generative model, receiving a command from the first client computing device to invoke the generative model, wherein the command to invoke the generative model includes a request for the output from the generative model; as well as The prompt is constructed based on the request for the output, wherein the generative model generates a query based on the prompt and provides the query to a search engine, wherein the generative model receives at least a portion of search results identified by the search engine based on the query, and wherein the generative model also generates the output based on the at least a portion of the search results identified by the search engine.

11. A method performed by a computing system executing a messaging application, the method comprising: receiving a first message in the group conversation from a first client computing device, the first client computing device communicating with the computing system via the messaging application, wherein the first client computing device is operated by a first user; receiving a second message in the group conversation from a second client computing device that is in communication with the computing system via the messaging application, wherein the second client computing device is operated by a second user; constructing a prompt based on the first message and the second message; providing the prompt to a generative model, wherein the generative model generates a third message based on the prompt; as well as The third message is transmitted to the first client computing device and the second client computing device for display as part of the group conversation, wherein the third message is identified in the group conversation as being generated by the generative model.

12. The method according to claim 11, further comprising: after transmitting the third message to the first and second client computing devices, receiving a command from the first client computing device to invoke the generative model, wherein the command to invoke the generative model includes a request for a fourth message from the generative model; constructing a second prompt based on the command, wherein the second prompt includes the third message; providing the second prompt to the generative model, wherein the generative model generates the fourth message based on the second prompt; as well as The fourth message is transmitted to the first and second client computing devices for presentation as part of the group conversation, wherein the fourth message is identified in the group conversation as being generated by the generative model.

13. The method of claim 11, wherein the prompt comprises an identifier for the first user assigned to the first message and an identifier for the second user assigned to the second message.

14. The method of claim 11 , wherein constructing the prompt comprises: constructing a record of the group conversation; as well as The transcript of the group conversation is truncated so that a number of tags in the truncated transcript is less than a predetermined threshold, wherein the prompt includes the truncated transcript.

15. The method according to claim 11, further comprising: after transmitting the third message to the first and second client computing devices, receiving a command from the first client computing device to invoke the generative model, wherein the command to invoke the generative model includes a request for a fourth message from the generative model; constructing a second prompt based on the command that invokes the generative model, wherein the second prompt includes the first message and the second message but does not include the third message; providing the second prompt to the generative model, wherein the generative model generates the fourth message based on the second prompt; as well as The fourth message is transmitted to the first client computing device and the second client computing device for presentation as part of the group conversation, wherein the fourth message is identified in the group conversation as being generated by the generative model.