Generating summary data from audio or video data in a group-based communication system

By incorporating non-verbal signals, the transcription and summarization of multimedia collaboration sessions in group-based communication platforms are improved, addressing the limitations of existing methods.

JP2026500105APending Publication Date: 2026-01-06SALESFORCE INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025530272
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2023-11-07
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing speech-to-text transcription methods fail to robustly contextualize audio and video data in the presence of non-verbal signals, making it difficult to efficiently summarize multimedia collaboration sessions in group-based communication platforms.

Method used

A combination of signals, including non-verbal signals, is used to transcribe and summarize synchronous or asynchronous multimedia collaboration sessions in group-based communication platforms, leveraging machine learning models to enhance the contextualization of audio-video data.

Benefits of technology

The solution effectively addresses the contextualization of multimedia collaboration sessions, enhancing the transcription and summarization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026500105000001_ABST
    Figure 2026500105000001_ABST
Patent Text Reader

Abstract

Techniques for transcribing and / or summarizing a multimedia collaboration session are discussed herein. For example, users may communicate within a teleconference conference associated with a channel. In some examples, a first machine learning model may be configured to receive audiovisual data and user interaction data (e.g., selected emojis, detected gestures, messages or text entered by users, message threads, etc.) and output a teleconference conference summary. In some examples, the teleconference conference summary may include one or more highlights, action items, and / or artificial intelligence (AI) notes generated using the machine learning model(s).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Related Applications] This application claims priority to U.S. Patent Application No. 18 / 385,219, filed October 30, 2023, entitled "GENERATING SUMMARY DATA FROM AUDIO DATA OR VIDEO DATA IN A GROUP-BASED COMMUNICATION SYSTEM," which claims priority to and benefit of U.S. Provisional Application No. 63,429,082, filed November 30, 2022, all of which are incorporated herein by reference.

[0002] [Technical field] Communication platforms have become increasingly popular for facilitating work-related communications, such as project collaboration within a single organization and between two or more organizations. To facilitate project collaboration, users often share information and engage in ad-hoc discussions within the communication platform. However, reviewing conversations within a communication platform can require a lot of time to extract important parts of the conversation. Existing systems may allow users to search for conversations based on keywords, search functions, etc. to access relevant parts of the conversation, or may provide transcripts of the ad-hoc discussions.

[0003] In speech-to-text transcription, neural networks classify spoken sounds as either containing phonemes (units of speech) or not, and then match the phonemes to known words, phrases, and sentences based on models previously trained on sets of audio clips of the speech and text transcripts of the speech. However, existing speech-to-text transcription methods may not robustly contextualize audio (and / or video) data in the presence of accompanying non-verbal signals. [Brief explanation of the drawings]

[0004] The detailed description will be set forth with reference to the accompanying drawings. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. When the same reference number is used in different figures, it indicates a similar or identical component or feature. The drawings are not drawn to scale.

[0005] [Figure 1] 1 illustrates an exemplary system for performing the techniques described herein. [Figure 2A] For a specific example, a user interface for a group-based communication system is shown. [Figure 2B] For a particular example, a user interface for a multimedia collaboration session within a group-based communication system is shown. [Figure 2C] For a specific example, a user interface for inter-organizational collaboration within a group-based communication system is shown. [Figure 2D] For a specific example, a user interface for a collaborative document within a group-based communication system is shown. [Figure 3A] 1 illustrates a user interface for a workflow within a group-based communication system. [Figure 3B] 1 illustrates a block diagram for implementing certain examples as discussed herein. [Figure 4] For a specific example, a diagram of a machine learning model that can be trained to receive an input and output a summary document using that input is shown. [Figure 5] 1 illustrates an exemplary process associated with generating a summary document for a synchronous multimedia collaboration session in a virtual space. [Figure 6]1 illustrates an exemplary user interface for requesting and generating a summary of a synchronized multimedia collaboration session and for generating notes during a synchronized multimedia collaboration session. [Figure 7] 1 illustrates an exemplary user interface for requesting and generating a summary of a message thread, as discussed herein. [Figure 8] 1 illustrates an exemplary process associated with generating a summary of a virtual space as discussed herein. DETAILED DESCRIPTION OF THE INVENTION

[0006] The following detailed description addresses the above deficiencies by using a combination of signals, including non-verbal signals, that contextualize audio-video data to more robustly transcribe and / or summarize synchronous or asynchronous multimedia collaboration sessions in a group-based communication platform.

[0007] The following detailed description of examples refers to the accompanying drawings, which show specific examples in which the present techniques may be practiced. These examples are intended to describe aspects of the systems and methods in sufficient detail to enable those skilled in the art to practice the techniques discussed herein. Other examples may be utilized and modifications may be made without departing from the scope of the present disclosure. Therefore, the following detailed description should not be construed in a limiting sense. The scope of the present disclosure is defined only by the appended claims, along with the full scope of equivalents to which such claims are entitled. Group-based communication system

[0008] FIG. 1 illustrates an exemplary environment 100 for performing the techniques described herein. In at least one example, the exemplary environment 100 may be associated with a communications platform that can leverage a network-based computing system to enable users of the communications platform to exchange data. In at least one example, the communications platform may be “group-based” such that the platform and associated systems, communication channels, messages, collaborative documents, canvases, audio / video conversations, and / or other virtual spaces have security (which may be defined by permissions) to limit access to a defined group of users. In some examples, such groups of users may be defined by a group identifier that may be associated with common access credentials, domains, etc., as described above. In some examples, the communications platform may be a hub that provides a secure, private virtual space for users to chat, meet, call, collaborate, transfer files or other data, or otherwise communicate among one another. As described above, each group may be associated with a workspace, enabling users associated with the group to chat, meet, call, collaborate, transfer files or other data, or otherwise communicate among one another within the secure, private virtual space. In some examples, members of a group, and thus a workspace, may be associated with the same organization. In some examples, members of a group, and thus a workspace, may be associated with different organizations (e.g., entities with different organization identifiers).

[0009] In at least one example, the exemplary environment 100 may include one or more server computing devices (or “server(s)”) 102. In at least one example, the server(s) 102 may include one or more servers or other types of computing devices that may be embodied in any number of ways. For example, in the server example, the functional components and data may be implemented on a single server, a cluster of servers, a server farm or data center, a cloud-hosted computing service, a cloud-hosted storage service, etc., although other computer architectures may additionally or alternatively be used.

[0010] In at least one example, the server(s) 102 can communicate with the user computing devices 104 via one or more network(s) 106. That is, the server(s) 102 and the user computing devices 104 can send, receive, and / or store data (e.g., content, information, etc.) using the network(s) 106 as described herein. The user computing devices 104 can be any suitable type of computing device, e.g., portable, semi-portable, semi-stationary, or stationary. Some examples of the user computing devices 104 can include tablet computing devices, smartphones, mobile communication devices, laptops, netbooks, desktop computing devices, terminal computing devices, wearable computing devices, augmented reality devices, Internet of Things (IoT) devices, or any other computing device capable of transmitting communications and performing functions in accordance with the techniques described herein. While a single user computing device 104 is shown, in practice, the exemplary environment 100 can include multiple (e.g., tens, hundreds, thousands, or millions) user computing devices. In at least one example, a user computing device, such as user computing device 104, may be operable by a user to, among other things, access communication services via a communications platform. A user may be an individual, a group of individuals, an employer, a business, an organization, and / or the like.

[0011] The network(s) 106 may include any type of network known in the art, such as, but not limited to, a local or wide area network, the Internet, a wireless network, a cellular network, a local wireless network, Wi-Fi and / or short-range wireless communication, Bluetooth®, Bluetooth Low Energy (BLE), Near Field Communication (NFC), a wired network, or any other such network, or any combination thereof. The components used for such communication may depend at least in part on the type of network, the selected environment, or both. Protocols for communicating over such network(s) 106 are well known and will not be described in detail herein.

[0012] In at least one example, server(s) 102 may include one or more processors 108 , computer-readable media 110 , one or more communication interfaces 112 , and / or input / output devices 114 .

[0013] In at least one example, each processor of processor(s) 108 can be a single processing unit or multiple processing units and can include single or multiple computing units or multiple processing cores. Processor(s) 108 can be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units (CPUs), graphics processing units (GPUs), state machines, logic circuits, and / or any device that manipulates signals based on operational instructions. For example, processor(s) 108 can be one or more hardware processors and / or logic circuits of any suitable type specifically programmed or configured to execute the algorithms and processes described herein. Processor(s) 108 can be configured to fetch and execute computer-readable instructions stored on a computer-readable medium, which can program the processor(s) to perform the functions described herein.

[0014] The computer-readable medium 110 can include volatile and nonvolatile memory and / or removable and non-removable media implemented in any type of technology for storing data, such as computer-readable instructions, data structures, program modules, or other data. Such computer-readable medium 110 can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, optical storage, solid-state storage, magnetic tape, magnetic disk storage, RAID storage systems, storage arrays, network-attached storage, storage area networks, cloud storage, or any other medium usable to store desired data and accessible by a computing device. Depending on the configuration of the server(s) 102, the computer-readable medium 110 can be a type of computer-readable storage medium and / or can be a tangible, non-transitory medium, insofar as non-transitory computer-readable medium, when referred to, excludes media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0015] The computer-readable medium 110 may be used to store any number of functional components executable by the processor(s) 108. In many implementations, these functional components include instructions or programs executable by the processor(s) 108 that, when executed, specifically configure the processor(s) 108 to perform the actions attributed above to the server(s) 102. The functional components stored on the computer-readable medium may optionally include a messaging component 116, an audio / video component 118, a summarization engine 120, an operating system 122, and a data store 124.

[0016] In at least one example, the messaging component 116 can process messages between users. That is, in at least one example, the messaging component 116 can receive an outgoing message from a user computing device 104 and transmit the message as an incoming message to a second user computing device. The message can include a direct message sent from the originating user to one or more designated users and / or a communication channel message sent from the originating user via the communication channel to one or more users associated with the communication channel. Additionally, the message can be transmitted in association with a collaborative document, canvas, or other collaborative space. In at least one example, the canvas can include a flexible canvas for curating, organizing, and sharing a collection of information between users. In at least one example, the collaborative document can be associated with a document identifier (e.g., a virtual space identifier, a communication channel identifier, etc.) configured to enable messaging functionality attributable to the virtual space (e.g., the communication channel) of the collaborative document. That is, the collaborative document can be treated as a virtual space, such as a communication channel, and can include functionality associated therewith. The virtual space or communication channel can be a data pathway used to exchange data between systems and devices associated with the communication platform.

[0017] In at least one example, the messaging component 116 can establish communication paths between various user computing devices, enabling the user computing devices to communicate and share data among one another. In at least one example, the messaging component 116 can manage such communication and / or data sharing. In some examples, data associated with the virtual space, such as collaborative documents, can be presented via a user interface. Additionally, metadata associated with each message sent through the virtual space can be stored in association with the virtual space, such as a timestamp associated with the message, a sending user identifier, a receiving user identifier, a conversation identifier and / or a root object identifier (e.g., a conversation associated with a thread and / or root object), and / or the like.

[0018] In various examples, the messaging component 116 can receive a message (e.g., a direct message instance, a communication channel, a canvas, a collaborative document, etc.) sent in association with the virtual space. In various examples, the messaging component 116 can identify one or more users associated with the virtual space and cause the rendering of a message in association with an instance of the virtual space on each user computing device 104. In various examples, the messaging component 116 can identify the message as an update to the virtual space and, based on the identified update, cause a notification associated with the update to be presented in association with a sidebar of a user interface associated with one or more of the user(s) associated with the virtual space. For example, the messaging component 116 can receive a message sent in association with the virtual space from a first user account. In response to receiving the message (e.g., interaction data associated with the first user's interaction with the virtual space), the messaging component 116 can identify a second user associated with the virtual space (e.g., another user who is a member of the virtual space). In some examples, the messaging component 116 can cause a notification of an update to the virtual space to be presented via a sidebar of a user interface associated with a second user account of the second user. In some examples, the messaging component 116 can cause the notification to be presented in response to determining that a sidebar of a user interface associated with the second user account includes an affordance associated with the virtual space. In such examples, the notification can be presented in association with the affordance associated with the virtual space.

[0019] In various examples, the messaging component 116 can be configured to identify a mention or tag associated with a message sent in association with a virtual space. In at least one example, the mention or tag can include an @mention (or other special character) of a user identifier associated with the communication platform. The user identifier can include a username, real name, or other unique identifier associated with a particular user. In response to identifying a mention or tag of a user identifier, the messaging component 116 can cause a notification to be presented on a user interface associated with the user identifier, e.g., in a sidebar of a user interface associated with the particular user and / or in a virtual space associated with mentions and reactions, associated with affordances associated with the virtual space. That is, the messaging component 116 can be configured to alert a particular user that they have been mentioned in a virtual space.

[0020] In at least one example, the audio / video component 118 may be configured to manage audio and / or video communications between users. In some examples, the audio and / or video communications may be associated with audio and / or video conversations. In at least one example, the audio and / or video conversations may include individual identifiers configured to uniquely identify the audio and / or video conversations. In some examples, the audio and / or video component 118 may store user identifiers associated with user accounts of members of particular audio and / or video conversations, for example, to identify user(s) who have appropriate permissions to access the particular audio and / or video conversations.

[0021] In some examples, communications associated with an audio and / or video conversation (“conversation”) can be synchronous and / or asynchronous. That is, a conversation can include a real-time audio and / or video conversation between a first user and a second user for a period of time, and after the first period, a third user associated with (e.g., a member of) the conversation can contribute to the conversation. Audio / video component 118 can be configured to store audio and / or video data associated with a conversation, for example, to allow users with appropriate permissions to listen to and / or view the audio and / or video data.

[0022] In some examples, the audio / video component 118 may be configured to generate a transcript of the conversation and may further store the transcript in association with the audio and / or video data. The transcript may include a text representation of the audio and / or video data. In at least one example, the audio / video component 118 may generate the transcript using known speech recognition techniques. In some examples, the audio / video component 118 may generate the transcript simultaneously or substantially simultaneously with the conversation. That is, in some examples, the audio / video component 118 may be configured to generate a text representation of the conversation while the conversation is taking place. In some examples, the audio / video component 118 may generate the transcript after receiving an indication that the conversation is complete. The indication that the conversation is complete may include an indication that an associated host or administrator has stopped the conversation, a threshold number of conference attendees have closed an associated interface, and / or the like. That is, the audio / video component 118 may identify the completion of a conversation and, based on that completion, generate a transcript associated therewith.

[0023] In at least one example, the audio / video component 118 can be configured to cause presentation of a transcript in association with a virtual space with which the audio and / or video conversation is associated. For example, a first user can initiate an audio and / or video conversation in association with a communication channel. The audio / video component 118 can process audio and / or video data between participants in the audio and / or video conversation and further generate a transcript of the audio and / or video data. In response to generating the transcript, the audio / video component 118 can cause the transcript to be published or otherwise presented via the communication channel. In at least one example, the audio / video component 118 can render one or more sections of the transcript selectable for commenting, for example, to enable members of the communication channel to comment on or further contribute to the conversation. In some examples, the audio / video component 118 can update the transcript based on the comments.

[0024] In at least one example, the audio / video component 118 can manage one or more audio and / or video conversations associated with a virtual space associated with a group (e.g., organization, team, etc.) management or command center. The group management or command center can be referred to herein as a virtual (and / or digital) headquarters associated with the group. In at least one example, the audio / video component 118 can be configured to coordinate with the messaging component 116 and / or other components of the server(s) 102 to send communications associated with other virtual spaces associated with the virtual headquarters. That is, the messaging component 116 can send data (e.g., messages, images, drawings, files, etc.) associated with one or more communication channels, direct messaging instances, collaborative documents, canvases, and / or the like associated with the virtual headquarters. In some examples, the communication channel(s), direct messaging instance(s), collaborative document(s), canvas(s), and / or the like can be associated with one or more audio and / or video conversations managed by the audio / video component 118. That is, audio and / or video conversations associated with a virtual headquarters may be further associated with or unrelated to one or more other virtual spaces of the virtual headquarters.

[0025] In at least one example, summarization engine 120 can manage transcription and / or summarization of synchronous and / or asynchronous multimedia collaboration sessions. Summarization engine 120 may be tasked with transcribing and / or summarizing a multimedia collaboration session and generating one or more summary documents representing actionable tasks and / or deadlines that characterize the multimedia collaboration session. Summarization engine 120 may utilize machine learning (ML) model(s) 142 (or MLM) that accepts input and uses the input to output such summary document(s). In various examples, virtual spaces (e.g., group DMs, direct DMs, communication channels, etc.) may be created or otherwise accessed to discuss tasks and deadlines related to various purposes, and within the virtual spaces, users may collaborate with one another via synchronous multimedia collaboration sessions regarding achieving the purposes. This may include users posting messages to the virtual spaces, posting content (e.g., audio, video, links, files, etc.) to the virtual spaces, etc. Thus, in at least some examples, input to ML model(s) 142 may include ambient data associated with a synchronous multimedia collaboration session involving real-time communications occurring within a virtual space between other users. However, ML model(s) 142 may have access to a variety of other data from which to generate one or more summary documents, such as relationship data between participants in the multimedia session, contextual data for the virtual space (e.g., virtual space data, channel data, message data, thread data, emoji data, reaction data, etc.), organizational data, etc.

[0026] The ML model(s) 142 may be trained to transcribe and / or summarize communications within the virtual space. For example, the ML model(s) 142 may extract significant portions of conversations within the virtual space based on contextual clues, the frequency of communications by the user(s), past feedback, communications (or other user interactions) above a threshold level, user communications marked as favorites, task assignments to the user, user ratings, user expertise, suggested user preferences for each role, user-specified parameters, user-specified permissions, heuristics from user activity, emojis, reactions, and / or other user interactions with previously generated summary documents. Furthermore, the ML model(s) 142 may be trained to identify one or more users associated with the task or assignment containing the summary document(s), such as users mentioned in the virtual space, users who posted to the virtual space, or users who are members of the virtual space. The ML model(s) 142 may also be trained to identify respective messages, contributions, posts, etc. for user(s) within the virtual space. Thus, ML model(s) 142 may infer actionable context from the inputs it takes in and output summary document(s) representing such actionable context. For example, ML model(s) 142 may be trained to infer from the frequency of communication between two users exceeding a threshold level and other contextual clues that a first user and a second user represent a task assignor and assignee.

[0027] In some examples, the communication platform can manage communication channels. In some examples, the communication platform can be a channel-based messaging platform that, in some examples, can be usable by a group(s) of users. Users of the communication platform can communicate with other users through communication channels. A communication channel or virtual space can be a data pathway used to exchange data between systems and devices associated with the communication platform. In some examples, a channel can be a virtual space where people can post messages, documents, and / or files. In some examples, access to a channel can be controlled by permissions. In some examples, a channel can be limited to a single organization, shared between different organizations, or a public, private, or specialized channel (e.g., a hosted channel with a guest account where guests can make posts but cannot perform certain actions such as inviting other users to the channel). In some examples, some users can be invited to a channel via email, channel invite, direct message, text message, etc. Examples of channels and related functionality are described throughout this disclosure.

[0028] In at least one example, the operating system 122 may manage the processor(s) 108, computer-readable media 110, hardware, software, etc. of the server(s) 102.

[0029] In at least one example, data store 124 may be configured to store accessible, manageable, and updatable data. In some examples, data store 124 may be integrated with server(s) 102, as shown in FIG. 1 . In other examples, data store 124 may be located remotely from server(s) 102 and may be accessible to server(s) 102 and / or user device(s), such as user computing device 104. Data store 124 may include multiple databases, which may include at least first training data 126 and second training data 128. Additional or alternative data may be stored in data store 124 and / or one or more other data stores. For example, data store 124 may alternatively or additionally store one or more summary documents (not shown) therein.

[0030] In at least one example, user / organization data 129 can include data associated with a user of the communication platform. In at least one example, user / organization data 129 can store data in a user profile (which may also be referred to as a “user account”), which can store data associated with a user, including, but not limited to, one or more user identifiers associated with multiple different organizations or entities with which the user is associated, one or more communication channel identifiers associated with communication channels to which the user is authorized to access, one or more group identifiers for groups (or organizations, teams, entities, etc.) with which the user is associated, an indication of whether the user is the owner or manager of any communication channels, an indication of whether the user has any communication channel restrictions, a plurality of messages, a plurality of emojis, a plurality of conversations, a plurality of conversation topics, an avatar, an email address, a real name (e.g., John Doe), a username (e.g., j doe), a password, a time zone, a status, a token, etc.

[0031] In at least one example, user / organization data 129 can include permission data associated with permissions of individual users of the communications platform. In some examples, permissions can be set automatically or by an administrator of the communications platform, an employer, company, organization, or other entity using the communications platform, a team leader, group leader, or other entity using the communications platform to communicate with team members, group members, etc., an individual user, etc. Permissions associated with individual users can be mapped to or otherwise associated with an account or profile within user / organization data 129. In some examples, permissions can indicate which users can communicate directly with other users, which channels a user is authorized to access, restrictions on individual channels, which workspaces a user is authorized to access, restrictions on individual workspaces, etc. In at least one example, permissions can support the communications platform by maintaining security for limiting access to defined groups of users. In some examples, such users can be defined by common access credentials, group identifiers, etc., as described above.

[0032] In at least one example, user / organization data 129 can include data associated with one or more organizations of the communication platform. In at least one example, user / organization data 129 can store data in an organization profile, which can store data associated with the organization, including, but not limited to, one or more user identifiers associated with the organization, one or more virtual space identifiers associated with the organization (e.g., workspace identifiers, communication channel identifiers, direct message instance identifiers, collaborative document identifiers, canvas identifiers, audio / video conversation identifiers, etc.), organization identifiers associated with the organization, one or more organization identifiers associated with other organizations authorized to communicate with the organization, etc.

[0033] In at least one example, the virtual space data may include data associated with one or more virtual spaces associated with the communication platform. The virtual space data may include text data, audio data, video data, images, files, and / or any other type of data configured to be transmitted in association with a virtual space. Non-limiting examples of virtual spaces include workspaces, communication channels, direct messaging instances, collaborative documents, canvases, and audio and / or video conversations. In at least one example, the virtual space data may store data associated with individual virtual spaces separately, for example, based on a distinct identifier associated with each virtual space. In some examples, a first virtual space may be associated with a second virtual space. In such examples, first virtual space data associated with the first virtual space may be stored in association with the second virtual space. For example, data associated with a collaborative document generated in association with a communication channel may be stored in association with the communication channel. In another example, data associated with an audio and / or video conversation occurring in association with a communication channel may be stored in association with the communication channel.

[0034] As described above, each virtual space in a communication platform may be assigned an individual identifier that uniquely identifies the virtual space. In some examples, the virtual space identifier associated with a virtual space may include a physical address within the virtual space where data related to the virtual space is stored. A virtual space may be “public,” which may allow any user within an organization (e.g., associated with an organization identifier) ​​to participate in data sharing via the virtual space, or it may be “private,” which may restrict data communication within the virtual space to specific users or users with appropriate viewing permissions. In some examples, a virtual space may be “shared,” which may allow users associated with different organizations (e.g., entities associated with different organization identifiers) to participate in data sharing via the virtual space. A shared virtual space (e.g., a shared channel) may be public, accessible to any user from either organization, or private, restricted to access by specific users from both organizations (e.g., users with appropriate permissions).

[0035] In some examples, the data store 124 may be divided into separate items of data (e.g., data shards) that can be accessed and managed individually. Data shards can simplify many technical tasks, such as data retention, deployment (e.g., detecting that message content contains links, crawling metadata for links, and determining a uniform summary of the metadata), and integration configuration. In some examples, data shards may be associated with organizations, groups (e.g., workspaces), communication channels, users, etc.

[0036] In some examples, individual organizations may be associated with database shards within the data store 124 that store data related to a particular organization identification. For example, a database shard may store electronic communication data associated with members of a particular organization, allowing members of that particular organization to communicate and exchange data with other members of the same organization in real time or near real time. In this example, the organization itself may be the owner of the database shard and control where and how the associated data is stored. In some examples, a database shard may store data related to two or more organizations (e.g., as in a shared virtual space).

[0037] In some examples, individual groups may be associated with database shards in data store 124 that store data related to a particular group identity (e.g., a workspace). For example, a database shard may store electronic communication data associated with members of a particular group, allowing members of that particular group to communicate and exchange data with other members of the same group in real time or near real time. In this example, the group itself may be the owner of the database shard and control where and how the associated data is stored.

[0038] In some examples, a virtual space may be associated with a database shard in data store 124 that stores data related to a particular virtual space identification. For example, the database shard may store electronic communication data associated with the virtual space, allowing members of that particular virtual space to communicate and exchange data with other members of the same virtual space in real time or near real time. As described above, communication through a virtual space can be synchronous and / or asynchronous. In at least one example, a group or organization may be the owner of a database shard and may control where and how associated data is stored.

[0039] In some examples, individual users may be associated with database shards within data store 124 that store data related to a particular user account. For example, a database shard may store electronic communication data associated with an individual user, allowing the user to communicate and exchange data with other users of the communications platform in real time or near real time. In some examples, the users themselves may be owners of database shards, controlling where and how their associated data is stored.

[0040] In some examples, for example, when a channel is shared between two organizations, each organization may be associated with its own encryption key. When a user associated with one organization posts a message or file to a shared channel, the message or file is encrypted in data store 124 with an encryption key specific to that organization, and the other organization can decrypt the message or file before accessing it. Furthermore, in examples where the organizations are in different geographic areas, data associated with a particular organization may be stored in a location corresponding to the organization and temporarily cached in a location closer to a client (e.g., associated with the other organization) when such a message or file is accessed. Data may be maintained, stored, and / or deleted in data store 124 according to data management policies associated with each particular organization.

[0041] The communication interface(s) 112 may include one or more interfaces and hardware components to enable communication with various other devices (e.g., user computing device 104), for example, over the network(s) 106 or directly. In some examples, the communication interface(s) 112 may facilitate communication via WebSockets, an application programming interface (API) (e.g., using API calls), Hypertext Transfer Protocol (HTTPs), etc.

[0042] The server(s) 102 may further comprise various input / output devices 114 (e.g., I / O devices). Such I / O devices 114 may include a display, various user interface controls (e.g., buttons, joysticks, keyboards, mice, touchscreens, etc.), audio speakers, connection ports, etc.

[0043] In at least one example, user computing device 104 may include one or more processors 130 , computer-readable media 132 , one or more communication interfaces 134 , and input / output devices 136 .

[0044] In at least one example, each processor of processor(s) 130 may be a single processing unit or multiple processing units and may further include single or multiple computing units or multiple processing cores. Processor(s) 130 may include any of the types of processors described above with reference to processor(s) 108, and may be the same as or different from processor(s) 108.

[0045] Computer-readable medium 132 may include any of the types of computer-readable medium 132 described above with reference to computer-readable medium 110, and may be the same as or different from computer-readable medium 110. Functional components stored on the computer-readable medium may optionally include at least one application 138 and an operating system 140.

[0046] In at least one example, the application 138 can be a mobile application, a web application, or a desktop application, which may be provided by the communications platform or may be another dedicated application. In some examples, individual user computing devices associated with the environment 100 may have an instance or versioned instance of the application 138 that may be downloaded from an application store, accessible via the internet, or otherwise executable by the processor(s) 130 to perform operations as described herein. That is, the application 138 may be an access point that allows the user computing device 104 to interact with the server(s) 102 to access and / or use communication services available through the communications platform. In at least one example, the application 138 may facilitate the exchange of data between various other user computing devices, for example, via the server(s) 102. In at least one example, the application 138 may present a user interface as described herein. In at least one example, a user may interact with the user interface via touch input, keyboard input, mouse input, verbal input, or any other type of input.

[0047] A non-limiting example of user interface 143 is shown in FIG. 1. As shown in FIG. 1, user interface 143 can present data associated with one or more virtual spaces, which may include one or more workspaces. That is, in some examples, user interface 143 can integrate data from multiple workspaces into a single user interface, thereby enabling a user (e.g., of user computing device 104) to access and / or interact with data associated with multiple workspaces with which the user is associated and / or otherwise communicate with other users associated with the multiple workspaces. In some examples, user interface 143 can include a first region 144 or pane that includes indicator(s) (e.g., user interface element(s) or object(s)) associated with workspace(s) with which the user (e.g., the user's account) is associated. In some examples, user interface 143 can include a second region 146 or pane that includes indicator(s) (e.g., user interface element(s), affordance(s), object(s), etc.) representing data associated with workspace(s) with which the user (e.g., the user's account) is associated. In at least one example, the second region 146 may represent a sidebar of the user interface 143 .

[0048] In at least one example, user interface 143 may include a third region 148 or pane that may be associated with a data feed (or “feed”) indicating messages posted and / or actions taken for one or more communication channels and / or other virtual spaces for facilitating communication as described herein (e.g., virtual spaces associated with direct message communication(s), virtual spaces associated with event(s) and / or action(s), etc.). In at least one example, data associated with third region 148 may be associated with the same or different workspaces. That is, in some examples, third region 148 may present data associated with the same or different workspaces via a unified feed. In some examples, data may be organized and / or sortable by workspace, time (e.g., when the associated data was posted or when the associated action was otherwise performed), type of action, communication channel, user, etc. In some examples, such data may be associated with an indication of which users (e.g., members of a communication channel) posted the message and / or performed the action. In examples where third region 148 presents data associated with multiple workspaces, at least some of the data may be associated with an indication of which workspace the data is associated with. In some examples, third region 148 may be resized or popped out as a separate window.

[0049] In at least one example, operating system 140 may manage processor(s) 130, computer-readable media 132, hardware, software, etc. of server(s) 102.

[0050] The communication interface(s) 134 may include one or more interfaces and hardware components to enable communication with various other devices (e.g., user computing device 104), for example, over the network(s) 106 or directly. In some examples, the communication interface(s) 134 may facilitate communication via WebSockets, an API (e.g., using API calls), HTTPs, etc.

[0051] The user computing device 104 may further include various input / output devices 136 (e.g., I / O devices). Such I / O devices 136 may include a display, various user interface controls (e.g., buttons, joystick, keyboard, mouse, touch screen, etc.), audio speakers, connection ports, etc.

[0052] Although the techniques described herein are described as being performed by the messaging component 116, the audio / video component 118, the summarization engine 120, and the application 138, the techniques described herein may be performed by any other component or combination of components that may be associated with the server(s) 102, the user computing device 104, or a combination thereof. A user interface for a group-based communication system

[0053] 2A illustrates a user interface 200 of a group-based communication system, which is useful for illustrating the operation of various examples discussed herein. The group-based communication system may include communication data such as messages, queries, files, mentions, users or user profiles, interactions, tickets, channels, applications integrated into one or more channels, conversations, workspaces, or other data generated by or shared among users of the group-based communication system. In some instances, the communication data may include data associated with a user, such as a user identifier, channels the user is authorized to access, groups the user is associated with, permissions, and other user-specific information.

[0054] User interface 200 includes multiple objects such as panes, text entry fields, buttons, messages, or other user interface components viewable by a user of the group-based communication system. As shown, user interface 200 includes a title bar 202, a workspace pane 204, a navigation pane 206, a channel 208, a document 210 (e.g., a collaborative document), a direct message 212, an application 214, a synchronized multimedia collaboration session pane 216, and a channel pane 218.

[0055] By way of example and not limitation, when a user opens the user interface 200, the user can select a workspace via the workspace pane 204. A particular workspace may be associated with data that is unique to the workspace and accessible via permissions associated with the workspace. Different sections of the navigation pane 206 may present different data and / or options to the user. Different graphical indicators may be associated with virtual spaces (e.g., channels) to summarize attributes of the channel (e.g., whether the channel is public, private, shared across organizations, locked, etc.). When a user selects a channel, the channel pane 218 may be presented. In some examples, the channel pane 218 may include a header, pinned items (e.g., documents or other virtual spaces), an “about” document providing an overview of the channel, etc. In some cases, channel members can search within the channel, access content associated with the channel, add other members, post content, etc. In some examples, depending on the permissions associated with the channel, users who are not members of the channel may be limited in their ability to interact with (or even view or otherwise access) the channel. As a user navigates within a channel, the user can view messages 222, react to messages (e.g., by selecting the reaction button 224), reply to a thread, start a thread, etc. Additionally, the channel pane 218 can include a composition pane 228 for composing message(s) and / or other data to associate with the channel. In some examples, the user interface 200 can include a thread pane 230 that provides an additional level of detail for the messages 222. In some examples, different panes can be resized, panes can pop out into separate windows, and / or separate windows can be merged into multiple panes of the user interface 200.In some examples, users may communicate with other users via a multimedia collaboration session pane 216, which may provide synchronous or asynchronous audio and / or video capabilities for communication. Of course, these are illustrative examples, and additional examples of the foregoing features are provided throughout this disclosure.

[0056] In some examples, title bar 202 includes a search bar 220. Search bar 220 may allow a user to search for content located in the current workspace of the group-based communication system, such as files, messages, channels, members, commands, and functions. A user may narrow a search by attributes such as content type, content creator, and by the user associated with the content. A user may optionally search within a specific workspace, channel, direct message conversation, or document. In some examples, title bar 202 includes navigation commands that allow a user to move back and forth between different panes and view a history of accessed content. In some examples, title bar 202 may include additional resources, such as links to help documentation and user configuration settings.

[0057] In some examples, the group-based communication system can include multiple separate workspaces, each associated with a different group of users and channels. Each workspace can be associated with a group identifier, and one or more user identifiers can be mapped to or otherwise associated with the group identifier. A user corresponding to such a user identifier can be referred to as a member of the group. In some examples, the user interface 200 includes a workspace pane 204 for navigating between, adding to, or removing from various workspaces within the group-based communication system. For example, a user can be part of a workspace for Acme, where the user is an employee of or otherwise affiliated with Acme. The user can also be a member of a local volunteer organization that also collaborates using the group-based communication system. To navigate between the two groups, the user can use the workspace pane 204 to change from the Acme workspace to the volunteer organization workspace. A workspace can include one or more channels specific to that workspace and / or one or more channels shared between one or more workspaces. For example, Acme Corporation can have a workspace for an Acme project, such as Project Zen, a workspace for social discussions, and an additional workspace for general company operations. In some examples, an organization such as a particular company may have multiple workspaces, and a user may be associated with one or more workspaces belonging to that organization. In yet other examples, a particular workspace may be associated with one or more organizations or other entities associated with the group-based communication system.

[0058] In some examples, the navigation pane 206 allows a user to navigate between virtual spaces, such as pages, channels 208, documents 210 (such as those illustrated in FIG. 2D ), applications 214, and direct messages 212, within a group-based communication system. For example, the navigation pane 206 may include an indicator representing a virtual space that can aggregate data associated with multiple virtual spaces of which the user is a member. In at least one example, each virtual space may be associated with an indicator in the navigation pane 206. In some examples, the indicator may be associated with an actuation mechanism (e.g., an affordance, also referred to as a graphical element) that, when actuated, causes the user interface 200 to present data associated with the corresponding virtual space. In at least one example, a virtual space may be associated with all unread data associated with each of the workspaces with which the user is associated. That is, in some examples, when a user requests access to a virtual space associated with “unreads,” all data that has not been read (e.g., not viewed) by the user may be presented, for example, in a feed. In such examples, different types of events and / or actions that may be associated with different virtual spaces may be presented via the same feed. In some examples, such data may be organized and / or sortable by associated virtual space (e.g., the virtual space through which the communication was transmitted), time, type of action, user, and / or the like. In some examples, such data may be associated with an indication of which user (e.g., a member of the associated virtual space) posted the message and / or performed the action.

[0059] In some examples, virtual spaces may be associated with the same type of events and / or actions. For example, "threads" may be associated with messages, files, etc. posted to a thread in response to a message posted to the virtual space, and "mentions and reactions" may be associated with messages or threads in which the user mentions (e.g., via tags) or another user reacts (e.g., via emojis, reactions, etc.) to a message or thread posted by the user. That is, in some examples, the same type of events and / or actions that may be associated with different virtual spaces may be presented via the same feed. Similar to the "unreads" virtual space, data associated with such virtual spaces may be organized and / or sortable by virtual space, time, type of action, user, and / or the like.

[0060] In some examples, the virtual space may be associated with facilitating communication between the user and other users of the communication platform. For example, "connect" may be associated with enabling the user to generate invitations to communicate with one or more other users. In at least one example, in response to receiving an indication of selection of the "connect" indicator, the communication platform may cause a connection interface to be presented.

[0061] In some examples, the virtual space may be associated with one or more boards or collaborative documents with which the user is associated. In at least one example, the document may include a collaborative document configured to be accessed and / or edited by two or more users with appropriate permissions (e.g., view permissions, edit permissions, etc.). In at least one example, when a user requests access to a virtual space associated with one or more documents with which the user is associated, the one or more documents may be presented via user interface 200. In at least one example, the documents may be associated with an individual (e.g., a user's private document), a group of users (e.g., a collaborative document), and / or one or more communication channels (e.g., members of a communication channel who have been granted access permissions to the document), as described herein, thereby enabling, for example, users of the communication platform to create, interact with, and / or view data associated with such documents. In some examples, the collaborative document may be a virtual space, board, canvas, page, etc. for collaborative communication and / or data organization within the communication platform. In at least one example, a collaborative document may support editable text and / or objects that can be ordered, added, deleted, modified, and / or the like. In some examples, a collaborative document may be associated with permissions that define which users of a communication platform can view and / or edit the document. In some examples, a collaborative document may be associated with a communication channel, where members of the communication channel can view and / or edit the document. In some examples, a collaborative document may be shareable such that data associated with the document is accessible and / or interactable by members of multiple communication channels, workspaces, organizations, and / or the like.

[0062] In some examples, a virtual space may be associated with a group (e.g., organization, team, etc.) headquarters (e.g., management or command center). In at least one example, a group headquarters may include a virtual or digital headquarters for management or command functions associated with a group of users. For example, an "HQ" may be associated with an interface including a list of indicators associated with virtual spaces configured to allow associated members to communicate. In at least one example, a user may associate one or more virtual spaces with an "HQ" virtual space, such as via a drag-and-drop operation. That is, a user may determine relevant virtual space(s) to associate with the virtual or digital headquarters, such as associating virtual space(s) important to the user.

[0063] In some examples, the virtual space may be associated with one or more boards or collaborative documents with which the user is associated. In at least one example, the document may include a collaborative document configured to be accessed and / or edited by two or more users with appropriate permissions (e.g., view permissions, edit permissions, etc.). In at least one example, when a user requests access to a virtual space associated with one or more documents with which the user is associated, the one or more documents may be presented via user interface 200. In at least one example, the documents may be associated with an individual (e.g., a user's private document), a group of users (e.g., a collaborative document), and / or one or more communication channels (e.g., members of a communication channel who have been granted access permissions to the document), as described herein, thereby enabling, for example, users of the communication platform to create, interact with, and / or view data associated with such documents. In some examples, the collaborative document may be a virtual space, board, canvas, page, etc. for collaborative communication and / or data organization within the communication platform. In at least one example, a collaborative document may support editable text and / or objects that can be ordered, added, deleted, modified, and / or the like. In some examples, a collaborative document may be associated with permissions that define which users of a communication platform can view and / or edit the document. In some examples, a collaborative document may be associated with a communication channel, where members of the communication channel can view and / or edit the document. In some examples, a collaborative document may be shareable such that data associated with the document is accessible and / or interactable with by members of multiple communication channels, workspaces, organizations, and / or the like.

[0064] Additionally or alternatively, in some examples, a virtual space may be associated with one or more canvases with which users are associated. In at least one example, a canvas may include a flexible canvas for curating, organizing, and sharing collections of information among users. That is, a canvas may be configured to be accessed and / or modified by two or more users with appropriate permissions. In at least one example, a canvas may be configured to enable the sharing of text, images, videos, GIFs, drawings (e.g., user-generated drawings via a canvas interface), game content (e.g., users manipulating game controls synchronously or asynchronously), and / or the like. In at least one example, modifications to a canvas may include adding, deleting, and / or modifying previously shared (e.g., transmitted, presented) data. In some examples, content associated with a canvas may be shareable via another virtual space, thereby making the data associated with the canvas accessible and / or interactable to members of the virtual space.

[0065] The navigation pane 206 may further include indicators representing communication channels (e.g., channels 208). In some examples, the communication channels may include public channels, private channels, shared channels (e.g., between groups or organizations), single-workspace channels, cross-workspace channels, combinations of the above, etc. In some examples, the represented communication channels may be associated with a single workspace. In some examples, the represented communication channels may be associated with different workspaces (e.g., cross-workspace). In at least one example, when a communication channel is cross-workspace (e.g., associated with different workspaces), a user may be associated with both workspaces or, alternatively, with only one of the workspaces. In some examples, the represented communication channels may be associated with a combination of communication channels associated with a single workspace and communication channels associated with different workspaces.

[0066] In some examples, navigation pane 206 may show some or all of the communication channels that the user has permission to access (e.g., as determined by permission data). In such examples, the communication channels are arranged alphabetically, based on most recent interaction, based on frequency of interaction, based on communication channel type (e.g., public, private, shared, cross-workspace, etc.), based on workspace, in user-specified sections, or similar. In some examples, navigation pane 206 may show some or all of the communication channels of which the user is a member, and the user may interact with user interface 200 to browse or view other communication channels of which the user is not a member but which are not currently displayed in navigation pane 206. In some examples, different types of communication channels (e.g., public, private, shared, cross-workspace, etc.) may be in different sections of navigation pane 206 or may have their own sub-areas or sub-panes within user interface 200. In some examples, communication channels associated with different workspaces may be in different sections of the navigation pane 206, or the communication channels may have their own area or pane within the user interface 200.

[0067] In some examples, the indicator may be associated with a graphical element that visually distinguishes the type of communication channel. For example, project_zen is associated with a lock graphical element. As a non-limiting example, for purposes of this discussion, the lock graphical element may indicate that the associated communication channel, project_zen, is private, with restricted access, while another communication channel, general, is public, with access available to any member of the organization with which the user is associated. In some examples, additional or alternative graphical elements may be used to distinguish between shared communication channels, communication channels associated with different workspaces, communication channels in which the user is currently a member or not, and / or the like.

[0068] In at least one example, the navigation pane 206 may include indicators representing communications with an individual user or multiple designated users (e.g., on behalf of all or a subset of the members of an organization). Such communications may be referred to as “direct messages.” The navigation pane 206 may include indicators representing virtual spaces associated with private messages between one or more users.

[0069] Direct messages 212 may be communications between a first user and a second user, or may be multi-person direct messages between a first user and two or more second users. Navigation pane 206 may be sorted and organized into hierarchies or sections according to user preferences. In some examples, all of the channels a user is authorized to access may be displayed in navigation pane 206. In other examples, a user may choose to hide certain channels or collapse sections containing certain channels. Items in navigation pane 206 may indicate when new messages or updates have been received or are currently unread, for example, by bolding text associated with channels in which unread messages are located or by adding an icon or badge (e.g., with a count of unread messages) to the channel name. In some examples, the group-based communication system may additionally or alternatively store permission data associated with the permissions of individual users of the group-based communication system, indicating which channels the user may view or participate in. Permissions may indicate, for example, which users can communicate directly with other users, which channels the user is permitted to access, restrictions on individual channels, which workspaces the user is permitted to access, and restrictions on individual workspaces.

[0070] Additionally or alternatively, navigation pane 206 may include subsections that are personalized subsections associated with teams of which the user is a member. That is, the "team" subsection may include affordance(s) of one or more virtual spaces associated with the team, such as communication channels, collaborative documents, direct messaging instances, audio or video synchronous or asynchronous conferences, and / or the like. In at least one example, a user may associate a selected virtual space with the team subsection by, for example, dragging and dropping, pinning, or otherwise associating the selected virtual space with the team subsection. Channels in a Group-Based Communication System

[0071] In some examples, the group-based communication system is a channel-based messaging platform, as shown in Figure 2A. Within a group-based communication system, communications can be organized into channels, each dedicated to a particular topic and set of users. A channel is generally a virtual space related to a particular topic that contains messages and files posted by members of the channel.

[0072] For purposes of this description, a “message,” as described herein, may refer to any electronically generated digital object provided by a user using a user computing device 104 and configured for display within a communication channel and / or other virtual space to facilitate communication (e.g., a virtual space associated with direct message communication(s)), such as a virtual space associated with direct message communication(s). A message may include any text, image, video, audio, or combination thereof provided by a user (using a user computing device). For example, a user may provide a message that includes not only text but also images and video within the message as message content. In such an example, the text, image, and video constitute the message. Each message sent or posted to a communication channel of a communication platform may include metadata including a sending user identifier, a message identifier, message content, a group identifier, a communication channel identifier, etc. In at least one example, each of the aforementioned identifiers may include American Standard Code for Information Interchange (ASCII) text, a pointer, a memory address, etc.

[0073] Channel discussions can persist for days, months, or years, providing a historical log of user activity. Members of a particular channel can post messages within that channel that are visible to other members of that channel, along with other messages in that channel. Users may select a channel to browse to see only messages related to that channel's topic, without seeing messages posted in other channels on different topics. For example, a software development company may have a different channel for each software product under development, allowing developers working on each specific project to generally talk about a single topic (e.g., the project) without the noise of unrelated topics. Because channels are generally persistent and targeted to a specific topic or group, users can quickly and easily refer to previous communications for reference. In some examples, the channel pane 218 may display information related to a channel a user selects in the navigation pane 206. For example, a user may select the project_zen channel to discuss Project Zen's ongoing software development efforts. In some examples, the channel pane 218 may include a header containing information about the channel, such as the channel name, a list of users in the channel, and other channel controls. Users may pin items to the header for later access and add bookmarks to the header. In some examples, links to collaborative documents may be included in the header. In a further example, each channel may have a corresponding virtual space that includes channel-related information such as a channel summary, tasks, bookmarks, pinned documents, and other channel-related links that may be editable by members of the channel.

[0074] A communication channel or other virtual space may be associated with data and / or content other than messages or data and / or content associated with messages. For example, non-limiting examples of additional data presentable via the channel pane 218 of the user interface 200 may include collaborative documents (e.g., documents that are collaboratively editable in real time or near real time), audio and / or video data associated with a conversation, members added to and / or removed from a communication channel, file(s) uploaded to and / or removed from a communication channel (e.g., attachment(s)), application(s) added to and / or removed from a communication channel, post(s) added to and / or removed from a communication channel (data that are collaboratively editable in near real time by one or more members of a communication channel), descriptions added to, modified from, and / or removed from a communication channel, modifications to properties of a communication channel, etc.

[0075] The channel pane 218 may include messages, such as message 222, which are content posted to a channel by a user. A user may post text, images, video, audio, or any other file as message 222. In some examples, specific identifiers (among messages or otherwise) may be indicated by preceding them with a predetermined character. For example, a channel may be preceded by a "#" character (such as #project_zen), and a username may be preceded by an "@" character (such as @@J_Smith or @User_A). A message, such as message 222, may include an indication of which user posted the message and when the message was posted. In some examples, a user may react to a message by selecting a reaction button 224. The reaction button 224 allows a user to select an icon, such as a thumbs-up (sometimes referred to in this context as a reaction), to associate with the message. A user may reply to another user's message, such as message 222, with a new message. In some examples, such conversations within a channel may be further divided into threads. Threads may be used to aggregate messages related to a particular conversation, making it easier to follow and reply to the conversation without cluttering the main channel with discussion. A thread reply preview 226 is displayed below the message that starts the thread. The thread reply preview 226 may show information related to the thread, such as the number of replies and the members who replied. The thread replies may be displayed in a thread pane 230, which may be separate from the channel pane 218, and may be viewed by other members of the channel by selecting the thread reply preview 226 in the channel pane 218.

[0076] In some examples, one or both of the channel pane 218 and the thread pane 230 may include a composition pane 228. In some examples, the composition pane 228 allows users to compose and send messages 222 to members of the channel or to members of the channel who are following the thread (if the message is sent in a thread). The composition pane 228 may have text editing features such as bold, strikethrough, and italics, and / or may allow users to format their messages or attach files such as collaborative documents, images, videos, or any other files for sharing with other members of the channel. In some examples, the composition pane 228 may enable additional formatting options, such as numbered or bulleted lists, either through the user interface or an API. The composition pane 228 may act as a workflow trigger to initiate a workflow associated with the channel or message. In further examples, links or documents sent via the composition pane 228 may include unfill instructions regarding how the content should be displayed. Transcription and / or summary of multimedia collaboration sessions

[0077] Conventional transcription algorithms simply convert audio signal data and apply a language model to output written language from the spoken language. However, conventional transcription models are inadequate in considering the robust contextual data inherent in group-based communication channels. Thus, conventional transcription models typically provide only text records lacking contextually actionable information. In contrast, ML model(s) as disclosed in various examples herein (e.g., ML model(s) 142) can be trained to ingest at least data including a multimedia collaboration session and output a summary document characterized by the most contextually relevant and / or actionable information disseminated during the multimedia collaboration session. That is, the ML model(s) of various examples can be trained to leverage the channel context associated with the virtual space in which the multimedia collaboration session is initiated. Such non-verbal context includes, but is not limited to, the identity of and relationships between users of the group-based communication platform, references to users (i.e., @-mentions), the submission of documents or other file types in conjunction with the discussion, etc. This channel context provides powerful data for training various example ML model(s) to extrapolate meeting purpose, goals, intent, targets, etc. Accordingly, various example ML model(s) may be advantageously trained to at least learn one or more relationships between instances of environmental data that constitute a multimedia collaboration session and transcripts that represent the synchronous multimedia collaboration session in various examples disclosed herein. In at least some examples, the machine learning model(s) may be trained to learn one or more relationships between instances of environmental data and transcripts that define tasks and deadlines, or to otherwise correlate actors with actions to be taken and timeframes within which the actions should be taken.

[0078] 2B illustrates a multimedia collaboration session (e.g., a synchronous multimedia collaboration session, also referred to herein as an audio and / or video conference) triggered from a channel, as shown in synchronized multimedia collaboration session pane 216. Synchronous multimedia collaboration sessions can provide ambient and ad-hoc multimedia collaboration in a group-based communication system. Users of the group-based communication system can quickly and easily join or leave these synchronized multimedia collaboration sessions at any time without interrupting the synchronized multimedia collaboration sessions for other users. In some examples, synchronized multimedia collaboration sessions can be based on a specific topic, a specific channel, a specific direct message or multiparty direct message, or a set of users to facilitate task completion according to a deadline.

[0079] The synchronized multimedia collaboration session pane 216 may be associated with a synchronized multimedia collaboration session conducted for multiple users in a channel, users in a multi-party direct message conversation, or users in a direct message conversation. Thus, a synchronized multimedia collaboration session may be initiated for a particular channel, a multi-party direct message conversation, or a direct message conversation by one or more members of that channel or conversation. A user may initiate a synchronized multimedia collaboration session in a channel as a means of communicating with other members of that channel who are currently online. For example, a user may need to make an urgent decision and desire immediate verbal feedback from other members of the channel. As another example, a synchronized multimedia collaboration session may be initiated with one or more other users of a group-based communication system through direct messaging. In some examples, the audience for a synchronized multimedia collaboration session may be determined based on the context in which the synchronized multimedia collaboration session is initiated. For example, initiating a synchronized multimedia collaboration session in a channel may automatically invite the entire channel to participate. As another example, initiating a synchronized multimedia collaboration session allows a user to initiate an instant audio and / or video conversation with other members of the channel without having to schedule or initiate a communication session through a third-party interface. In some examples, users may be directly invited to participate in a synchronous multimedia collaboration session via a message or notification.

[0080] In some examples, the content of a synchronized multimedia collaboration session may be automatically persisted within a channel associated with the synchronized multimedia collaboration session. Members of a particular synchronized multimedia collaboration session may post messages within a thread associated with the synchronized multimedia collaboration session that are visible to other members of the synchronized multimedia collaboration session, along with other messages within the thread.

[0081] The multimedia in a synchronous multimedia collaboration session may include collaboration tools such as any or all of audio, video, screen sharing, collaborative document editing, whiteboarding, co-programming, or any other form of media. The synchronous multimedia collaboration session may also allow users to share their screens with other members of the synchronous multimedia collaboration session. In some examples, members of the synchronous multimedia collaboration session may mark up, comment, draw, or otherwise annotate the shared screen. In further examples, such annotations may be saved and persist even after the synchronous multimedia collaboration session has ended. Canvases may be created directly from the synchronous multimedia collaboration session to further enhance collaboration between users.

[0082] In some examples, a user may start a synchronized multimedia collaboration session via a toggle in the synchronized multimedia collaboration session pane 216 shown in FIG. 2B . When a synchronized multimedia collaboration session is started, the synchronized multimedia collaboration session pane 216 may expand to provide information about the synchronized multimedia collaboration session, such as how many members are present, which user is currently speaking, which user is sharing their screen, and / or a screen share preview 232. In some examples, users in the synchronized multimedia collaboration session may be displayed with an icon indicating they are participating in the synchronized multimedia collaboration session. In further examples, the expanded view of the participants may indicate which users are active and which users are not active in the synchronized multimedia collaboration session. The screen share preview 232 may show a user's screen or a desktop view of a user sharing a particular application or presentation. Changes to a user's screen, such as a user advancing to the next slide in a presentation, are automatically shown in the screen share preview 232. In some examples, the screen share preview 232 may be activated to expand the screen share preview 232 so that it appears as its own pane within the group-based communication system. In some examples, the screen share preview 232 can be activated to pop out into a new window or application that is separate and distinct from the group-based communication system. In some examples, the synchronized multimedia collaboration session pane 216 can include tools for the synchronized multimedia collaboration session that allow a user to mute the user's microphone or invite other users. In some examples, the synchronized multimedia collaboration session pane 216 can include a screen share button 234 that can allow a user to share the user's screen with other members of the synchronized multimedia collaboration session pane 216.In some examples, the screen share button 234 may provide a user with additional controls during a screen share. For example, a user sharing their screen may be provided with additional screen share controls to specify which screen to share, to annotate the shared screen, or to save the shared screen.

[0083] In some cases, the synchronized multimedia collaboration session pane 216 may persist within the navigation pane 206 regardless of the state of the group-based communication system. In some examples, when no synchronized multimedia collaboration session is active and / or depending on which item is selected from the navigation pane 206, the synchronized multimedia collaboration session pane 216 may be hidden or removed from presentation via the user interface 200. In some instances, when the synchronized multimedia collaboration session pane 216 is active, the synchronized multimedia collaboration session pane 216 may be associated with the currently selected channel, direct message, or multi-party direct message, thereby initiating a synchronized multimedia collaboration session associated with the currently selected channel, direct message, or multi-party direct message.

[0084] The list of synchronized multimedia collaboration sessions may include one or more active synchronized multimedia collaboration sessions selected for recommendation. For example, the synchronized multimedia collaboration session may be selected from a plurality of currently active synchronized multimedia collaboration sessions. Furthermore, the synchronized multimedia collaboration sessions may be selected based in part on user interaction with the session or some association of the instant user with the session or users involved in the session. For example, recommended synchronized multimedia collaboration sessions may be displayed in part based on the instant user being invited to the respective synchronized multimedia collaboration session or having previously collaborated with the user in the recommended synchronized multimedia collaboration session. In some examples, the list of synchronized multimedia collaboration sessions further includes additional information for each synchronized multimedia collaboration session, such as an indication of the participating users or number of participating users, the topic of the synchronized multimedia collaboration session, and / or an indication of associated group-based communication channels, multi-party direct message conversations, or direct message conversations.

[0085] In some examples, the list of recommended active users may include multiple group-based communication system users that are recommended based on at least one of user activity, user interactions, or other user information. For example, the list of recommended active users may be selected based on the user's active status within the group-based communication system, past, recent, or frequent user interactions with the instant user (e.g., communications within a group-based communication channel), or similarity between the recommended user and the instant user (e.g., determining that the recommended user shares a shared channel membership with the instant user). In some examples, machine learning techniques such as cluster analysis may be used to determine the recommended users. The list of recommended active users may include status user information for each recommended user, such as whether the recommended user is active, in a meeting, idle, in a synchronous multimedia collaboration session, or offline. In some examples, the list of recommended active users further includes multiple actuatable buttons corresponding to some or all of the recommended users (e.g., recommended users having a status indicating availability), which, when selected, may be configured to initiate at least one of a text-based communication session (e.g., a direct message conversation) or a synchronous multimedia collaboration session.

[0086] In some examples, one or more recommended asynchronous multimedia collaboration sessions or conferences may be displayed in the asynchronous conference section. In contrast to synchronous multimedia collaboration sessions (described above), asynchronous multimedia collaboration sessions allow each participant to collaborate at their convenience. This collaboration participation is then recorded for later consumption by other participants, which may generate additional multimedia replies. In some examples, replies are aggregated into a multimedia thread (e.g., a video thread) corresponding to the asynchronous multimedia collaboration session. For example, an asynchronous multimedia collaboration session may be used for an asynchronous conference where a topic is posted with a message at the start of the conference thread and conference participants can reply by posting messages or video replies. The resulting thread includes any documents, videos, or other files related to the asynchronous conference. In some examples, a preview of a subset of the video replies may be shown in the asynchronous collaboration session or thread. This may allow, for example, a user to jump to a relevant segment of the asynchronous multimedia collaboration session or pick up where the user previously left off.

[0087] In other examples, the multimedia collaboration session may be a short, transient session where the data does not persist. In yet some other examples, the multimedia collaboration session may be recorded, transcribed, and / or summarized for later review and interaction.

[0088] Additional discussion regarding the determination of summary data is provided in connection with FIGS. 4 and 5, and throughout this disclosure. Connecting within a group-based communication system

[0089] 2C shows user interface 200 displaying connect pane 252. Connect pane 252 may provide tools and resources for users to connect across different organizations, each of which may have its own (usually private) instance of the group-based communication system or may not already belong to the group-based communication system. For example, a first software company may want to collaborate to jointly develop a new software application and form a joint venture with a second software company. Connect pane 252 allows users to determine which other users and organizations are already in the group-based communication system and invite those users and organizations currently outside the group-based communication system to join.

[0090] The connect pane 252 may include a connect search bar 254, recent contacts 256, connections 258, a create channel button 260, and / or a start direct message button 262. In some examples, the connect search bar 254 may allow a user to search for users within the group-based communication system. In some examples, only users from organizations connected to the user's organization are shown in the search results. In other examples, users from any organization that uses the group-based communication system may be displayed. In still other examples, users from organizations that are not yet using group-based communication may also be displayed, allowing the searching user to invite them to join the group-based communication system. In some examples, users may be searched for via the user's group-based communication system username or the user's email address. In some examples, email addresses may be suggested or auto-completed based on external data sources, such as email directories or the searching user's contact list.

[0091] In some examples, external organizations and individual users may be shown in response to a user search. Because search results may include organizations that have not yet joined the group-based communication system, external organizations may be matched based on organization name or Internet domain (similar to searching and matching specific users described above). External organizations may be ranked in part based on how many users from the user's organization are connected to users of the external organization. In response to selecting an external organization in the search results, the searching user may invite the external organization to connect via the group-based communication system.

[0092] In some examples, recent contacts 256 may display users with whom the instant user has recently interacted. Recent contacts 256 may display the user's name, company, and / or status indication. Recent contacts 256 may be ordered based on which contacts the instant user interacts with most frequently or based on which contacts the instant user most recently interacted with. In some examples, each recent contact in recent contacts 256 may be an actuatable control that allows the instant user to quickly start a direct message conversation with the recent contact, invite them to a channel, or take any other appropriate user action with that recent contact.

[0093] In some examples, connections 258 may display a list of companies (e.g., organizations) with which the user has interacted. For each company, the company's name may be displayed along with the company's logo and an indication of how many times the user has interacted with that company, e.g., the number of conversations. In some examples, each connection in connections 258 may be an actuatable control that enables an instant user to quickly invite an external organization to a shared channel, view recent connections with that external organization, or take any other appropriate organizational action with that connection.

[0094] In some examples, the create channel button 260 allows a user to create a new shared channel between two different organizations. By selecting the create channel button 260, the user can further name the new connect channel and enter a description of the connect channel. In some examples, the user may select one or more external organizations or one or more external users to add to the shared channel. In other examples, the user may add external organizations or external users to the shared channel after the shared channel is created. In some examples, the user may select whether to make the connect channel private (e.g., accessible only by invitation from current members of the private channel).

[0095] In some examples, the direct message initiation button 262 allows a user to quickly initiate a direct message (or a multi-party direct message) with an external user of an external organization. In some examples, an external user identifier at the external organization may be provided by the instant user as the external user's group-based communication system username or as the external user's email address. In some examples, analysis of the email domain of the external user's email address may influence messages between the user and the external user. For example, the external user's identifier may indicate that the user's organization and the external user's organization are already connected (e.g., based on the email address domain). In some such examples, the email address may be converted to a group-based communication system username. Alternatively, the external user's identifier may indicate that the external user's organization belongs to the group-based communication system but is not connected to the instant user's organization. In some such examples, an invitation to connect to the instant user's organization may be generated in response. As another alternative, the external user may not be a member of the group-based communication system, and in response, an invitation to join the group-based communication system as a guest or a member may be generated. joint document

[0096] FIG. 2D shows the user interface 200 displaying a collaboration document pane 264. A collaborative document may be any file type, such as a PDF, video, audio, or word processing document, but is not limited to a word processing document or spreadsheet. A collaborative document may be modified and edited by two or more users. A collaborative document may also be associated with different user permissions, such that a user may be selectively authorized to view, edit, or comment on a collaborative document (or a section of a collaborative document) based on the user's permissions for the document (or a section of a document, as described below). Thus, users within a set of users with access to a document may have different permissions to view, edit, comment, or otherwise interface with the collaborative document. In some examples, permissions may be automatically determined and / or assigned based on how the document(s) are created and / or shared. In some examples, permissions may be determined manually. A collaborative document may allow users to create and modify documents simultaneously or asynchronously. Collaborative documents may be integrated with group-based communication systems and may be used to initiate workflows and store workflow results, as discussed further below with respect to Figures 3A and 3B.

[0097] In some examples, the user interface 200 may include one or more collaborative documents (or one or more links to such collaborative documents). A collaborative document (also called a document or canvas) may include a flexible workspace for curating, organizing, and sharing a collection of information among users. Such a document may be associated with a synchronous multimedia collaboration session, an asynchronous multimedia collaboration session, a channel, a multi-party direct message conversation, and / or a direct message conversation. A shared canvas may be configured to be accessed and / or modified by two or more users with appropriate permissions. Alternatively or additionally, a user may have one or more private documents that are not associated with any other user.

[0098] Additionally, such documents can be @mentioned so that a particular document can be referenced within a channel (or other virtual space or document) and / or other users can be @mentioned within such documents. For example, @mentioning a user within a document can provide instructions to the user and / or provide the user with access to the document. In some examples, tasks can be assigned to a user via @mention, and such task(s) can be populated into a pane or sidebar associated with the user.

[0099] In some examples, a channel and a collaborative document 268 may be associated so that when a comment is posted to the channel, the comment is populated into the document 268, and vice versa.

[0100] In some examples, when a first user interacts with a collaborative document, the communication platform may identify a second user account associated with the collaborative document and present an affordance (e.g., a graphical element) in a sidebar (e.g., navigation pane 206) indicating the interaction. Additionally, the second user may select an affordance and / or notification associated with or representative of the interaction to access the collaborative document to efficiently access the document and view updates thereto.

[0101] In some examples, when one or more users interact with a collaborative document, indications (e.g., icons or other user interface elements) may be presented via a user interface with the collaborative document to represent such interactions. For example, if a first instance of the document is currently open on a first user computing device of a first user and a second instance of the document is currently open on a second user computing device of a second user, one or more presence indicators may be presented on the respective user interfaces to indicate various interactions with the document and by which users. In some examples, the presence indicators may have attributes (e.g., appearance attributes) that indicate information about each user, such as, but not limited to, permission level (e.g., edit permission, read-only access, etc.), virtual space membership (e.g., whether a member belongs to a virtual space associated with the document), and how the user is interacting with the document (e.g., currently editing, viewing, open but not active, etc.).

[0102] In some examples, a preview of the collaborative document can be provided. In some examples, the preview can include a summary of the collaborative document and / or a dynamic preview that displays various content (e.g., changing text, images, etc.) to enable a user to quickly understand the context of the document. In some examples, the preview can be based on user profile data associated with the user viewing the preview (e.g., permissions associated with the user, content viewed, edited, created, etc. by the user), etc.

[0103] In some examples, collaborative documents may be created independently of or in association with virtual spaces and / or channels, and may be posted within channels and edited or interacted with as discussed herein, with various affordances or notifications indicating the presence of users associated with the document and / or various interactions.

[0104] In some examples, a machine learning model can be used to determine a summary of the channel's content and create a collaborative document that includes the summary for posting to the channel. In some examples, the communication platform may identify users in the virtual space, actions associated with the users, and other contributions to the conversation to generate the summary document. As such, the communication platform can enable users to create documents (e.g., collaborative documents) to summarize content and events that occurred in the virtual space.

[0105] In some examples, a document may be configured to allow sharing of content, including (but not limited to) text, images, videos, GIFs, drawings (e.g., user-generated drawings via a drawing interface), or game content. In some examples, a user accessing the canvas may add new content or delete (or modify) previously added content. In some examples, appropriate permission may be required for a user to add content or delete or modify content added by a different user. Thus, for example, some users may be able to access part or all of a document only in a view-only mode, while other users may be able to access part or all of the document in an edit mode that allows those users to add or modify its content. In some examples, a document may be shared via a message in a channel, a multi-party direct message, or a direct message, such that data associated with the document is accessible and / or interactable with by members of the channel or recipients of the multi-party direct message or direct message.

[0106] In some examples, collaboration document pane 264 may include a collaborative document toolbar 266 and a collaborative document 268. In some examples, collaborative document toolbar 266 may provide the ability to edit or format a post, as discussed herein.

[0107] In some examples, a collaborative document may include free-form unstructured sections and workflow-related structured sections. In some examples, an unstructured section may include an area of ​​a document where a user can freely modify the collaborative document without constraints. For example, a user may freely type text to describe the purpose of the document. In some examples, a user may add a workflow or structured workflow section by typing (or otherwise mentioning) the name of a workflow. In a further example, typing an "at" sign (@), a previously selected symbol, or a predetermined special character or symbol may provide the user with a list of workflows from which the user can select to add to the document. For example, a user may type "!Marketing Approval" to indicate that a marketing team member needs to sign off on the proposal and initiate a workflow that results in the marketing team member approving the proposal. Placing an exclamation point in front of the "Marketing Approval" group name initiates a request for a specific action, in this case, routing the proposal for approval. In some examples, a structured section may include a text entry, a selection menu, a table, a checkbox, a task, a calendar event, or any other document section. In a further example, a structured section may include a text entry space that is part of a workflow. For example, a user may enter text into a text entry field detailing the reason for approval and then select a submit button that advances the workflow to the next step in the workflow. In some examples, a user may be able to add, edit, or remove structured sections of the document that make up the workflow components.

[0108] In an example, sections of a collaborative document may have individual permissions associated with them. For example, a collaborative document having sections with individual permissions may provide a first user with permission to view, edit, or comment on the first section, but a second user does not have permission to view, edit, or comment on the first section. Alternatively, a first user may have permission to view the first section of the collaborative document, and a second user may have permission to view and edit the first section of the collaborative document. Permissions associated with a particular section of a document may be assigned by the first user through various methods, including manual selection of the particular section of the document by the first user or another user with permission to assign permissions, typing or selecting an “assign” indicator such as an “@” symbol, or selecting the section by its name. In a further example, permissions may be assigned to multiple collaborative documents in a single instance through these methods. For example, multiple collaborative documents each have a section titled “Group Information,” and a first user with permission to assign permissions desires that an entire group of users access the information in the “Group Information” section of the multiple collaborative documents. In an example, a first user can select multiple collaborative documents and a “Group Information” section to effectuate permission for the entire user group to access (or view, edit, etc.) the “Group Information” section of each collaborative document of the multiple collaborative documents. Automation in group-based communication systems

[0109] FIG. 3A illustrates a user interface 300 for automation in a group-based communication system. Automation, also referred to as workflow, allows users to automate functions within the group-based communication system. A workflow builder 302 is shown that allows users to create new workflows, modify existing workflows, and review workflow activities. The workflow builder 302 may include a workflow tab 304, an activity tab 306, and / or a settings tab 308. In some examples, the workflow builder may include a publish button 314 that allows users to publish new or modified workflows.

[0110] By selecting the Workflows tab 304, a user can create a new workflow or modify an existing workflow. For example, a user may want to create a workflow to automatically welcome new users who join a channel. A workflow may include workflow steps 310. Workflow steps 310 may include at least one trigger that starts the workflow and at least one function that takes an action when the workflow is triggered. For example, a workflow may be triggered when a user joins a channel, and the workflow's function may be to post in the channel welcoming the new user. In some examples, a workflow may be triggered from a user action, such as a user reacting to a message, joining a channel, or collaborating on a collaborative document, from a scheduled date and time, or from a web request from a third-party application or service. In further examples, a workflow function may include sending a message or form to a user, a channel, or any other virtual space, modifying a collaborative document, or interfacing with an application. Workflow functions may include workflow variables 312. For example, a welcome message may include the user's name via a variable to allow for a customized message. A user may edit existing workflow steps or add new workflow steps depending on the desired workflow function. Once a workflow is complete, the user may publish it using the publish button 314. A published workflow waits until it is triggered, at which point it performs its function.

[0111] The activity tab 306 may display information related to the activities of the workflow. In some examples, the activity tab 306 may indicate how many times the workflow has run. In further examples, the activity tab 306 may include information related to each workflow run, including the status, the date of the last activity, the run time, the user who started the workflow, and other relevant information. The activity tab 306 may allow a user to sort and filter the workflow activities to find useful information.

[0112] The settings tab 308 may allow a user to modify the settings of a workflow. In some examples, a user may change the title or icon associated with a workflow. A user may also manage collaborators associated with a workflow. For example, a user may add additional users to a workflow as collaborators so that the additional users can modify the workflow. In some examples, the settings tab 308 may also allow a user to delete a workflow.

[0113] 3B illustrates workflow-related elements within a group-based communication system, generally referred to as reference numeral 316. In various examples, trigger(s) 318 may be configured to invoke the execution of function(s) 336 in response to a user command. The trigger initiates function execution and may take the form of one or more schedule(s) 320, webhook(s) 322, shortcut(s) 324, and / or slash command(s) 326. In some examples, schedule 320 acts like a timer, scheduling the trigger to fire periodically or once at a predetermined time in the future. In some examples, an end user of an event-based application sets an arbitrary schedule for the firing of the trigger, such as once an hour or at 9:15 a.m. daily.

[0114] Additionally, the trigger 318 may take the form of a webhook 322. The webhook 322 may be a software component that listens on a webhook URL and port. In some examples, the trigger fires when a suitable HTTP request is received on the webhook URL and port. In some examples, the webhook 322 requires proper authentication, such as with a bearer token. In other examples, the trigger depends on the payload content.

[0115] Another source of a trigger for one of trigger(s) 318 is a shortcut within shortcut(s) 324. In some examples, shortcut(s) 324 may be global to the group-based communication system and not specific to a group-based communication system channel or workspace. A global shortcut may trigger a function that can be performed without the context of a particular group-based communication system message or group-based communication channel. In contrast, a message-based or channel-based shortcut is specific to a group-based communication system message or channel and operates in the context of the group-based communication system message or group-based communication channel.

[0116] A further source of trigger for one of triggers 318 may be provided by slash command 326. In some examples, slash command(s) 326 may serve as an entry point for group-based communication system functionality, integration with external services, or group-based communication system message responses. In some examples, slash command 326 may be entered by a user of the group-based communication system to trigger the execution of an application function. The slash command may be followed by slash command line parameters, which may be passed to any group-based communication system function invoked in connection with the trigger of the group-based communication system function, such as one of functions 336.

[0117] An additional way a function is invoked is when an event (such as one of events 328) matches one or more conditions predetermined in a subscription (such as subscription 334). Event 328 can be subscribed to by any number of subscriptions 334, with each subscription specifying different conditions and triggering different functions. In some examples, events are implemented as group-based communication system messages received in one or more group-based communication system channels. For example, all events may be posted as non-user-visible messages in the associated channel, which are monitored by subscription 334. App event 330 can be a group-based communication system message with associated metadata created by an application in a group-based communication system channel. Event 328 can also be a direct message received by one or more group-based communication system users, who can be actual users or technical users such as bots. Bots are technical users of the group-based communication system used to automate tasks. Bots can be programmatically controlled to perform various functions. Bots can monitor and help process group-based communication system channel activity as well as post messages in group-based communication system channels and react to members' in-channel activity. Bots can post messages and upload files as well as be invited or removed from both public and private channels in the group-based communication system.

[0118] Events 328 can be any events associated with the group-based communication system. Such group-based communication system events 332 include events related to the creation, modification, or deletion of a user account in the group-based communication system, or events related to messages in a group-based communication system channel, such as creating a message, editing or deleting a message, or reacting to a message. Events 328 can also be related to the creation, modification, or deletion of a group-based communication system channel or channel membership. Events 328 can also be related to changing a user profile or creating a group, managing members, or deleting a group.

[0119] As described above, subscriptions 334 indicate one or more conditions that, when matched with an event, trigger a function. In some examples, a set of event subscriptions is maintained in association with the group-based communication system, and when an event occurs, information about the event is matched against the set of subscriptions to determine which function 336 (if any) should be invoked. In some examples, the events to which a particular application may subscribe are managed by an authorization framework. In some cases, the event types matched against subscriptions are managed by OAuth permission scopes, which may be maintained by an administrator of a particular group-based communication system.

[0120] In some examples, function 336 can be triggered by triggers 318 and events 328 to which it subscribes. Function 336 takes zero or more inputs, performs processing (potentially including accessing external resources), and returns zero or more results. Function 336 can be implemented in a variety of forms. First, there are group-based communication system built-ins 338 associated with core functionality of a particular group-based communication system. Some examples include creating a group-based communication system user or channel. Second, there are no-code builder functions 340 that can be developed by a user of a group-based communication system user in association with an automation user interface, such as a workflow builder user interface. Third, there are hosted code functions 342 implemented by a group-based communication system application developed as software code in association with a software development environment.

[0121] These various types of functions 336 may then be integrated with APIs 344. In some examples, the APIs 344 are associated with third-party services that the functions 336 employ to provide custom integration between the particular third-party service and the group-based communications system. Examples of third-party service integrations include video conferencing, sales, marketing, customer service, project management, and engineering application integration. In such an example, one of the triggers 318 may be a slash command 326 used to trigger a hosted code function 342 that makes an API call to a third-party video conferencing provider via one of the APIs 344. As shown in FIG. 3B , the APIs 344 may themselves also be the source or event 328 of any number of triggers 318. Continuing with the example above, successful completion of a video conference triggers one of the functions 336 to send a message that initiates a further API call to the third-party video conferencing provider to download and archive the video conference recording and save it to a group-based communications system channel.

[0122] In addition to integrating with API 344, function 336 may persist and access data in tables 346. In some examples, tables 346 are implemented in association with a database environment associated with a serverless execution environment in which a particular event-based application is running. In some instances, tables 346 may be provided in association with a relational database environment. In other examples, tables 346 are provided in association with a database mechanism that does not employ relational database techniques. As shown in FIG. 3B , in some examples, the reading or writing of specific data to one or more of tables 346, or data in a table that matches predefined conditions, is itself the source of a number of triggers 318 or events 328. For example, if table 346 is used to maintain discovery data in an incident management system, the number of open tickets may exceed a predetermined threshold, triggering a message to be posted to an incident management channel in a group-based communication system. Determining summary data

[0123] FIG. 4 shows, for a particular example, a pictorial representation of an ML model that can be trained to receive an input and output a summary document using that input.

[0124] The server(s) 102 (not shown) may be configured to implement the ML model(s) 142 that ingests input data 405. The input data 405 may be environmental data comprising a synchronized multimedia collaboration session, including raw audiovisual data 405A, user reaction data 405B, contextual data, etc. As used herein, the phrase “raw audiovisual data” refers to audiovisual data as received or captured directly from a providing device. Thus, raw audiovisual data may include, for example, audiovisual data presented in a device-specific format or having processing steps already performed within the providing device or before reaching the providing device. However, such raw audiovisual data does not exclude environmental audiovisual data that has been or has been processed, filtered, or otherwise modified during or substantially simultaneously with the synchronized multimedia collaboration session that includes the audiovisual data. Examples of raw audiovisual data 405A include, without limitation, video streams, augmented / virtual reality streams, audio streams including uncompressed audio, lossless audio, lossy compressed audio, and modified or enhanced video or audio stream data including, for example, video or images modified by filters or the like. Additionally, server(s) 102 may, in some cases, be configured to extract audio data from video data and parse and analyze the audio and video data separately. Examples of user reaction data 405B include, without limitation, emojis, perceived physical expressions, text (and / or keyboard strokes), messages, etc.

[0125] As described above with respect to FIG. 2A , a communication channel or other virtual space may be associated with data and / or content. For example, the user interface 200 described above may include a composition pane in which a user may send a message including text (i.e., ASCII characters). Accordingly, various example ML model(s) 142 may be trained to receive and / or retrieve input data 405 to transcribe and / or summarize a synchronous multimedia collaboration session. That is, various example ML model(s) 142 may be trained to receive input data 405 and analyze raw audiovisual data and user reaction data to infer context related to the multimedia collaboration session and / or retrieve prior contextual data associated with the virtual space. For example, the ML model(s) 142 may be trained to receive and analyze messages or other data sent or posted to the virtual space to identify metadata associated with the messages, including one or more of a sending user identifier, a message identifier, a group identifier, a communication channel identifier, etc.

[0126] The ML model(s) may also be configured to recognize specific users within the virtual space. For example, users within the virtual space may have specific roles (e.g., CEO, CTO, etc.). The ML model(s) 142 may also be trained to learn relationships between users based on these roles. For example, the ML model(s) 142 may be trained to learn relationships between authorship and editorial metadata associated with messages and various file or document types to infer supervisor-supervisee relationships. The ML model(s) 142 may be trained to apply specific predetermined weights to users' communications based on their roles / relationships. For example, the channel owner's communications may be given a greater weight compared to other users in the channel. In some examples, the ML model(s) 142 may be trained to identify or receive input regarding weights to be applied to these individuals' communications, which may be prioritized when generating a summary document. In at least some examples, ML model(s) 142 may be trained to identify communications as constituting one or more of an assignment of a task to a user and a deadline for completing the task, and therefore assign a greater weight to such communications. However, ML model(s) 142 may additionally or alternatively prioritize particular user communications based on the recency of the user's communications, the frequency of the user's communications, communications (or other user interactions) above a threshold level, the user's communications being marked as favorites, the user's ratings, the user's expertise, the user's suggested preferences for their respective roles, user-specified parameters, user-specified permissions, heuristics from user activity, and / or the user's other interactions with previously generated summary documents.

[0127] Returning now to FIG. 4 , in some cases, the machine learning model 142 may include a single ML model that is trainable to fuse various data types, including the input data 405. For example, the machine learning model 142 may include a deep neural network (DNN) capable of multimodal fusion of audio, video, and text data. Thus, some example machine learning models 142 may include a DNN architecture that includes a separate fully connected layer for each data modality. That is, the ML model(s) 142 may be trained to learn relationships between instances of the same type of input data (e.g., by implementing a pair of fully connected layers for each data modality). The outputs of each fully connected layer may then be merged (e.g., by concatenation). Alternatively, the ML model 142 may include a first model and a second model (not shown), each trainable to learn relationships between different types of input data and transcripts. That is, the first model and the second model that make up some example ML models 142 may implement an ensemble learning algorithm. In some examples, the ML model(s) 142 may include a neural model for summarization such as a GPT-3 (Generative Pre-trained Transformer 3) model, an abstraction or generative summarization model, a machine learning summarization model, natural language processing, machine learning, and / or other techniques for identifying meaning and / or sentiment in messages within a virtual space and / or communication session.

[0128] The output data 415 may include a summary document 415A summarizing the posted content and communications that took place within the virtual space during the synchronized multimedia collaboration session. In this way, the communication platform can treat the summary document 415A as a summary or outline of the synchronized multimedia collaboration session and the tasks, deadlines, etc. discussed. In this way, functionality of the virtual space may be embedded in the summary document. The summary document 415A may accordingly identify the action(s), upload(s), conversation(s) (e.g., conversation threads), etc. that occurred within the virtual space. In particular, the summary document 415A may include data characterizing one or more objectives of the multimedia collaboration session with respect to tasks and / or deadlines. The data may include text, references or links to user objects, documents, and / or files, calendar objects, relationship data, etc. The text may be organized in any suitable format, including, but not limited to, bulleted lists, paragraphs, tables, etc. For example, the text may identify one or more users and their respective actions to be taken by them by username or other identifier. In some cases, summary document 415A may support editable text. In at least some examples, ML model(s) 142 may be trained to rank identified topics of conversation according to one or more relevance regimes. ML model(s) 142 may implement any suitable ranking metric, including average precision, discounted cumulative gain, etc. Thus, summary document 415A may categorize text or other content by relevance score or other indicator of overall importance.

[0129] The summary document may also identify the action(s), upload(s), conversation(s), etc. that occurred within the virtual space. For example, the summary document may include message timing documents (e.g., timestamps) associated with messages between users, audio and / or video communications, file uploads, notifications, new messages posted, mentions or tags, mention or tagging functionality, event propagation, workflow management and tracking, access permissions, and / or other functionality that occurred within the virtual space or that is enabled based on the type of virtual space.

[0130] Additionally, summary document 415A may include various objects that can be ordered, added, deleted, modified, and / or the like. Such objects may include, but are not limited to, files (e.g., text, messages previously sent via a communication platform, audio, video, applications, local objects, remote objects (e.g., third-party objects), etc.) or links thereto, and calendar objects indicating tasks or other events to be completed (e.g., scheduled events, calendar invitations, etc.). Thus, various example ML model(s) 142 may be tasked with using input data 405 to learn to automatically schedule meetings with participants in a synchronous multimedia collaboration session, automatically assign tasks from a list of tasks entered via a user interface (i.e., typed into a chat interface within a virtual space), automatically send links to documents referenced in a synchronous multimedia collaboration session, etc. Thus, some example ML model(s) 142 may be trained to use input data 405 to define a series of executable steps necessary to accomplish tasks discussed during a synchronous multimedia collaboration session. The executable steps may, in some examples, be defined by one or more summary objects (i.e., object code, etc.). In examples where summary document 415A is configured to be displayed within a user interface, execution of a series of steps may be initiated by user interaction. ML model(s) 142 may be trained to receive the user interaction and generate one or more summary interface elements associated with the summary object(s) to thereby initiate completion of the task.

[0131] For example, the ML model(s) 142 can receive audio / video data (e.g., from text / emoji-based messages by and between participants in a conference), audio / video data associated with context data (e.g., channel data, virtual space data, organizational data, relationship data, etc.), messages, emojis, etc., and can output one or more summary documents summarizing the content of the conference and / or provide any tasks, deadlines, and / or action items from the conference. The ML model(s) 142 can classify text based on comparing embeddings associated with the text from the conference with embeddings associated with the channel (e.g., intermediate output data of the ML model(s) 142, such as vectors that encode portions of the data) to determine relevance data. For example, the distance between embeddings can indicate that discussions in the conference are more similar to discussions in the channel and, therefore, may be more relevant in determining action items, tasks, deadlines, etc. If the distance between embeddings meets or exceeds a threshold, the text from the conference can be downweighted when determining the summary in the summary document.

[0132] FIG. 5 illustrates an exemplary process 500 associated with generating a summary document characterizing a synchronized multimedia collaboration session in a virtual space.

[0133] In some examples, process 500 may begin at operation 502, which may include receiving teleconference conference data associated with a channel of the group-based communication platform. In some examples, the teleconference conference data may be environmental data associated with a synchronous multimedia collaboration session occurring within the channel of the group-based communication platform. Thus, in some examples, the teleconference conference data may include live audiovisual data and user reaction data. Additionally, in some examples, the user reaction data includes one or more of an emoji selected by a user, a perceived physical expression of a user (e.g., a gesture detected from the video data using machine vision techniques), a message or text entered by a user (e.g., during the conference), or a thread of messages entered by multiple users (e.g., associated with a channel or virtual space). In some examples, the user reaction data includes a message or text entered into a user interface proximate to the video data substantially simultaneously during the generation of the audiovisual data.

[0134] In at least some examples, receiving raw audiovisual data includes capturing and / or converting audiovisual signals recorded by an analog or digital device. In some examples, receiving user reaction data includes at least capturing keystroke or mouse click data recorded by an analog or digital device. In other words, various examples of server(s) 102 may be configured to interpret selection of a user interface object representing a reaction (i.e., an emoji) or keystroke combination (e.g., a series of uppercase letters, a series of repeated letters, etc.) as representing a user reaction. In various other examples, receiving user reaction data includes at least capturing audiovisual data including a user reaction (e.g., voice inflection, gaze patterns, etc.).

[0135] In some examples, process 500 may include, at operation 504, performing natural language processing (NLP) on the raw audiovisual data using a first trained ML model, such as ML model(s) 142 described above, to generate transcript data associated with the teleconference conference. In some cases, performing NLP includes transforming and / or interpreting keystroke data recorded by an analog or digital device and may perform one or more of keystroke identification analysis, keystroke verification analysis, keystroke logging, etc. That is, ML model(s) 142 may be trained to interpret the keystroke data to identify participants in the synchronous multimedia collaboration session, verify that participants have provided particular input to the synchronous multimedia collaboration session, label syntactic patterns, etc. In some examples, the ML model(s) may be further trained to analyze and interpret the keystroke data to infer that the received corpus of keystroke data includes computer code. For example, ML model(s) 142 may be trained to recognize predefined keystroke combinations as constituting the beginning or end of coding syntax. In at least some cases, an ML model may be trained to rank the contextual relevance of individual instances of teleconference meeting data. Thus, in the above example, ML model(s) 142 may be trained to rank coding syntax as having higher contextual relevance within a communication channel dedicated to the launch of a new computer software product than audio data determined to be a conversation related to "lunch today."

[0136] In some examples, the ML model(s) 142 may include a single ML model, or in some examples, may include several ML models with separate tasks (e.g., a first ML model may convert speech to text, a second ML model may determine a summary, a third ML model may determine action items / associated details based on contextual data from the channel, etc.).

[0137] In some cases, generation of transcript data may occur in real time. Thus, in some examples, the ML model(s) may be trained to transform and / or interpret received audio-visual signals and perform one or more of speech recognition, speaker recognition, environmental sound recognition, or other such audio analysis. In some examples, performing NLP includes statistically correlating one or more interpreted emotions and / or gaze patterns with text data to infer context. In at least one non-limiting example, performing NLP may include determining semantic distances associated with recognized terms spoken or otherwise input by participants in the synchronous multimedia collaboration session. Thus, for example, if a synchronous multimedia collaboration session is initiated in a project_zen channel (dedicated to discussing ongoing software development work for Project Zen), the ML model(s) 142 may be trained to recognize any of the frequently occurring terms, such as “development,” “testing,” and “scripts,” as being particularly contextually relevant. In this example, the embeddings associated with these terms indicate that the discussion in the channel is more relevant when generating action items.

[0138] In some examples, process 500 may include, at operation 506, inputting the raw audiovisual data and user reaction data into a second trained ML model (which may be similar to ML model(s) 142 described above).

[0139] In some examples, process 500 may include, at operation 508, receiving a teleconference conference summary associated with the teleconference conference data as output from the second trained ML model. In at least some cases, the teleconference conference summary may be configured to be displayed within a user device display interface. Further, various example summaries may therefore include one or more summary objects configured for execution by the user device. As described above, the ML model may be configured to output a contextually actionable record or “summary document” representing the synchronous multimedia collaboration session. That is, the ML model(s) 142 of various examples described herein may be configured to create one or more data structures that can perform tasks within the group-based communication platform. Thus, in at least some cases, the summary document may include one or more collaborative documents, as described in further detail with reference to FIG. 2D . In at least some cases, the ML model(s) 142 may further be configured to associate one or more summary interface elements with one or more respective summary objects. Each summary interface element may be configured to receive user interaction and thereby initiate execution of the task(s) associated with the summary object. Returning momentarily to the Project Zen example, as shown in FIG. 2D , one or more summary objects may include a task list associated with the dates June 13 and 14. The corresponding summary interface elements include checkboxes configured to receive user interaction to indicate completion of particular tasks.

[0140] In some examples, process 500 may end at operation 510, which may include causing the display of the teleconference meeting summary. In at least some examples, causing the display of the teleconference meeting summary may be initiated by a user request.

[0141] 6 illustrates an example user interface 600 associated with a communications platform for enabling a user to request and generate a summary of a teleconference meeting. FIG. 6 also illustrates the generation of artificial intelligence (AI) notes or annotations during the teleconference meeting.

[0142] The exemplary user interface 600 illustrates an exemplary multimedia collaboration session (also referred to herein as a teleconference conference). Users of the group-based communication system can join or leave a synchronous multimedia collaboration session at any time without interrupting the session for other users. As described above in connection with FIG. 2A , a synchronous multimedia collaboration session may be based on a particular topic, a particular channel, a particular direct message, or a multi-party direct message to facilitate the completion of a task according to a deadline. In certain examples, a multimedia collaboration session is based on a channel (e.g., the #Team-native-ai channel). In some examples, a user may initiate a synchronous multimedia collaboration session within a channel as a means of communicating with other members of that channel who are currently online.

[0143] In at least one example, user interface 600 may include multiple sections, subsections, and / or panes that present various data or content. For example, first section 602 may include an agenda subsection 604, an action items subsection 606, a files subsection 608, an artificial intelligence (AI) notes (AI) subsection 610, and / or the like. In some examples, user interface 600 may include live video 620 of one or more users, a transcript 622 (generated in real time or near real time) of the teleconference meeting, threads 626, and / or tools 628 that allow users to control various functions of the collaboration session. While user interface 600 is shown as including certain fields or text, other content is contemplated.

[0144] In some examples, the content presented in the first section 602 may have been added or uploaded from another virtual space, such as a canvas, communication channel, workspace, or collaborative document associated with the communication platform. The content uploaded or pulled from the virtual space may include text data, audio data, video data, images, files, and / or any other type of data that may be configured to be presented in the first section 602 of a synchronous multimedia collaboration session. In some examples, the content items presented in the first section 602 may have been entered and / or edited by one or more users during a teleconference meeting. Alternatively, or additionally, one or more of the content items in the first section 602 may be generated and / or updated using an ML model.

[0145] In some examples, one or more ML model(s) may generate and output information that may be presented via user interface 600. For example, a first ML model may convert speech to text as the teleconference meeting progresses (e.g., transcript 622), a second ML model may further process or filter the text to remove redundant or irrelevant data, a third ML model may generate (i.e., output) a summary of the teleconference meeting from the filtered data input to the third ML model, a fourth ML model may determine action items / relevant details based on contextual data from the channel, a fifth ML model may output AI notes in near real time during the teleconference meeting, and / or a sixth ML model may be configured to predict the likelihood of receiving a request to generate a summary. In some examples, information generated by the ML models (e.g., transcript, teleconference meeting summary, AI notes, etc.) may be downloaded and / or uploaded to different parts of the communications platform (e.g., channels or workspaces).

[0146] In some examples, the ML model(s) may be trained based in part on analyzing and learning relationships (e.g., as learned relationship data) between (i) first data including prior audiovisual data and (ii) second data including prior summary data. In some examples, the ML model(s) may be configured to analyze channel context data based in part on channel data associated with the channel from which the teleconference meeting was initiated (e.g., as shown in FIG. 6, the meeting was initiated from the #Team-native-ai channel). In some examples, the channel context data may be input to an ML model configured to output a summary of the teleconference meeting.

[0147] Agenda subsection 604 may represent a list of items that the user(s) intend to discuss during the teleconference meeting. Action items subsection 606 may include completed or incomplete action items from previous teleconference meeting action item(s) copied or uploaded from a virtual space (e.g., a channel, a thread, a canvas, etc.) and / or action item(s) manually entered by the user as opposed to action items generated by an ML model. Files subsection 608 may include documents and / or links to documents, web pages, channels, etc. that may be related to the teleconference meeting. Information presented in agenda subsection 604, action items subsection 606, and / or files subsection 608 may be input to one or more ML models configured to output a transcript of the teleconference meeting, a summary of the teleconference meeting, AI notes during the teleconference meeting, and / or action items presented in the AI ​​notes. In some examples, this additional information may provide contextual data and assist the ML model(s) in determining the relevance of the information.

[0148] The transcript 622 may be generated by an ML model using known speech recognition techniques. The transcript 622 may include textual representations of the audio and / or video data, user reactions (e.g., emojis), messages typed in the thread 626, etc. In some examples, the audio / video component may generate the transcript simultaneously or substantially simultaneously with the teleconference conference. In some examples, the audio / video component may generate the transcript upon receiving an indication of selection of a “generate transcript” indicator (not shown), upon receiving an indication of selection of a record indicator, automatically upon initiation of the teleconference conference (e.g., based on permissions and / or settings associated with a user account, channel, workspace, communication platform, etc.). In at least one example, the audio / video component may be configured to cause presentation of the transcript 622 in association with a virtual space with which the audio and / or video conversation is associated. For example, a first user may initiate an audio and / or video conversation in association with a conference or communication channel.

[0149] In some examples, the audio / video component can process audio and / or video data between attendees of a teleconference meeting and generate a transcript of the audio, video, and / or message data. For example, the transcript can include time-stamped text typed in thread 626 (i.e., the transcript can include that User 3 typed "Brb one minute" into thread 626 at 9:30 AM). In some examples, the transcript can include reactions and / or emojis selected by one or more users during the teleconference meeting in response to what users said, their edits to the agenda, their edits to action items, their edits to AI notes, their thread messages, or their replies (e.g., a timestamp of a thumbs-up for User 3's thread message at 9:30 AM). In at least one example, the audio / video component can render one or more sections of the transcript selectable for commenting, such as to allow members of a communication channel to comment on or further contribute to the conversation. In some examples, the audio / video component can update the transcript based on the comments.

[0150] In some examples, the audio / video component may stop generating the transcript upon receiving an indication that the conversation is complete (e.g., an indication that the host or an administrator associated therewith has ended the collaboration conference, an indication that a threshold number of attendees (e.g., more than 50% of the attendees, more than 75% of the attendees, etc.) have closed the associated interface, an indication that the teleconference conference is no longer being recorded, and / or the like). In some examples, the audio / video component may identify the completion of the conversation and then generate a transcript associated therewith based on the completion.

[0151] In some examples, the ML model(s) may be configured to receive as input raw audiovisual data, user reaction data (e.g., emojis selected by a user), detected gestures in a video (e.g., a user shaking their head “no,” nodding “yes,” waving goodbye, etc.), messages or text entered by a user (e.g., a user making edits to an AI note 610), a thread of messages entered by multiple users (e.g., in thread 626), and / or other data, and use the text to filter the data to remove redundant and / or irrelevant data. Removing redundant and / or irrelevant data or information reduces the amount of data that an ML model configured to output a summary and / or AI notes needs to process. Output from a first ML model configured to filter the audiovisual data and / or user interaction data may be input to another ML model (e.g., a summarization model such as a large-scale language model (LLM)) that outputs a summary of the teleconference meeting and / or AI notes from the teleconference meeting. The use of additional ML models may reduce latency issues and / or error rates that may be associated with using a single ML model or using an ML model that processes unfiltered data. In some examples, determining which data is irrelevant is based at least in part on determining and comparing the context of all data associated with the teleconference meeting.

[0152] The user interface 600 may also include a “summarize” affordance 624 (e.g., a graphical indicator) that initiates a summary of a teleconference meeting or a portion of a teleconference meeting. For example, to initiate generation of a summary of a teleconference meeting (e.g., an audio and / or video conference in channel #Team-native-ai), a user may select the summarize affordance 624. In response, the user interface may display a dialog window that allows the user to enter details or criteria for the summary to be generated. For example, the user may request a summary of what a particular user or group of users said, a topic discussed, a discussion of keywords (e.g., emergency, deadline, rollout, etc.), a summary of a particular time of the meeting (e.g., the first five minutes and / or last five minutes), a summary of action items, a general theme, or any combination thereof. In such an example, the requested summary may be presented within a section or pane of the user interface 600, and a second window may prompt the user to download the summary to the user's computing device.

[0153] In some examples, a teleconference meeting summary may be automatically generated at the end of a conference. In some examples, a teleconference meeting summary may be automatically generated for a conference that lasts a threshold time (e.g., at least 10 minutes, 15 minutes, etc.). In some examples, an ML model may be configured to predict the likelihood that a user(s) will request a teleconference meeting summary at the end of a conference. Predicting when a teleconference meeting summary may be requested enables the ML model to pre-compute at least a portion of the teleconference meeting summary before the teleconference conference ends. For example, transcripts, AI notes, audiovisual data, and / or user reaction data may be input into ML model(s) configured to predict or estimate how likely a user(s) will request a meeting summary during a conference. For example, determining that a teleconference meeting summary is likely to be requested may be made based at least in part on transcripts, AI notes, audiovisual data, etc. indicating that a user had to leave the teleconference meeting early; based on a threshold number of action items (e.g., three or more) being generated or discussed during the meeting; based on the status or identity of one or more users in the meeting (e.g., a teleconference meeting summary is more likely to be requested when managers and / or new employees attend the meeting); based on the duration of the teleconference meeting (e.g., a summary is more likely to be requested when the teleconference meeting lasts a threshold amount of time); based on the number of users present in the meeting (e.g., users are more likely to request a summary when a threshold number of users attend and / or participate); based on topics discussed; based on the number of questions and / or answers discussed; etc.

[0154] In some examples, the AI ​​notes 610 may be generated using an ML model as the teleconference meeting progresses (i.e., in real time or near real time). In some examples, the ML model is configured to generate the AI ​​notes by processing data as it becomes available or based on a predetermined dataset or time frame (e.g., data received every minute, every two minutes, etc.). In some examples, the ML model may be prompted to generate a dataset when the transcript reaches a conclusion (e.g., when it determines there is a change in topic, detects that a user has joined / left the meeting, detects a number of reactions above a threshold or average, etc.). That is, the ML model may be configured to measure similarities between texts and predict whether the conversation has changed significantly. In at least one example, a dedicated ML model may be used to logically divide the transcript into datasets that can be input into a large-scale language model (LLM).

[0155] The generated AI notes 610 may include information related to the teleconference meeting. For example, the AI ​​notes 610 subsections may include user descriptions (e.g., "User 1 summarized the previous conversation about using markdown as output from the model"), responses, suggestions (e.g., "User 1 thinks we should continue using markdown for now because it easily captures bullet points," or "User 4 suggested transforming the markdown from the model into formatted headings and text styles"), questions (e.g., "User 2 asked what can be done about the tagging issue"), notes of when users left the meeting (e.g., "User 3 left the meeting at 9:30," as determined based on the transcript, threaded messages, replies, video camera powered off, user out of camera view, etc.). AM (User 3 left the meeting at 9:30 AM)").

[0156] In some examples, the AI ​​notes may include generated action items, takeaways, highlights, etc. of the teleconference meeting. In examples, the action items, short summaries, takeaways, and / or highlights may be presented in bulleted list format, paragraph format, table format, etc. Individual action items, short summaries, takeaways, and / or highlights may be associated with links that, when selected, direct the user to particular portions of the transcript and / or video associated with that particular action item, short summary, takeaway, or highlight. For example, selecting link 616 associated with the bullet point indicating "User 5 suggested giving the model semantically meaningful tags" may direct the user to portions(s) of transcript 622 where User 5 discussed this topic and / or highlight multiple portions of the transcript associated with this topic.

[0157] An action item may represent a recommended task that a user can perform at some point after the teleconference meeting ends. For example, as shown in FIG. 6, an AI memo may include a first action item 618(1) associated with user 6 and a second action item 618(2) associated with user 7. In some examples, an action item may be associated with due date. An action item may be associated with or directed to a single user, a group of users, or may be directed to users associated with a communication channel or other virtual space (e.g., a task may state, "Upload document A to Channel A"). In some examples, the action item(s) may be associated with an urgency level. For example, a first action item 618(1) may be associated with a first urgency level represented by a first symbol (e.g., a single exclamation mark, a yellow symbol), and a second action item 618(2) may be associated with a second urgency level represented by a second symbol (e.g., three exclamation marks, a red symbol, etc.). The urgency levels may be based at least in part on a due date associated with the action item, a user associated with the action item, the status or role of the user who proposed the action item (e.g., a manager suggested that an employee complete the action item), the context or relevance of the action item, etc.

[0158] In some examples, one or more users may collaborate on an AI note during a teleconference meeting (e.g., comment on and / or edit the AI ​​note 210 once the AI ​​note 610 is generated). For example, a user may select the edit affordance 612 to add or delete text, edit the urgency of an action item, or edit the assigned user for the action item. In some examples, one or more users may edit the text of an AI note without having to first select the edit affordance. In some examples, before starting an audio and / or video conference and / or during the conference, a user may indicate what information should be included in the AI ​​note. For example, a user may indicate that the AI ​​note provides a list of questions and answers discussed during the meeting, action items, suggestions offered by particular users, etc.

[0159] In some examples, summaries and / or AI notes generated during the teleconference meeting may be automatically uploaded to another virtual space (e.g., a canvas, a channel, etc.). In some examples, a user may select an export affordance 614 associated with the AI ​​note 610. By selecting the export affordance 614, the user may be presented with one or more options (e.g., via a pop-up window or a sidebar) for selecting a location to export the AI ​​note. In some examples, the AI ​​note may be automatically exported to a designated location within the communications platform (e.g., to the channel where the audio and / or video conference was initiated or to the canvas associated with the channel) at some point after the teleconference meeting has ended.

[0160] In some examples, the user interface may present tools 628. The tools 628 allow a user to perform actions and / or control features of the teleconference meeting. For example, selectable elements may allow a user to mute the user's microphone, turn the user's camera on / off, share the user's screen, send reactions (e.g., emojis), update user settings, leave the teleconference meeting, and / or invite other users. A screen share button not only allows a user to share their screen with other attendees of the teleconference meeting, but may also provide the user with additional controls during screen sharing. For example, a user sharing their screen may be provided with additional screen share controls to specify which screen to share, to annotate the shared screen, or to save the shared screen.

[0161] FIG. 7 shows an example user interface 700 for requesting and generating a summary of message threads in a virtual space (e.g., channel #Team-native-ai). The example user interface 700 may include a navigation pane 702. As described above in connection with FIG. 2A , the navigation pane 702 may include different graphical indicators that may be associated with various virtual spaces, such as channels 704. The example user interface 700 may include any number of channels 704 that may be used to organize conversations between users according to topic. In some examples, the example user interface 700 may include channels 704 such as a general channel, a social channel, a Team-native-ai channel 706, and / or any other channel. When a user selects a channel 704, a channel pane 708 may be presented. In some examples, the channel pane 708 may include access to content associated with the channel in addition to allowing the user to add other members, post content, etc. The channel pane 708 may include a user interface object (not shown) that, when activated (e.g., selected), may display a canvas (e.g., virtual space) associated with the selected channel.

[0162] In some examples, the exemplary user interface 700 may include a pane 708 for viewing virtual space data posted and / or entered by user(s) of the virtual space. As described above, some or all channels may include associated (or corresponding) canvases (e.g., virtual spaces). That is, the Team-native-ai channel 706 may include an associated canvas configured to curate, organize, and / or share a collection of information among users of the Team-native-ai channel 706. The exemplary user interface 700 may include a canvas pane 710. In this example, the canvas pane 710 may indicate that the user is viewing a portion of the Team-native-ai channel canvas.

[0163] In some examples, the channel pane 708 may include data (e.g., text) specific to the channel. While in a channel, a user can comment on messages in a “thread.” In some examples, a thread may be a message associated with another message that is not posted to a channel but is maintained within an object associated with the original message. Messages and / or threads may be associated with file(s), emoji(s), app(s), link(s), image(s), video(s), audio, and / or other attachments. The exemplary user interface 700 displays a thread 722 posted by user 3 that includes several replies.

[0164] The user interface 700 may include a messaging box 726 that allows users in the channel to post to the virtual space. In some examples, a user may request the generation of a thread summary. For example, a user may select an affordance associated with generating a thread summary (not shown). In another example, the summarization engine may use an ML model to generate a thread summary in response to receiving a command in the messaging box 726. The command may be used to summarize all or part of the data with a particular thread (e.g., replies, emojis, thread-specific attachments, links, etc.). Here, a user may enter (e.g., type) various commands into the messaging box 726 to invoke a bot, API, or other program of the communication platform to generate a thread summary. In some examples, the command may be entered as if it were a reply to thread 722. Exemplary commands include "@bot / summarize this thread" to summarize data in a particular thread, and " / summarize data in thread posted by" to generate a thread summary specifically of what User 1 posted.<User 1> ", or " / summarize data in thread regarding tags" or other topics or keywords. Although specific commands 724 are shown in Figure 7, other commands are contemplated.

[0165] In response to entering command 724 and submitting command 724 (e.g., pressing submit, send, enter, etc.), thread summary 728 may be presented. In some examples, thread summary 728 may be presented within user interface 700 as a reply to the thread, in a pop-up window, a sidebar, etc. The thread summary may be presented as a public post or a reply to the summarized thread, and other users with access (e.g., permissions) to the channel may view the summary. Alternatively, the thread summary may be private and viewable only to the user who requested the thread summary.

[0166] In some examples, the thread summary may be updated if additional communication occurs within the thread. For example, additional replies may be posted to thread 722 after thread summary 728 is generated, and in response, the summarization engine may utilize ML model(s) to generate a second thread summary or update the original thread summary. In some examples, the user(s) who requested the generation of the thread summary may be notified (e.g., receive an automated message, email, etc.) that a new thread summary is available for review. In some examples, before generating a new thread summary, the communication platform may send a notification to the user(s) confirming whether the user(s) would like an updated thread summary.

[0167] The thread summary 728 may include specific fields, text, and / or affordances. For example, the thread summary 728 may include summary analysis fields, topics and / or keywords discussed in the thread, thread highlights, action items for the user to perform, etc. In some examples, the thread summary 728 may present the user with a “send an updated summary” option if the thread receives additional replies, reactions, etc. In this example, the communication platform may send an updated thread summary to the user(s) based in part on using an ML model to determine that the additional replies substantially change the thread summary (e.g., changes to due dates, action items, topics discussed, etc.). However, if the additional replies include information that is redundant or irrelevant to the original summary thread (e.g., a reply such as “Ok. Have a good weekend!” or a thumbs-up emoji), the ML model may determine that the new replies, reactions, etc. are not substantially related to the original thread summary and may not automatically generate an updated thread summary.

[0168] In some examples, thread summary 728 may include customization affordance 730. Upon receiving an indication of a selection of customization affordance 730, the communication platform may render (e.g., in pane 708) a window or list (not shown) that includes various actions by which the communication platform may modify, update, and / or replace the data presented in the thread summary analysis. That is, customization affordance 730 may allow a user to customize the information presented in thread summary 728. For example, the user may update the output location of the thread summary, update permissions associated with the thread summary (public vs. private, identifying specific users who can and / or cannot view the thread summary), change the level of detail (e.g., verbosity) of the summary, identify the type of data in the thread that the user wants summarized (e.g., include summaries of documents in the thread, but not links), etc.

[0169] In some examples, thread summary 728 may include an add doc. affordance 732. In response to receiving an indication of a selection of add doc. affordance 732, the communication platform may cause a dialog window (not shown) to be presented in the virtual space. The dialog window may include several fields, information inputs, etc. associated with providing the user with options to include additional documents, threads, virtual spaces, etc. to create the thread summary.

[0170] In some examples, the thread summary 728 may include a feedback affordance 734. To improve the performance of the ML model(s), a user can provide feedback regarding the accuracy of the thread summary analysis. For example, a user may edit the thread summary, and the edited version may be used to retrain the ML model. In some examples, a user may rate the accuracy of the thread summary analysis (e.g., on a scale of 1 to 5, etc.).

[0171] In some examples, the canvas pane 710 may present a summary of a teleconference meeting, such as that described in connection with FIG. 6. Of course, this is merely an example, and the canvas may be directed to one or more alternative topics. In some examples, the canvas pane 710 may include a teleconference meeting summary that may be generated to provide a user with context regarding the conversations that took place during the teleconference meeting. The teleconference meeting summary may include a title 712 and indicate the date and / or time the summary was generated.

[0172] The teleconference meeting summary may include a short summary 714 that represents a brief description or overview of the teleconference meeting and is intended to provide users with insight into what was discussed during the teleconference meeting. In some examples, the teleconference meeting summary may include teleconference meeting highlights 716. For example, meeting highlights may include the more important or significant moments, points, or outcomes of the meeting. In some examples, highlights may include important takeaways or noteworthy aspects that effectively summarize what occurred during the meeting. Meeting highlights may depend on the purpose and nature of the meeting but may include, for example, major discussion points (major topics or issues discussed), decisions made during the meeting that resulted in action items, or resolutions. In some examples, highlights may include achievements or milestones. For example, if the meeting is centered around a project or initiative, achievements or milestones reached may be identified as highlights (e.g., project progress, goals achieved, obstacles overcome, etc.). In some examples, highlights may include important new information or insights that may impact the organization or project. In some examples, highlights may include challenges, problems, or issues discussed at the meeting that are important to the success of the project and proposed solutions or plans to address them. In some examples, highlights may include future plans or goals that define a strategy or objectives. In some examples, highlights may include key speakers or attendees and key takeaways and insights that key speakers provided. In some examples, highlights may include any important questions or concerns raised during the meeting and / or any formal agreements or resolutions (e.g., policy changes, etc.) reached during the meeting.

[0173] In some examples, the teleconference meeting summary may include action items 718 or tasks identified during the teleconference meeting. For example, the ML model may be configured to detect keywords or phrases indicative of action items (e.g., “to-do,” “action,” “task,” “follow-up,” “assign,” etc.). In some examples, the action items may be organized based on urgency level, importance, due date, department, user role or identity, etc. Action items may be associated with a particular user, a group of users, a virtual space (e.g., a channel), etc.

[0174] In some examples, the user interface 700 may include a messaging box 720 that allows users in a channel to post to the canvas. For example, a user may type in the messaging box 720, upload images, documents, suggest edits to a summary, react to a summary, etc. In response to a message, the messaging component may cause the message to be updated and displayed in the virtual space. In some examples, the messaging component may identify mentions or tags (e.g., @mention, or other special characters) of usernames, documents, channels, etc. associated with the communication platform. The mentions or tags may be associated with a user identifier. In response to identifying the mention or tag, the messaging component may cause a notification to be presented on the user interface associated with the user identifier.

[0175] In some examples, the teleconference meeting summary may include other or additional information about the teleconference meeting. For example, the meeting summary may include questions and answers discussed during the meeting, a list of users who attended the meeting and / or the level of participation of individual users in the meeting, important topics, phrases, or keywords discussed during the meeting, summaries or overviews of previous meetings related to the topic, dates of follow-up meetings (e.g., "A follow meeting is scheduled for next week, Sept. 25, 2023"), etc. In some examples, the meeting summary may include AI notes generated during the meeting (e.g., AI notes 610), portions of the AI ​​notes, or an affordance that, when selected, prompts the user device to download the AI ​​notes.

[0176] 8 illustrates an example process 800 associated with generating a summary of a virtual space (e.g., a summary of a teleconference meeting and / or a summary of a message thread). As described above, the virtual space may include means for transmitting written communications between users (e.g., direct message instances, communication channels, canvases, collaborative documents, canvases, audio calls, video calls, etc.).

[0177] In some examples, operation 802 of process 800 may include receiving one or more of message data or teleconference meeting data. That is, the process may include obtaining data associated with generating a summary, AI memo, thread summary, etc. of the teleconference meeting. The message data and teleconference meeting data may be associated with a virtual space of the communication platform and may represent exemplary virtual space data. In some examples, the virtual space data may include live audiovisual data of the teleconference meeting, user reaction data, emojis, perceived user expressions (e.g., gestures detected using machine vision techniques from video data), messages or text entered by users during the meeting, threads of messages entered by multiple users (e.g., associated with a channel or virtual space), file attachments (e.g., PDFs, text, videos, photos, etc.), links uploaded to the virtual space, etc.

[0178] In some examples, operation 804 of process 800 may include performing natural language processing (NLP) on the raw audiovisual data using a first trained ML model, such as the ML model(s) 142 described above, to generate transcript data associated with the teleconference conference. In some cases, generating the transcript data may occur in real time. Thus, some example ML model(s) may be trained to transform and / or interpret the received audiovisual signals and perform one or more of speech recognition, speaker recognition, environmental sound recognition, or other such audio analysis. In some examples, performing NLP includes statistically correlating one or more interpreted emotions and / or gaze patterns with text data to infer context. In at least one non-limiting example, performing NLP may include determining semantic distances associated with recognized terms spoken or otherwise input by participants in the synchronous multimedia collaboration session. Thus, for example, if a synchronous multimedia collaboration session is initiated within the Team-native-ai channel (dedicated to discussing ongoing machine learning model development), the ML model(s) 142 may be trained to recognize any of the frequently occurring terms, such as "development," "model," or "artificial intelligence," as being particularly contextually relevant. In this example, the embeddings associated with these terms indicate that the discussion within the channel is more relevant when generating action items.

[0179] In some examples, the ML model(s) 142 can include a single ML model, or in some examples, can include several ML models with discrete tasks (e.g., a first ML model can convert speech to text, a second ML model can process or filter the text to remove irrelevant text, a third ML model can determine a summary from the filtered text, a fourth ML model can determine action items / relevant details based on contextual data from the channel, a fifth ML model can output notes in near real time during a synchronous multimedia collaboration session, a sixth ML model can be configured to predict when a summary is likely to be requested by user(s) before the user(s) request the summary to be generated, etc.). In at least one example, the ML models can be configured to output different types of output (e.g., AI notes and summaries of the synchronous multimedia collaboration session) based on receiving different prompts.

[0180] In some examples, operation 806 of process 800 may include inputting one or more of the transcript, raw text data, audiovisual data, user reaction data, and / or other virtual space data generated in operation 804 into a second trained ML model. The second ML model may be configured to filter out irrelevant text to improve the efficiency and performance of the ML model(s) used to generate the summary and / or AI notes. Removing redundant or irrelevant information can reduce the size of the dataset and, therefore, the computational complexity of the ML model algorithm. In some examples, the second ML model may filter the virtual space data using embedding techniques (e.g., word embeddings, user and item embeddings, etc.), by incorporating feature selection into the training process of the ML algorithm (e.g., lasso and ridge regression, decision trees, neural networks, etc.), using filter-based techniques that include selecting features based on statistical or mathematical criteria (e.g., variance thresholds, correlation-based feature selection, chi-square tests, etc.), or using wrapper methods that include selecting features by training the ML model multiple times on different subsets of features (e.g., forward selection, backward elimination, recursive feature elimination, etc.), etc.

[0181] In some examples, operation 808 of process 800 may include receiving filtered text data, audiovisual data, user reaction data, and / or other virtual space data as output from the second trained ML model.

[0182] In some examples, operation 810 of process 800 may include inputting the filtered text data, audiovisual data, user reaction data, and / or other virtual space data into a third ML model. In some examples, the third ML model may be configured to output one or more of a teleconference meeting summary, an AI memo, a thread summary, an action item, or the like. In at least some cases, the teleconference meeting summary, the AI ​​memo, the thread summary, and the action item may be configured to be displayed within an interface of a display of a user device. Furthermore, the teleconference meeting summary, the AI ​​memo, the thread summary, and / or the action item may include one or more objects configured for execution by the user device. As described above, the ML model(s) may be configured to output a contextually actionable record or “summary document” representing a synchronous multimedia collaboration session or thread. That is, the ML model(s) 142 of various examples described herein may be configured to create one or more data structures that can perform tasks within a group-based communication platform. Thus, in at least some cases, the generated summary may include one or more collaborative documents, as described in further detail with reference to FIG. 2D . In at least some cases, the ML model(s) 142 may be further configured to associate one or more summary interface elements (e.g., check boxes) with one or more respective summary objects, each of which may be configured to receive user interaction and thereby initiate execution of a task(s) associated with the summary object.

[0183] In some examples, operation 812 of process 800 may include receiving a request to generate a summary of the teleconference meeting. Alternatively, process 800 may proceed to operation 814, described below. In some instances, the summarization component may receive a request to generate a summary from a user. In some examples, the request may be received in association with a virtual space. In some examples, the request may specify content to be used when generating the teleconference meeting summary or context. For example, the user interface may display a dialog window that allows a user to enter details or criteria for the summary to be generated via the dialog window (e.g., generate a summary of what a particular user said, a group of users, a topic discussed, mentions of keywords or phrases, summary information received during a particular time or times of the meeting, etc.).

[0184] In some examples, operation 814 of process 800 may include generating a summary of the teleconference conference using an ML model (e.g., a third ML model trained to output a teleconference conference summary). For example, summarization engine 120 may utilize one or more ML models to summarize content discussed, presented, and / or associated with the teleconference conference. In some examples, the teleconference conference summary may be configured to be displayed within a display interface of a user interface. In some examples, the teleconference conference summary may include one or more highlights, action items, a participant list, AI notes, keywords, a list of questions and answers asked during the teleconference conference, etc.

[0185] In some examples, operation 816 of process 800 may include receiving a request to generate a summary of a thread. For example, a user may select an affordance associated with generating a thread summary. In another example, the summarization engine may generate a thread summary using an ML model in response to receiving a command in a messaging box. The command may request a summary of all or some of the data associated with a particular thread (e.g., replies, emojis, thread-specific attachments, links, etc.). For example, a user may enter (e.g., type) various commands into a messaging box to invoke a bot, API, or other program of a communication platform to generate a thread summary using an ML model.

[0186] In some examples, operation 818 of process 800 may include generating a summary of the thread. That is, the thread summary may be received as output from an ML model trained to output thread summaries. In some instances, the thread summary may be configured to be displayed within an interface on a display of a user device. Example clauses

[0187] A: A computer-implemented method for summarizing a teleconference conference of a group-based communication platform, the computer-implemented method including: receiving teleconference conference data associated with a channel of the group-based communication platform, the teleconference conference data including audiovisual data and user reaction data, the user reaction data including one or more of an emoji selected by a user, a detected gesture associated with the user, a message or text entered by the user, and a thread of messages entered by multiple users; performing natural language processing (NLP) on the audiovisual data using a first trained machine learning model (MLM) to generate transcript data associated with the teleconference conference; inputting the audiovisual data and the user reaction data into a second trained MLM; receiving a teleconference conference summary associated with the teleconference conference data as output from the second trained MLM, the teleconference conference summary configured to be displayed within an interface of a display of a user device, and further, the teleconference conference summary including one or more summary objects configured to be executed by the user device; and causing display of the teleconference conference summary.

[0188] B. The computer-implemented method of paragraph A, wherein receiving teleconference meeting data includes receiving one or more of audiovisual data and user reaction data in real time.

[0189] C. The computer-implemented method of paragraph A or B, wherein causing the display of the teleconference meeting summary includes at least causing the display of one or more summary interface elements representing one or more summary objects.

[0190] D. The computer-implemented method of paragraph C, further comprising: in response to detecting user interaction with one or more summary interface elements representing one or more summary objects, identifying one or more executable steps associated with the one or more summary objects; and causing execution of the one or more executable steps associated with the one or more summary objects.

[0191] E. The computer-implemented method of any of paragraphs A-D, wherein the second trained MLM is trained to determine relationship data between the first data and the second data based on (i) the first data including previous audiovisual data and (ii) the second data including previous summary data, whereby the second trained MLM is configured to learn to use the relationship data to generate a teleconference meeting summary upon input of the audiovisual data.

[0192] F. The computer-implemented method of any of paragraphs A-E, further including: determining contextual data based in part on channel data associated with a channel of the group-based communications platform; and inputting the contextual data into a second trained MLM, wherein output from the second trained MLM is further based on the contextual data.

[0193] G. The computer-implemented method of any of paragraphs A-F, wherein the context data includes one or more messages generated asynchronously with the audiovisual data.

[0194] H. The computer-implemented method of any of paragraphs A-G, further including: generating, by the second trained MLM, an action item associated with the teleconference meeting summary; generating a user identifier based on the action item; and transmitting the action item to a computing device associated with the user identifier so that it is presented on a display associated with the instance of the group-based communication platform associated with the user identifier.

[0195] I. The computer-implemented method of any of paragraphs A-H, wherein generating an action item includes classifying text associated with the teleconference meeting and determining a relevance weight associated with the text based on classifying the text.

[0196] J. The computer-implemented method of any of paragraphs A-I, wherein a message or text entered by a user is entered into a user interface proximate to the video data substantially simultaneously while the audiovisual data is being generated.

[0197] K. A system, comprising: one or more processors; and non-transitory memory storing processor-executable instructions that, when executed by the one or more processors, cause the system to perform operations including receiving teleconference conference data associated with a teleconference conference of a group-based communications platform, the teleconference conference data including audiovisual data and user reaction data; performing natural language processing (NLP) on the audiovisual data using a first trained machine learning model (MLM) to generate transcript data associated with the teleconference conference; inputting the audiovisual data and the user reaction data into a second trained MLM; and receiving as output from the second trained MLM a teleconference conference summary associated with the teleconference conference data, the teleconference conference summary configured to be displayed within an interface of a display of a user device, and further, the teleconference conference summary including one or more summary objects configured to be executed by the user device; and causing display of the teleconference conference summary.

[0198] L. The system of paragraph K, wherein the second trained MLM is trained to learn a relationship between the first data and the second data based on (i) the first data including previous audiovisual data and (ii) the second data including previous summary data, whereby the second trained MLM is configured to learn to generate a teleconference meeting summary upon input of the audiovisual data using the learned relationship.

[0199] M. The system of paragraph K or L, wherein the teleconference meeting data is associated with a channel of a group-based communications platform, and the operations further include determining contextual data based in part on channel data associated with the channel of the group-based communications platform, and inputting the contextual data into a second trained MLM, wherein output from the second trained MLM is further based on the contextual data.

[0200] N. The system of any of paragraphs K-M, wherein the context data includes one or more messages generated asynchronously with the audiovisual data.

[0201] O. The system of any of paragraphs K-N, wherein the operations further include generating, by the second MLM, an action item associated with the teleconference meeting summary; generating a user identifier based on the action item; and transmitting the action item to a computing device associated with the user identifier for presentation on a display associated with the instance of the group-based communications platform associated with the user identifier.

[0202] P. The system of any of paragraphs K-O, wherein receiving teleconference meeting data includes receiving one or more of audiovisual data and user reaction data in real time.

[0203] Q. One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform operations including receiving teleconference conference data associated with a teleconference conference of a group-based communications platform, the teleconference conference data including audiovisual data and user reaction data; performing natural language processing (NLP) on the audiovisual data using a first trained machine learning model (MLM) to generate transcript data associated with the teleconference conference; inputting the audiovisual data and the user reaction data into a second trained MLM; and receiving as output from the second trained MLM a teleconference conference summary associated with the teleconference conference data, the teleconference conference summary configured to be displayed within an interface of a display of a user device, and further, the teleconference conference summary including one or more summary objects configured to be executed by the user device; and causing display of the teleconference conference summary.

[0204] R. The one or more non-transitory computer-readable media described in paragraph Q, wherein the operations further include generating, by the second MLM, an action item associated with the teleconference meeting summary; generating a user identifier based on the action item; and transmitting the action item to a computing device associated with the user identifier for presentation on a display associated with the instance of the group-based communications platform associated with the user identifier.

[0205] S. The one or more non-transitory computer-readable media described in paragraph Q or R, wherein the user reaction data includes messages or text entered into a user interface proximate to the video data substantially simultaneously during the generation of the audiovisual data.

[0206] T. The one or more non-transitory computer-readable media of any of paragraphs Q to S, wherein the teleconference meeting data is associated with a channel of a group-based communication platform, and the operations further include determining contextual data based in part on channel data associated with the channel of the group-based communication platform, and inputting the contextual data into a second trained MLM, wherein output from the second trained MLM is further based on the contextual data.

[0207] Although the example clauses described above are described with respect to one particular implementation, it should be understood that in the context of this document, the content of the example clauses may also be implemented via a method, a device system, a computer-readable medium, and / or another implementation. Additionally, any of Examples A-T may be implemented alone or in combination with any other one or more of Examples A-T. conclusion

[0208] One or more examples of the techniques described herein have been described, and various modifications, additions, permutations, and equivalents thereof fall within the scope of the techniques described herein.

[0209] In describing the examples, reference is made to the accompanying drawings, which form a part of this specification and which show, by way of illustration, specific examples of the claimed subject matter. It is understood that other examples can be used and changes or modifications, such as structural changes, can be made. Such examples, changes, or modifications do not necessarily depart from the intended scope of the claimed subject matter. While the steps herein may be presented in a particular order, in some cases the order can be changed so that certain inputs are provided at different times or in a different order without changing the functionality of the systems and methods described. The disclosed procedures may also be performed in a different order. Additionally, the various calculations herein need not be performed in the order disclosed, and other examples using alternative orders of calculations can be readily implemented. In addition to being reordered, calculations can also be decomposed into sub-calculations that have the same result.

Claims

1. 1. A computer-implemented method for summarizing a teleconference meeting of a group-based communication platform, comprising: receiving teleconference meeting data associated with a channel of the group-based communications platform, the teleconference meeting data comprising: Audiovisual data; User reaction data and wherein the user reaction data includes emoji selected by the user, a detected gesture associated with the user; a message or text entered by the user, and Message threads entered by multiple users and performing natural language processing (NLP) on the audiovisual data using a first trained machine learning model (MLM) to generate transcript data associated with the teleconference meeting; inputting the audiovisual data and the user reaction data into a second trained MLM; receiving, as output from the second trained MLM, a teleconference meeting summary associated with the teleconference meeting data; the teleconference meeting summary configured to be displayed within an interface of a display of a user device, the teleconference meeting summary further comprising one or more summary objects configured to be executed by the user device; causing a display of said teleconference meeting summary; 10. A computer-implemented method comprising:

2. The computer-implemented method of claim 1 , wherein receiving the teleconference meeting data comprises receiving one or more of the audiovisual data and the user reaction data in real time.

3. The computer-implemented method of claim 1 , wherein causing the display of the teleconference meeting summary comprises at least causing the display of one or more summary interface elements representing the one or more summary objects.

4. identifying one or more executable steps associated with the one or more summary objects in response to detecting user interaction with the one or more summary interface elements representing the one or more summary objects; causing execution of said one or more executable steps associated with said one or more summary objects; The computer-implemented method of claim 3 further comprising:

5. 2. The computer-implemented method of claim 1, wherein the second trained MLM is trained to determine relationship data between first data and second data based on (i) first data including previous audiovisual data and (ii) second data including previous summary data, whereby the second trained MLM is configured to learn to use the relationship data to generate a teleconference meeting summary upon input of the audiovisual data.

6. determining context data based in part on channel data associated with the channel of the group-based communications platform; inputting the context data into the second trained MLM, wherein the output from the second trained MLM is further based on the context data; The computer-implemented method of claim 1 , further comprising:

7. The computer-implemented method of claim 6 , wherein the contextual data includes one or more messages generated asynchronously with the audiovisual data.

8. generating, by the second trained MLM, action items associated with the teleconference meeting summary; generating a user identifier based on the action item; transmitting the action item to a computing device associated with the user identifier so that it is presented on a display associated with the instance of the group-based communication platform associated with the user identifier; The computer-implemented method of claim 1 , further comprising:

9. The step of generating an action item comprises: classifying text associated with said teleconference meeting; determining a relevance weight associated with the text based on classifying the text; The computer-implemented method of claim 8 , comprising:

10. The computer-implemented method of claim 1 , wherein the message or text entered by the user is entered into a user interface proximate to video data substantially simultaneously during the generation of the audiovisual data.

11. 1. A system comprising: one or more processors; a non-transitory memory for storing processor-executable instructions; wherein the processor-executable instructions, when executed by the one or more processors, cause the system to: receiving teleconference conference data associated with a teleconference conference of a group-based communication platform, the teleconference conference data including audiovisual data and user reaction data; performing natural language processing (NLP) on the audiovisual data using a first trained machine learning model (MLM) to generate transcript data associated with the teleconference meeting; inputting the audiovisual data and the user reaction data into a second trained MLM; receiving, as output from the second trained MLM, a teleconference meeting summary associated with the teleconference meeting data; receiving the teleconference meeting summary configured to be displayed within an interface of a display of a user device, the teleconference meeting summary further comprising one or more summary objects configured to be executed by the user device; causing a display of said teleconference meeting summary; A system that causes an operation including

12. 12. The system of claim 11, wherein the second trained MLM is trained to learn a relationship between (i) first data including previous audiovisual data and (ii) second data including previous summary data based on the first data and the second data, whereby the second trained MLM is configured to learn to generate a teleconference meeting summary upon input of the audiovisual data using the learned relationship.

13. The teleconference meeting data is associated with a channel of the group-based communications platform, and the operation comprises: determining context data based in part on channel data associated with the channel of the group-based communications platform; inputting the context data into the second trained MLM, wherein the output from the second trained MLM is further based on the context data; and The system of claim 11 further comprising:

14. The operation is generating, by the second trained MLM, action items associated with the teleconference meeting summary; generating a user identifier based on the action item; transmitting the action item to a computing device associated with the user identifier so that the action item is presented on a display associated with the instance of the group-based communication platform associated with the user identifier; The system of claim 11 further comprising:

15. 15. One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the operations of any one of claims 1 to 14.

Citation Information

Patent Citations

  • Computer-implemented method, computer-implemented system and computer program product for generating summaries and next actions in real time for multiple users from interaction records in natural language

    JP2022105273A

  • Meeting summary service

    US11095468B1

  • Generating improved digital transcripts utilizing digital transcription models that analyze dynamic meeting contexts

    US20200403818A1

  • Virtual meeting

    US20210320811A1