Techniques for inferring context for online meetings
By combining speech-to-text and computer vision algorithms to process verbal and non-verbal communication information in online video conferences, and using generative language models to generate meeting summary descriptions, the problem of incomplete and inaccurate meeting summaries in existing technologies is solved, achieving more efficient meeting recording and information extraction.
Patent Information
- Application Number
- CN202480031017.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-20
- Filing Date
- 2024-06-11
- Publication Date
- 2025-12-09
AI Technical Summary
Existing online video conferencing systems are unable to effectively capture and integrate nonverbal communication information, resulting in incomplete and inaccurate meeting summaries, making it difficult to provide accurate meeting summaries for those who miss the meeting.
By combining speech-to-text and computer vision algorithms to process video streams, verbal and nonverbal communication information of conference participants is extracted, and a generative language model is used to generate a summary description of the conference.
It provides a more complete and accurate meeting summary, capturing nonverbal communication information during the meeting and improving the efficiency and accuracy of meeting minutes.
Smart Images

Figure CN121100522A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] This application relates to video-based conferencing over computer networks such as the Internet. More specifically, this application describes methods, systems, and computer program products for generating text using computer vision algorithms to describe or represent non-verbal communication identified within a video, and extracting text (e.g., word groups) from content that has been shared during an online meeting, and then providing these text elements as input to a meeting analyzer service that utilizes a generative language model to generate a summary description of the online meeting. BACKGROUND
[0002] Online video conferencing and meeting services have revolutionized the way individuals and teams connect, enabling seamless collaboration regardless of distance. These platforms and services provide virtual spaces where participants can join from anywhere, whether it be one-on-one, small team huddles, or large-scale meetings. With features such as content sharing, users can effortlessly present documents, slides, and multimedia files, facilitating engagement and boosting productivity. These online meeting services facilitate real-time communication through high-quality video and audio, ensuring clear interaction and enabling participants to see and hear each other clearly. With intuitive interfaces, simple scheduling options, and ubiquitous chat functionality, online video conferencing and meeting services provide a comprehensive solution for effective and dynamic remote collaboration. BRIEF DESCRIPTION OF DRAWINGS
[0003] Embodiments of the application are illustrated by way of example, and not by way of limitation, in the accompanying drawings in which: Figure 1 FIG. 1 is a diagram illustrating an example of a computer network environment with which an online meeting service is deployed, in accordance with some embodiments.
[0004] Figure 2 FIG. 2 is a diagram illustrating how a media processing service of an online meeting service receives and processes a combination of video and content sharing streams to generate text used as input to a generative language model, and more specifically to derive a summary description of an online meeting, in accordance with some embodiments.
[0005] Figure 3 FIG. 3 is a diagram illustrating an example of a user interface of a client application for an online meeting service, in accordance with embodiments of the application, where the user interface is presenting a live video stream of a meeting participant performing a first gesture (e.g., a nod gesture).
[0006] Figure 4FIG. 1 is a diagram illustrating an example of a user interface of a client application for an online meeting service, according to embodiments of the application, in which the user interface is presenting a shared content (e.g., a content presentation), and in particular a first slide or page of the presentation.
[0007] Figure 5 FIG. 2 is a diagram illustrating an example of a user interface of a client application for an online meeting service, according to embodiments of the application, in which the user interface is presenting a live video stream of a meeting participant performing a second gesture (e.g., a thumbs up gesture).
[0008] Figure 6 FIG. 3 is a diagram illustrating an example of a user interface of a client application for an online meeting service, according to embodiments of the application, in which the user interface is presenting a content share stream, and in particular a second slide or page of the presentation.
[0009] Figure 7 FIG. 4 is a diagram illustrating an example of an annotated text-based transcript for an online meeting, according to embodiments of the application, that is generated to include text describing non-verbal communication and word groups extracted from a presentation.
[0010] Figure 8 FIG. 5 is a block diagram illustrating an example of functional components of a media processing service that includes an online meeting service, according to some embodiments.
[0011] Figure 9 FIG. 6 is a diagram illustrating an example of various functional components for an online meeting analyzer that includes at least one generative AI model, such as a general purpose language model, for generating a summary description of an online meeting, according to some embodiments.
[0012] Figure 10 FIG. 7 is a diagram illustrating a software architecture that can be installed on any of various computing devices to perform methods according to those described herein.
[0013] Figure 11 FIG. 8 is a system diagram illustrating an example of a computing device that can implement embodiments of the application. DETAILED DESCRIPTION
[0014] Methods, systems, and computer program products are described herein for inferring meeting context for an online meeting, for example, by analyzing video to identify nonverbal communication and by extracting word groups from content shared during the online meeting. The meeting context is generated in the form of text that is then used in combination with text from other information sources to generate insights about the meeting. For example, the text can be provided as input to a software-based meeting analyzer service that allows end users to submit queries and ask questions to gain insights about what happened during the online meeting. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the various aspects of the present invention. It will be apparent, however, to one skilled in the art that the present invention can be practiced without all these specific details.
[0015] Online meeting services, sometimes referred to as video conferencing platforms, allow meeting participants to connect and communicate in real-time through audio, video, and collaboration content presentation and screen sharing tools. For many online meeting services, each meeting participant will register to create an account, typically providing an email address and establishing a username and password that are then used for authentication purposes. A meeting participant is able to schedule a meeting in advance, for example, by specifying a date, time, and meeting duration, and then sharing a unique meeting identifier that other meeting participants will use to join or enter the online meeting.
[0016] Once a meeting participant has joined an online meeting, the meeting participant typically has the option to select and configure (e.g., enable / disable) their audio and video devices. These devices can include, for example, a microphone, speaker(s), and a camera with video capture capabilities. These devices can be built into a client computing device on which a client software application for the online meeting service is installed and executed. Alternatively, a meeting participant can choose to use an external microphone, external speaker(s) (including headphones with a combined microphone and speaker), or an external video camera, which can provide a higher quality experience.
[0017] As in Figure 1As illustrated in the middle, the online meeting service 100 establishes real-time audio and video connections between the meeting participants. A client software application executing at each meeting participant's client computing device uses audio and video codecs to compress audio and video data and transmit it over a network 102, such as the Internet. The meeting participants are able to communicate with one another by speaking and listening through microphones and speakers of their client computing devices, while viewing video streams via displays of their devices. The client-based software application for the online meeting service can allow each participant to select from several different visual arrangements via a user interface. For example, as illustrated with reference number 104, a meeting participant can select a user interface that displays all meeting participants in a grid arrangement, while other meeting participants can select a user interface 106 that displays only active speakers. Various other user interface arrangements are possible.
[0018] Most online meeting services provide additional collaboration tools to enhance the meeting experience. For example, some online meeting services provide content sharing or presentation tools. These tools or meeting features allow each meeting participant to share content that is being presented on his or her screen or display, or content that is associated with a particular software application executing on his or her computing device. In some cases, only one meeting participant is allowed to share content at a time, while in other cases multiple meeting participants can be able to share content simultaneously.
[0019] With the rise of remote work, the popularity and use of online meeting services has skyrocketed. As the total number of meetings conducted increases, the likelihood that any one person will miss a meeting increases significantly. Moreover, with so many different meetings, it can be difficult to keep track of what happened on one particular online meeting and what can have happened on a different online meeting. Thus, there is a growing need for technology and tools that allow people, and in particular people who can have missed some or all of an online meeting, to obtain relevant information about what happened during the online meeting.
[0020] To this end, some online meeting services provide the ability to automatically record an online meeting, enabling the recording (e.g., a video file) to be played back later by someone who can have missed the meeting. Some online meeting services process the audio portion of the video stream by using an item commonly referred to as a speech-to-text algorithm or model to analyze the audio and convert spoken conversations that occur during the meeting into a text-based transcript. The resulting text-based transcript generally identifies the speaker and can include a time at which each spoken message occurred. Additionally, if meeting participants exchange text-based messages via a chat or messaging feature, many online meeting services make this chat transcript available at the end of the online meeting.
[0021] There are several problems with these approaches. First, with the approaches described above, there is no single source of information from which a person can obtain a complete and accurate understanding of what happened at a meeting. For example, if a person reviews the complete video recording of an online meeting, the person will also need to review the chat transcript to ensure that he or she has not missed some critical information. Second, reviewing the complete video recording of an online meeting is a very time-consuming task. Often, the person's information need is represented by only some small portion of the total conversation that occurred during the online meeting or a single slide or page in a multi-slide or multi-page content presentation. Thus, more often, reviewing the entire recording of an online meeting is an extremely inefficient way to obtain the desired information. Third, reviewing a text-based transcript derived via a text-to-speech algorithm is also a time-consuming and inefficient task. The text-based transcript is itself an incomplete representation of the online meeting. The text-based transcript lacks contextual information that can be critical to obtaining a complete and accurate understanding of what happened during the online meeting. For example, if a meeting participant shares a content presentation (e.g., via a screen share or application share feature), the text-based transcript will not reflect any information from the shared content presentation. In addition, the text-based transcript can lack other contextual information such as non-verbal communications or cues such as the meeting participant's facial expressions, posture, and body language. These non-verbal communications can provide important insights into the meeting participant's mood, intent, agreement or disagreement, and overall level of engagement during the online meeting.
[0022] One potential new approach for generating summary information for an online meeting involves using a generative language model, such as a large language model (LLM). For example, the audio portion of a video can be processed using speech-to-text to generate a text-based transcript that represents the conversation between the meeting participants. Portions of the text can then be used within a prompt that is provided as input to a generative language model, where the prompt is designed to serve as a question or instruction for generating a summary description of some portion of the online meeting. However, because the text-based transcript that is derived by converting the audio of a spoken conversation to text is not a complete and accurate representation of the online meeting, the resulting text-based responses generated by the generative language model can be incomplete, inaccurate, or misleading. This is due at least in part to the input text not being a complete and accurate representation of the meeting. As an example - if a first meeting participant asks a non-question and several meeting participants nod (e.g., up and down) - a gesture that is widely interpreted to mean that a person agrees - because these gestures are not included in the text-based transcript, the generative language model can not generate an appropriate response to the particular prompt. In fact, the generative language model can generate a completely inaccurate response by suggesting that the answer to the particular question is “no” when in fact the answer is “yes.” Thus, when a generative language model is provided with a text-based transcript, the generative language model can provide an inaccurate, incomplete, or misleading summary description due to the fact that gestures made by multiple meeting participants are not reflected in the input text available to the generative language model.
[0023] According to embodiments of the present invention, the above technical problems are addressed by utilizing various machine learning and natural language processing techniques to generate a complete and accurate digital representation of an online meeting, from which an accurate and complete summary description of the online meeting is generated. A software-based media processing service receives and processes a combination of video streams - including a first type of video stream that originates from a camera device of a meeting participant and that generally depicts a video image of the meeting participant, and a second type of video stream that includes shared content (e.g., that can be generated by a screen share or application share feature) that originates from a computing device of a meeting participant. Each video stream is independently analyzed using various pre-trained machine learning models of the media processing service to generate text. The text generated by the media processing service is then provided as input to a software-based meeting analyzer that uses the text to derive accurate insights about what occurred during the online meeting using various techniques.
[0024] According to some embodiments, each video stream originating from a camera device of a meeting participant is received and processed by a first type of pre-trained machine learning model that converts the audio portion of the video stream representing the spoken messages of the meeting participant into text. This type of machine learning model is commonly referred to as a speech-to-text model. In addition, each video stream is also processed by one or more additional pre-trained machine learning models, specifically one or more computer vision models, that identify or detect non-verbal communications made by the meeting participant. According to some embodiments, these models can output text that describes or represents these non-verbal communications. In some embodiments, in addition to detecting and generating text to represent the detected gestures, one or more pre-trained machine learning models can also be used to detect or identify facial expressions. Similarly, one or more pre-trained machine learning models can be used to detect or identify emotions or body language expressed by the meeting participant. The output of these models is text that describes or otherwise represents the detected facial expressions, emotions, or body language, as well as timing data (e.g., a timestamp, or a combination of a start time and an end time) for each detected non-verbal communication.
[0025] In addition to generating text from the video streams originating from each meeting participant's camera device, the video streams containing shared content (e.g., as a result of a meeting participant sharing his or her screen or sharing an application user interface) are also analyzed for the purpose of generating or extracting text. For example, a video stream that includes shared content can be processed by one or more pre-trained machine learning models to identify specific regions of interest that include collections of text. These collections of text are then further analyzed with additional pre-trained machine learning models to determine their layout or structure, enabling relevant groups of words to be extracted from the collections of text while preserving the structure and order of the relevant text.
[0026] As various video streams are processed by a media processing service, text output by various machine learning models can be associated with timing data. In particular, timing data can be generated for each detected non-verbal communication (e.g., head and hand poses, body language, facial expressions, word groups extracted from shared content, etc.) and stored with a textual description of the non-verbal communication. This timing data can be expressed relative to a start time of a meeting, but can also be expressed using some other standard measure for measuring time, including, for example, actual time of day. In this way, when an event or action (e.g., representing a non-verbal communication) is detected as recorded in a video and results in the generation of a corresponding textual description, the time at which the event or action occurred is also captured. For example, if a pose or other non-verbal communication is detected in a video stream, a machine learning model that identifies the non-verbal communication and generates a textual description of the non-verbal communication can also output timing data to indicate the time during a meeting when the non-verbal communication occurred. Thus, textual descriptions of various non-verbal communications detected by one or more pre-trained machine learning models can be further processed to arrange the text and store the text in a data structure (e.g., a table of a database) such that various text elements can be invoked, queried, searched, etc. based on the time at which the underlying action or event occurred. According to some embodiments, data generated from one or more models can be stored in fields of a database table, where separate fields are used to store a textual description of a non-verbal communication, a source (e.g., a person or meeting participant) of the detected non-verbal communication, and a time at which the non-verbal communication occurred. This improves the ability of a software-based meeting analyzer to generate accurate and complete insights about an online meeting.
[0027] In other cases, a text-based transcription (e.g., a text file) can be exported, in which each textual depiction of a non-verbal communication is stored in chronological order relative to other communications (e.g., conversations, as represented by text generated via speech-to-text translation) that occur during a meeting based on when the corresponding non-verbal communication occurred. In particular, each textual element that has been derived or extracted based on detecting or identifying a particular action or event, relative to other textual elements, is ordered or positioned relative to the time during the meeting in which the action or event occurred. Additionally, the text can be annotated or stored with metadata that reflects not only the relevant time, but also the source of the text - e.g., the name of the meeting participant associated with the event or action from which the text was derived. For some embodiments, an annotated text-based transcription can be generated, in which the text-based transcription includes all textual elements derived from various sources arranged in chronological order to reflect the time in which an event or action occurred during a meeting and from which a portion of the text was derived. Such an arrangement of text can also include text extracted from a text-based chat that occurred during an online meeting, as facilitated by a chat service or messaging service that is an integral part of the online meeting service. The output of the media processing service is a data structure of a text-based digital representation of an online meeting derived from multiple data sources. This data structure can be temporarily or more permanently stored in memory (e.g., random access memory) by writing the data structure as a text-based file to a non-volatile storage device, which can be referred to as a text-based, annotated transcription. In any case, this text-based digital representation of an online meeting is provided as input to a software-based meeting analyzer service.
[0028] In some embodiments, the timing data can also include time code data for video, which can allow for synchronizing the actual time in which an event occurred with the portion of a video file that depicts the non-verbal communication. Thus, if the meeting analyzer service provides textual depictions as a response to a query, the text-based response can also include a link to one or more relevant portions of the actual video file from which an answer to the end user was derived. In this way, an end user can verify particular information by quickly and efficiently viewing the portion of the video (e.g., a video clip) that is relevant to a particular answer or reply that has been derived based on a default or customized end user query.
[0029] According to some embodiments, the meeting analyzer service uses a text-based digital representation of an online meeting with a generative language model, such as a large language model (“LLM”), to generate an accurate and complete summary description of the online meeting. Generative language models are sophisticated artificial intelligence systems that are capable of understanding and generating human-like text. The models are trained on large data sets of text to learn grammar, semantics, and contextual relationships. When used to provide summary descriptions, and in particular to provide summary descriptions of online meetings, these models are capable of analyzing a portion of text and condensing it into a succinct and coherent summary while retaining key information. By leveraging the ability of these models to generate language, generative language models enable efficient and accurate summarization of online meetings. By providing a text-based representation (a complete and accurate representation of what happened during an online meeting) to the generative language model, the summary description generated by the generative language model is generally improved - i.e., more accurate and more complete.
[0030] The text elements representing an online meeting can be used in a variety of ways to generate an accurate summary description of the online meeting. First, according to some embodiments, various portions of the text can be selected to be included in a prompt that is provided as input to the generative language model. For example, in some cases, one or more prompt templates can be developed such that additional text extracted or selected from the text representing an online meeting can be injected into the prompt template. In some cases, the text can be formatted in the form of an annotated transcript such that the text can be selected in a variety of ways. For example, the text can be selected based on a time range, by speaker, by source, etc.
[0031] In another example, the text representing an online meeting can be provided as context or prior context. For example, any text-based input provided as input to a generative language model can generally be referred to as context, where the context provided prior to a prompt can be specifically referred to as prior context. Thus, according to some examples, the entire text representing an online meeting can be provided as prior context for any of a plurality of prompts.
[0032] Finally, in a third scenario, some portion of the text representing an online meeting can be used in a post-processing operation, for example, to confirm or verify that the output generated by a generative language model (e.g., a summary description) is accurate. For example, the text representing an answer generated by the generative language model can be subjected to a verification process that uses information from the text-based transcript to confirm or verify a particular answer before the answer is displayed or presented. In this way, having an accurate and complete text-based representation of an online meeting can help prevent a phenomenon commonly referred to as hallucination from occurring with a generative language model.
[0033] One of the several advantages of the approaches described herein is a more complete and accurate summary description of an online meeting as generated by a generative language model. This advantage is achieved by deriving or generating text from non-verbal communication and accurately extracting text from shared content. Because the resulting text is a more complete and accurate digital representation of an online meeting, the text provided as input to a generative language model, whether as a prompt or as additional context to a prompt, allows the generative language model to provide or generate a more complete and accurate summary description of the online meeting. Other aspects and advantages of the present subject matter will be readily apparent to those of ordinary skill in the art in view of the following description of several drawings.
[0034] Figure 2 FIG. 1 is a diagram illustrating how a media processing service 200 receives and processes various video streams generated during an online meeting. In this example, the online meeting has four meeting participants, each of whom is broadcasting a video stream. According to some embodiments, each video stream represents one of two types of video streams. The first type of video stream is a video stream generated using a video camera device of a client computing device executing a client software application for the online meeting service. Typically, this first type of video stream will depict a video image of the meeting participant. The second type of video stream, referred to herein as a content sharing stream, is a video stream representing content that a meeting participant is sharing with other meeting participants via the online meeting service. A content sharing stream is initiated at the meeting participant’s client computing device, not via a video camera. Instead, a content sharing stream is initiated when a meeting participant shares content by sharing an application user interface or sharing his or her screen or some portion thereof. In this example, the lines having reference numerals 202-A, 204-A, 206-A, and 208-A each represent a live video stream being transmitted from a meeting participant’s video camera to the media processing service 200 over time. The second type of video stream, referred to herein as a content sharing stream, is a video stream representing content that a meeting participant is sharing with other meeting participants via the online meeting service. A content sharing stream is initiated at the meeting participant’s client computing device, not via a video camera. Instead, a content sharing stream is initiated when a meeting participant shares content by sharing an application user interface or sharing his or her screen or some portion thereof. In this example, the lines having reference numerals 202-B and 204-B represent content sharing streams for content shared by the first and second meeting participants, respectively. For example, in this example, during the online meeting, the second participant (e.g., “Meeting Participant #2”) shares content with the other meeting participants, as indicated by the line having reference numeral 204-B. After the second meeting participant shares or presents content, the first meeting participant (e.g., Meeting Participant #1) shares content with the others, as indicated by the line having reference numeral 202-B. Figure 2
[0035] As in Figure 2 As shown in the figure, the passage of time is represented by a timeline having reference numeral 210. Thus, as shown by the timeline 210, the rightmost portion of each line representing a video stream (e.g., 202-A, 202-B, 204-A, 204-B, etc.) corresponds to time T=0, while the leftmost portion of each line represents the current time. Thus, for example, for the purposes of this example, assume that T is measured in minutes, the right portion of each line, at time T=0, represents the beginning of the online meeting, while the left portion of each line, at time T=40, represents the passage of forty minutes from the beginning of the online meeting. Thus, movement from right to left represents the passage of time.
[0036] According to some embodiments, as each video stream is received at the media processing service 200, each video stream is processed by the media processing service 200 to generate text. In particular, the audio portion of each video stream of the first type is processed using a speech-to-text algorithm or model to derive text representing spoken messages from each meeting participant. In addition, the video portion of each video stream of the first type is also processed using various pre-trained computer vision algorithms or models to detect or identify non-verbal communications made by the meeting participants, including gestures, facial expressions, emotions, and in some cases body language. The output of these computer vision algorithms or models is text that describes or otherwise represents the identified non-verbal communications. In some embodiments, each text element is associated with metadata indicating the time during the online meeting that the non-verbal communication occurred from which the text was derived, where the non-verbal communication is the action or event from which the text was derived. Similarly, metadata will be associated with each text element to indicate the source, e.g., the particular meeting participant from which the text was derived.
[0037] Each video stream of the second type, i.e., each content sharing stream, is also processed using various computer vision algorithms or models to identify and extract text and graphics. For example, a pre-trained object detection algorithm is utilized to analyze the content sharing stream to identify regions of interest or graphics that depict text. When a region of interest that depicts text is identified, a layout analysis algorithm or model is further utilized to process the text to identify the layout or structure of the text. Using the identified layout and structure, word groups are then extracted from the content. For example, an optical character recognition (OCR) algorithm can be applied to the structured collection of each identified text to extract actual word groups. Here, each word group is a collection of text extracted from the content where the ordering and grouping of individual words is preserved via analysis of the text layout and structure. The description of the media processing service 200 below describes this in more detail. Figure 8 The media processing service 200 is described in more detail below in connection with the description of
[0038] The media processing service 200 independently processes each video stream and then combines the output of the various analyses (e.g., text) to generate a complete and accurate text-based representation of the online meeting. For some embodiments, this digital representation of the online meeting is a text-based transcript 212 in which each text element inserted into the text-based transcript is chronologically positioned based on the time during the online meeting at which the particular action or event from which the text was derived occurred. Each text element can be annotated as entered in the transcript to identify the time during the online meeting at which the particular action or event from which the text was derived occurred. Similarly, the source of the text, i.e., the identity of the meeting participant from which the text was generated, can also be provided as an annotation for each text element.
[0039] As shown in Figure 2 , the text-based transcript 212 for the online meeting is provided as input to a meeting analyzer service 214. The meeting analyzer service 214 utilizes a generative language model 216 to generate a summary description of the online meeting based in part on the text-based transcript 212. The summary description for the online meeting can be derived in a variety of different ways. The operation of the meeting analyzer service 214 is described in more detail below in connection with the description of Figure 9 .
[0040] According to the example presented in Figure 2 , when the online meeting begins (e.g., at time T=0), all four meeting participants are broadcasting live video streams as represented by lines 202-A, 204-A, 206-A, and 208-A. Shortly after the online meeting begins, the second meeting participant (e.g., meeting participant #2) begins sharing a content stream as represented by the line having reference number 204-B. While the second meeting participant is sharing content, the second meeting participant is speaking, e.g., to explain, describe, or discuss the content being shared by the meeting participant. At some point, the second meeting participant asks a question or makes a statement, and the third meeting participant (e.g., meeting participant #3) conveys to all of the meeting participants that he or she agrees with the statement made by the second meeting participant, e.g., by nodding his head up and down. The detected head nod is represented in Figure 2 by the label “gesture” having reference number 216. An example of a user interface for the online meeting is shown in Figure 3 , showing the meeting participant nodding his head.
[0041] Figure 3is a diagram illustrating an example of a user interface 300 of a client software application for an online meeting service according to an embodiment of the application, where the user interface is presenting a live video stream 302 depicting a meeting participant 304 performing a first gesture (e.g., a nod gesture) 306. The user interface 300 includes two main components. The first component is a control panel 308, which provides various meeting control elements, such as, for example, a first button (e.g., “Leave Meeting”) 310, which when selected will remove the meeting participant from the online meeting, and a second button (e.g., “Share Content”) 312, which when selected will provide the meeting participant with various options for sharing content. For example, the meeting participant can choose to share content that is being presented on a display (e.g., screen sharing) or content associated with a particular application (e.g., application sharing). The control panel 308 portion of the user interface 300 includes additional control elements, such as a list of contacts 314 from which a meeting participant can select and add to the online meeting. The control panel 308 also includes a chat window 316 that provides text-based chat functionality. Finally, the control panel includes a button (e.g., “Device Settings”) that allows the meeting participant to select and configure various input and output devices (e.g., microphones, speakers, cameras, displays, etc.) that will be used with the software application.
[0042] In this example, the bounding box 320 and the annotation (e.g., “<human head>: 95%”) 322 are shown to merely convey how various computer vision algorithms or pre-trained models can operate to identify gestures made by a meeting participant. The bounding box 320 and the annotation 322 are not actually displayed as part of the user interface 300 presented to the meeting participant. Rather, in this example, the bounding box 320 represents a region of interest that has been detected by an object detection algorithm. In this case, the object detection algorithm has detected the head of the meeting participant within the region of interest identified by the bounding box 320. The annotation 322 indicates the object that has been detected - a human head - and the number (“95%”) represents a confidence score that indicates how likely the detected object is to be the one indicated by the annotation. According to some embodiments, the output of the object detection algorithm - specifically, the coordinates identifying the location of the region of interest, the class label of the identified object similar to the annotation 320, and the confidence score - will be provided as input to a second algorithm or model. Specifically, the output of the object detection algorithm or model is provided as input to a gesture detection algorithm. The gesture detection algorithm processes the portion of the video identified by the coordinates of the bounding box 320 to identify a gesture that can have been made by the meeting participant. In this example, the gesture detection algorithm or model detects that the meeting participant has performed a gesture by, for example, nodding his head, as in Figure 3The pose detection algorithm can be trained to generate text that specifically describes the pose that has been detected. This type of machine learning model can be referred to as a pose-to-text detection model. In other examples, the pose detection algorithm or model can simply identify the pose, e.g., as a category of pose. In this case, additional post-processing logic can be applied to map the detected or identified pose to a text-based description of the pose.
[0043] Referring again to Figure 2 Shortly after the third conference participant performs the nodding pose, i.e., after time T = 10, the second conference participant who is sharing content (as represented by the line with reference number 204-B) shares a slide of a presentation that includes text. In processing the shared content stream 204-B, the media processing service 200 detects the text and generates a text element to include in the final output (e.g., a text-based transcription). In this example, the text element is generated to include the text "I am sorry that I could not hear you clearly. Can you please repeat your question?" Figure 4 An example of a user interface that shares content in which text is detected is illustrated in
[0044] Figure 4 is a diagram illustrating an example of a user interface 400 for a client application for an online meeting service, where the user interface 400 is presenting a content sharing stream 402, and specifically, a slide or page of a presentation that includes text. As illustrated in Figure 4 The user interface 400 includes two main components - a content sharing panel 402, in which shared content is presented; and a control panel 408, which shows various control elements for the online meeting application.
[0045] In some embodiments, the content sharing stream is analyzed using an object detection algorithm or model. The object detection algorithm or model is trained to identify regions of interest - specifically, portions of the user interface 400 that include text being shared as part of a content sharing stream or presentation. Here, the output of the object detection algorithm is represented by the bounding box having reference numeral 404 in this example. For purposes of conveying an understanding of the computer vision algorithm or model, the bounding box 404 is illustrated here, and the bounding box 404 would not actually be present in the user interface displayed to the meeting participants. In any case, in some embodiments, the object detection algorithm or model is trained to identify text being shared via the content sharing feature of the online meeting service. More specifically, the model is trained to ignore or exclude those portions of each video frame of a shared video stream that represent portions of the user interface of the online meeting application, or the user interface of the particular application from which content is being shared. For example, if a document editor application or a slide editor and presentation application is used to share content, some portion of the actual shared content can represent the user interface of the application, rather than the actual content that the meeting participant intends to share. The object detection algorithm can be trained to identify only the content that the meeting participant intends to share, and to exclude any text or graphics that are detected as part of the actual user interface of the application used to present the content. As shown in Figure 4 the bounding box 404 has identified the location of the text being shared within the user interface 400.
[0046] As shown in Figure 4 the output of the object detection algorithm or model is provided as input to a layout analysis algorithm or model. The layout analysis algorithm or model is trained to identify the structure or layout of the text, enabling word groups to be extracted without mixing text from individual sets of text groups. Here, each word group is shown with its own bounding box to indicate the output of the layout analysis algorithm or model. Once the structure or layout of the text has been determined, an optical character recognition (OCR) algorithm can be used to recognize the characters, and ultimately extract the text. As described below and shown in conjunction with Figure 7 each text element can be added to a text-based transcript, annotated to indicate the time at which the text was presented, and the meeting participant who shared the content that included the text.
[0047] Referring again to Figure 2As the online meeting continues, shortly before time T = 30, the first meeting participant (e.g., meeting participant #1) begins a second content sharing presentation 202-B and begins discussing the shared content he or she is presenting. During the content sharing presentation, the third meeting participant (e.g., meeting participant #3) makes a gesture 220 that is detected in the video stream 206-A. In conjunction with Figure 5 Examples of detected gestures are shown and described.
[0048] Figure 5 is a diagram illustrating an example of a user interface 500 of a client application for an online meeting service in accordance with an embodiment of the application, where the user interface 500 is presenting a live video stream of a meeting participant performing a second gesture (e.g., a thumbs up gesture) 502. In this example, the object detection algorithm or model has identified a human hand within a bounding box or region of interest 504, and the gesture detection algorithm or model has identified a hand gesture, specifically, a gesture commonly referred to as a "thumbs up" gesture 502, indicating that the meeting participant has expressed agreement or given a positive response to something said during the online meeting. As with the gesture shown in Figure 3 the text description of the hand gesture 502 will be output and added to the text-based transcription for subsequent input to the meeting analyzer service.
[0049] Referring again to Figure 2 As the first meeting participant is sharing content 202-B, the first meeting participant shares a slide or page of a presentation that includes a combination of text and graphics 222 that are detected by the media processing service 200. In Figure 6 an example of such a content sharing user interface is presented in
[0050] Figure 6 is a diagram illustrating an example of a user interface 600 of a client application for an online meeting service in accordance with an embodiment of the application, where the user interface 600 is presenting a content sharing stream, and specifically, a slide or page of a presentation that includes a combination of text and graphics. In this example, the content being shared includes both text 602 and graphics 604. The text 602 is detected by the media processing service 200 and is output to the meeting analyzer service for analysis. The graphics 604 are also detected by the media processing service 200 and are output to the meeting analyzer service for analysis. Figure 4The text 602 is processed in the same way described in the text above. However, in this example, the object detection algorithm or model has already identified the graph 604. In some embodiments, the graph may be analyzed by one or more pre-trained machine learning models that generate one or more tags or category labels for the graph. The graph is then captured, tagged, and stored for subsequent retrieval. For example, the meeting analyzer may present links to the graph, allowing meeting participants to easily select a link to view the graph within the context of a specific summary description of the online meeting.
[0051] Figure 7 This diagram illustrates a portion of an annotated text-based transcription for an online meeting, generated by a media processing service according to an embodiment of the invention and including text describing nonverbal communication and word groups extracted from a presentation. According to some embodiments, the output of the media processing service is a file comprising a text-based transcription with multiple text elements. In some embodiments, each text-based element may be annotated, for example, by including various metadata elements along with the text, such as data indicating the source from which the text was derived, and data indicating the time when an action or event occurred during the online meeting, wherein the derived text is associated with said action or event. Furthermore, the text elements are inserted chronologically into the transcription and are ordered within the transcription.
[0052] For example, as in Figure 7 As shown, the text element with reference numeral 702 represents the output generated by converting recorded audio into text using a speech-to-text algorithm or model. In this example, the recorded audio is associated with a video stream from a first meeting participant (e.g., meeting participant #1). As shown in transcription 700, the information preceding the actual text includes information indicating the time the spoken message was recorded, the identifier of the meeting participant, and the name of the meeting participant.
[0053] As indicated by reference numeral 704, a text description has been input into the transcription of the detected pose. Here, again, the time and source of the detected pose are provided as annotations (e.g., metadata) in the text-based transcription. In this example, the pose-to-text detection algorithm or model has output text describing the detected pose, including the name of the participant associated with the video stream in which the pose was detected. In another example, the text element with reference numeral 710 is a text description of a hand pose detected in the video stream.
[0054] As shown with reference numbers 706 and 708, various text elements associated with the content sharing presentation have been added to the transcript. For example, the annotation or metadata with reference number 706 indicates the text identified during a screen sharing session during the online meeting. As shown in Figure 7 the time at which the text was detected. However, in some alternative embodiments, the annotation can indicate the duration of time that the presentation content or text was presented. Following the annotation or metadata, the actual extracted text from the presentation 708 is presented in the transcript.
[0055] As shown in Figure 7 in some embodiments, the output of the media processing service can be a single text-based file in which the text is arranged as a transcript. However, according to some alternative embodiments, each text element can be stored as a structured object, such as a list or array, in which the element can include a first element for the text itself, a second element for the timestamp, and a third element to identify the source of the text.
[0056] Figure 8 is a block diagram illustrating an example of the functional components of a media processing service including an online meeting service according to some embodiments. As shown in Figure 8 the media manager 802 is a component of the online meeting service that redirects the various input video streams to a queue processor 804. The queue processor 802 receives each video stream and temporarily stores the incoming data as a video file.
[0057] The media processing service 800 has two main components. The first component, referred to in Figure 8 as the video processor 806, is used to process the video originating from the video cameras of the computing devices of the meeting participants. As shown in Figure 8 the video processor 806 will process each video file in a parallel processing path. For example, each video file is processed by a speech-to-text algorithm or model 808 to convert spoken audio messages to text. At the same time, each video file is also processed by an object detection algorithm 810 to detect objects within the individual frames of the video that can be associated with non-verbal communication. For example, the object detection algorithm or model can detect the face, head, hands, and / or body of a human - i.e., the meeting participant depicted in the video. The output of the object detection algorithm or model 810 typically includes coordinates defining a region of interest or bounding box for the detected object and a class label for the detected object. The output of the object detection algorithm or model is then provided as input to downstream algorithms or models. In some cases, the particular algorithm or model invoked can depend on the class of the detected object.
[0058] For example, as shown in Figure 8As shown in the middle, two downstream machine learning algorithms or models are shown - a first model 812 for identifying or detecting a gesture, and a second model 814 for identifying or detecting an emotion. Of course, in various alternative embodiments, the functionality of the two models can be combined into a single model, or similarly, there can be additional models beyond the two depicted in the middle. Figure 8 As depicted in the middle, two models are depicted. In this example, when the object detection algorithm or model 810 detects an object such as a human head or hand, the output of the object detection algorithm or model is provided as input to the gesture detection algorithm or model 812. The gesture detection algorithm or model 812 then analyzes the region of interest defined by the bounding box in which the object was detected for identifying or detecting a gesture, which can be a hand gesture where the detected object is a hand, or a head gesture where the detected object is a human head. Of course, in various alternative embodiments, other objects and gestures can be detected. For example, in some embodiments, the detected object can be a body, and the gesture can be a shrug or some other detectable body gesture or body language. In some embodiments, when a human face is detected by the object detection algorithm or model 810, the relevant region of interest can be analyzed by a model trained to detect or identify emotions, such as the emotion detection algorithm or model 814.
[0059] According to some embodiments, the downstream models 812 and 814 can be trained to output a textual description of what has been detected. For example, in some embodiments, each model can be trained using a supervised training technique, where the training data includes a large amount of annotated video data, where the annotations are textual descriptions of what is depicted in the video. Thus, each model can be trained to identify particular gestures, facial expressions, body language, etc., and generate a textual description of the behavior or action that has been detected. However, in other embodiments, one or more of the downstream models can be trained to be categorical, which simply outputs a class label corresponding to the action or behavior that has been detected. For example, instead of generating a textual description of a thumbs up hand gesture, the output of the model can simply be a label or similar identifier indicating the particular type of gesture that was detected (e.g., a thumbs up gesture). In this case, some post-processing logic can be used to map the detected gesture to a textual description. For example, the post-processing logic can use a templated description of the source of the gesture (e.g., the name of a conference participant) as well as possibly the time at which the gesture was detected, and other information related to the context of the conference. Thus, the detected gesture can result in a detailed description using text such as "after John Doe made a statement, Jill Smith expressed agreement by giving a thumbs up gesture."
[0060] A second major component of the media processing service 800 is the content sharing processor 816. Whereas the video processor 806 processes video originating from a camera device, the content sharing processor 816 processes video files generated as a result of a meeting participant using a content sharing tool or feature of an online meeting service. The content sharing processor 816 includes an object detection algorithm or model 818 trained to identify or detect text, diagrams, graphics, pictures, and the like. In some embodiments, the object detection algorithm or model 818 is specifically trained to ignore user interfaces of applications being used to share content, whether those are content editing applications or those portions of the online meeting application’s own user interface. When the object detection algorithm 818 detects shared text, the output of the object detection algorithm or model 818 is provided as input to a layout analysis algorithm or model 820. The layout analysis algorithm or model 820 is trained to identify the structure of the text elements included in the regions of interest identified by the object detection algorithm. As an example, if the layout analysis algorithm or model detects text formatted in columns, the structure of the detected text can be used to ensure that when the text is extracted via optical character recognition (OCR) 822, the text associated with each column is kept together rather than being incorrectly mixed.
[0061] When the object detection algorithm 818 detects a diagram, graphic, or picture, the coordinates of the region of interest in which the object was detected are passed to a graphic marker 822. The graphic marker will analyze the region of interest to generate a label or tag identifying the detected object and generate an image (e.g., snippet) of the detected object. The image and associated tag(s) are then stored with metadata indicating the source and time, allowing the meeting asset to be linked to and subsequently called by the meeting analyzer service.
[0062] As shown in Figure 8 , various outputs of the video processor 806 and the content sharing processor 816, e.g., annotated text elements, are provided as input to a serializer 826, which temporarily stores the data before generating a final output in the form of an annotated text-based transcript 828. Although not shown in Figure 8 , the serializer 826 can also receive as input text that is part of a text-based chat session facilitated by the online meeting service and occurring during the online meeting. Individual text-based messages from the chat session can be inserted into the text-based transcript in temporal order with all other text elements from the various different sources. Thus, the end result of the media processing service 800 is a complete and accurate representation of what occurred during the online meeting.
[0063] Figure 9is a diagram illustrating an example of various functional components for an online meeting analyzer 900 including at least one generative language model 902, such as a large language model (LLM), for generating a summary description of an online meeting, in accordance with some embodiments. The generative language model can use a generative pre-trained transformer (GPT) and be pre-trained on a large dataset, and then fine-tuned 910 with data specific to online meetings.
[0064] As shown in Figure 9 The meeting analyzer 900 includes pre-processing logic 904 and post-processing logic 906. In addition, as indicated by dashed boxes 908 and 910, in some embodiments, the prompts or prompt templates can be derived through prompt engineering 908, and the model 902 can be fine-tuned through a supervised tuning process involving training data specific to online meetings.
[0065] Generally, a text-based transcription 828 of an online meeting can be used to generate prompts as part of a pre-processing stage. For example, the pre-processing logic 904 can include rules or instructions for extracting portions of text from the text-based transcription to generate one or more prompts, which can then be provided as input to the model 902. In some embodiments, a user interface (not shown) for the meeting analyzer service 900 provides an end user with the ability to simply select one or more graphical user interface elements (e.g., buttons) to invoke a prompt. The output from the model 902 can be processed with the post-processing logic 906 before being presented to the end user. For example, in some embodiments, portions of text from the text-based transcription 828 can be provided to the post-processing logic in order to verify the answer generated by the model for a particular prompt.
[0066] In addition to using text from a text-based transcription to generate prompts, in some embodiments, various prompts can reference text of the transcription 828 or a portion thereof, such that text from the transcription is provided as pre-prompt context for generating output by the model 902.
[0067] In addition to generating summary descriptions of online meetings, the meeting analyzer service can also provide a wide variety of other features and functionality. In particular, in some embodiments, a generative language model can be used to process a digital representation of an online meeting to identify and automatically formulate action items assigned to meeting participants or others. For example, one or more prompts can be constructed to identify action items based on a digital representation of an online meeting. One or more prompts can be constructed to identify a person who conveys particular knowledge or a topic, etc.
[0068] In contrast to existing techniques for generating summary descriptions for online meetings, the generative language model 902 is less likely to generate erroneous results because the input to the model 902 is more accurate and complete as a result of the processing done by the media processing service. As an example, because the digital representation of the online meeting includes text descriptions of various non-verbal communications, as well as timing data for indicating when those non-verbal communications occurred, the meeting analyzer service is able to infer answers to various questions that might otherwise not be possible. As an example, consider the following scenario: in which a first meeting participant shares content via an online meeting collaboration tool (e.g., a screen or application sharing feature), and a conversation occurs regarding the topic presented via the shared content. The first meeting participant can ask a question related to the content being shared, and one or more other meeting participants can communicate agreement or disagreement with the presenter (e.g., the first meeting participant) by making non-verbal communications such as hand or head gestures. According to embodiments of the invention, because the verbal and non-verbal communications are captured and represented in the digital representation of the online meeting, the meeting analyzer service is able to generate accurate answers to various questions based on the text descriptions and timestamps associated with the non-verbal communications. For example, when John presents a 2024 budget or financial plan, a person who was unable to attend the meeting can submit a query asking whether Jane Doe (meeting participant #3) agreed with John Baily (meeting participant #1). If Jane Doe is detected to make a thumbs-up gesture while John Baily is presenting the budget or financial plan during the meeting, the text description of the non-verbal communication would enable the meeting analysis service to generate an accurate answer to the query.
[0069] Machine and Software Architecture Figure 10 FIG. 1 illustrates an example of a software architecture 1002, which can be installed on any of various computing devices to implement methods according to those described herein. Figure 10 This is merely a non-limiting example of a software architecture, and it will be appreciated that many other architectures can be implemented to facilitate the functionality described herein. In various embodiments, the software architecture 1002 is implemented by hardware such as that described in FIG. 10. Figure 11 The machine 1100 of FIG. 10 includes a processor 1110, a memory 1130, and an input / output (I / O) component 1150. In this example architecture, the software architecture 1002 can be conceptualized as a stack of layers, each of which provides a particular functionality. For example, the software architecture 1002 includes layers such as an operating system 1004, libraries 1006, frameworks 1008, and applications 1010. Operationally, according to some embodiments, the applications 1010 invoke API calls 1012 through the software stack, and receive messages 1014 in response to the API calls 1012.
[0070] In various implementations, the operating system 1004 manages hardware resources and provides common services. The operating system 1004 includes, for example, a kernel 1020, services 1022, and drivers 1024. According to some embodiments, the kernel 1020 acts as an abstraction layer between the hardware and the other software layers. For example, the kernel 1020 provides memory management, processor management (e.g., scheduling), component management, networking, and security settings, among other functionality. The services 1022 enable other software layers to perform tasks such as
[0071] In some embodiments, the libraries 1006 provide a low-level common infrastructure used by the applications 1010. The libraries 1006 can include system libraries 1030 (e.g., C standard library) that can provide functions such as memory allocation functions, string manipulation functions, mathematical functions, and the like. In addition, the libraries 1006 can include API libraries 1032 such as media libraries (e.g., libraries to support presentation and manipulation of various media formats such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), graphics libraries (e.g., an OpenGL framework used to render in two and three dimensions in a two-dimensional (2D) and three-dimensional (3D) graphics context on a display), database libraries (e.g., SQLite to provide various relational database functions), web libraries (e.g., WebKit to provide web browsing functionality), and the like. The libraries 1006 also include a wide variety of other libraries 1034 to provide many other APIs to the applications 1010.
[0072] According to some embodiments, the frameworks 1008 provide a high-level common infrastructure that can be utilized by the applications 1010. For example, the frameworks 1008 provide various GUI functionalities, high-level resource management, high-level location services, and so forth. The frameworks 1008 can provide a broad spectrum of other APIs that can be utilized by the applications 1010, some of which can be specific to a particular operating system 1004 or platform.
[0073] In example embodiments, application 1010 includes a home application 1050, a contacts application 1052, a browser application 1054, a book reader application 1056, a location application 1058, a media application 1060, a messaging application 1062, a game application 1064, and a wide variety of other applications, such as a third-party application 1066. According to some embodiments, application 1010 is a program that performs functions defined in a program. One or more applications in application 1010 can be created using various programming languages, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C or assembly language). In a particular example, third-party application 1066 (e.g., an entity other than a platform-specific vendor using Android) TM or iOS TM Applications developed using a Software Development Kit (SDK) can run on mobile operating systems such as iOS. TM ANDROID TM Mobile software running on (Windows® Phone or another mobile operating system). In this example, third-party application 1066 is able to invoke API call 1012 provided by operating system 1004 to facilitate the functions described herein.
[0074] Figure 11 The illustration shows a graphical representation of a machine 1100 in the form of a computer system according to an example embodiment, within which a set of instructions can be executed to cause the machine to perform any or more of the methods discussed herein. Specifically, Figure 11FIG. 11 illustrates a diagrammatic representation of a machine in the example form of a computer system 1100 within which instructions 1116 (e.g., software, a program, an application, an applet, an app, or other executable code) for causing the machine 1100 to perform any one or more of the methodologies discussed herein can be executed. For example, the instructions 1116 can cause the machine 1100 to execute any one of the methods or algorithmic techniques described herein. Additionally or alternatively, the instructions 1116 can implement any one of the systems described herein. The instructions 916 transform the general, non-programmed machine 1100 into a particular machine 1100 programmed to carry out the described and illustrated functions in the manner described. In alternative embodiments, the machine 1100 operates as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1100 can operate in the capacity of a server machine or a client machine in server-client network environments, or as a peer machine in peer-to-peer (or distributed) network environments. The machine 1100 can comprise, but not be limited to, a server computer, a client computer, PC, a tablet computer, a laptop computer, a netbook, a set-top box (STB), a PDA, an entertainment media system, a cellular telephone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing the instructions 1116, sequentially or otherwise, that specify actions to be taken by machine 1100. Further, while only a single machine 1100 is illustrated, the term “machine” shall also be taken to include a collection of machines 1100 that individually or jointly execute the instructions 916 to perform any one or more of the methodologies discussed herein.
[0075] The machine 1100 can include processors 1110, memory 1130, and I / O components 1150, which can be configured to communicate with one another by way of a bus 1102, In an example embodiment, the processors 1110 (e.g., a Central Processing Unit (CPU), a Reduced Instruction Set Computer (RISC) processor, a Complex Instruction Set Computer (CISC) processor, a Graphics Processing Unit (GPU), a Digital Signal Processor (DSP), an ASIC, a Radio-Frequency Integrated Circuit (RFIC), another processor, or any suitable combination thereof) can include, for example, a processor 1112 and a processor 1114 that can execute the instructions 1116. The term “processor” is intended to include multiple processors 1110 that can be present in a Figure 11Multiple processors 1110 are shown, but the machine 1100 can include a single processor with single cores, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0076] The storage unit 1136 can also include a removable memory that can be loaded into the storage unit 1136. The storage unit 1136 can include virtual storage. The storage unit 1136 can include memory that is shared with another component of the machine 1100, such as storage that is shared with the main memory 1132.
[0077] The I / O components 1150 can include a wide variety of components to receive input, provide output, produce output, transmit information, exchange information, capture measurements, and so on. The specific I / O components 1150 that are included in the machine 1100 will depend on the type and Figure 11 configuration of the machine 1100. For example, portable machines such as mobile phones will likely include a touch input device or other such input mechanisms, while a headless server machine will likely not include such a touch input device. It will be appreciated that the I / O components 1150 can include many other components that are not shown in FIG. 1. The I / O components 1150 are grouped according to the functionality provided in the figure, but this does not require or imply that the components are physically grouped or even that they are on the same machine. In various example embodiments, the I / O components 1150 can include output components 1152 and input components 1154. The output components 1152 can include visual components (e.g., a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibratory motor, resistance mechanisms), other signal generators, and so forth. The input components 1154 can include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or another pointing instrument), tactile input components (e.g., a physical button, a touch screen that provides location and / or force of touches or touch gestures, or other tactile input components), audio input components (e.g., a microphone), and the like.
[0078] In further example embodiments, the I / O components 1150 can include biometric components 1156, motion components 1158, environmental components 1160, or position components 1162, among a myriad of other components. For example, biometric components 1156 can include components to detect expressions (e.g., hand expressions, facial expressions, vocal expressions, body gestures, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, perspiration, or brain waves), identify a person (e.g., voice identification, retinal identification, facial identification, fingerprint identification, or electroencephalogram-based identification), and the like. The motion components 1158 can include acceleration sensor components (e.g., accelerometer), gravitation sensor components, rotation sensor components (e.g., gyroscope), and so forth. The environmental components 1160 can include, for example, illumination sensor components (e.g., photometer), temperature sensor components (e.g., one or more thermometers that detect ambient temperature), humidity sensor components, pressure sensor components (e.g., barometer), acoustic sensor components (e.g., one or more microphones that detect background noise), proximity sensor components (e.g., infrared sensors that detect nearby objects), gas sensors (e.g., gas detection sensors to detect concentrations of hazardous gases for safety or to measure pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to a surrounding physical environment. The position components 1162 can include location sensor components (e.g., a GPS receiver component), altitude sensor components (e.g., altimeters or barometers that detect air pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), and the like.
[0079] Communication can be implemented using a wide variety of technologies. The I / O components 1150 can include communication components 1164 operable to couple the machine 1100
[0080] Moreover, the communication components 1164 can detect identifiers or include components operable to detect identifiers (e.g., NFC smart tag detection components, optical reader components (e.g., an optical sensor to detect one- dimensional bar codes such as Universal Product Code (UPC) bar code, multi-dimensional bar codes such as Quick Response (QR) code, Aztec code, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar code, and other optical codes), or acoustic detection components (e.g., microphones to identify tagged audio signals). In addition, a variety of information can be derived via the communication components 1164, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, location via detecting an NFC beacon signal that can indicate a particular location, and so forth.
[0081] Executable instructions and machine-storage media The various memories (i.e., 1130, 1132, 1134, and / or memory of the processor(s) 1110) and / or storage units 1136 can store one or more sets of instructions and data structures (e.g., software) embodying or utilized by any one or more of the methodologies or functions described herein. These instructions (e.g., instructions 1116), when executed by the processor(s) 1110, cause various operations to implement the disclosed embodiments.
[0082] As used herein, the terms "machine-storage medium," "device-storage medium," "computer-storage medium," and "device-storage media" mean the same thing and can be used interchangeably in this disclosure. The terms refer to a single or multiple storage devices and / or media (e.g., a centralized or distributed database, and / or associated caches and servers) that store executable instructions and / or data structures. Thus, the term should be understood to include, but not be limited to, tangible storage devices and media such as solid-state memories, and optical and magnetic media, including memory internal or external to processors. Specific examples of machine-storage media, computer-storage media, and / or device-storage media include non-volatile memory, including by way of example semiconductor memory devices, e.g., erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGA, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The term "machine-storage media," "computer-storage media," and "device-storage media" specifically exclude carrier waves, modulated data signals, and other such media, at least some of which are covered under the term "signal medium" as discussed below.
[0083] Transmission media In various example embodiments, one or more portions of network 980 can be a self-organizing network, intranet, extranet, VPN, local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless wide area network (WW AN), metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the PSTN, ordinary old telephone service (POTS) network, cellular telephone network, wireless network, a fiber optic-based network, or other communication network, another type of network, or a combination of two or more such networks. For example, network 1180 or a portion of network 1180 can include a wireless or cellular network, and coupling 1182 can be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or another type of cellular or wireless coupling. In this example, coupling 1182 can implement any of a variety of types of data transfer technology, such as Single Carrier Radio Transmission Technology (lxRTT), Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data rates for GSM Evolution (EDGE) technology, third Generation Partnership Project (3GPP) including 3G, fourth generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Worldwide Interoperability for Microwave Access (Wi MAX), Long Term Evolution (LTE) standard, others defined by various standards setting organizations, other long range protocols, or other data transfer technology.
[0084] The instructions 1116 can be transmitted or received over the network 1180 using a transmission medium via a network interface device and utilizing any one of a number of well-known transfer protocols (e.g., HTTP). Similarly, the instructions 1116 can be transmitted or received using a transmission medium via the coupling 1172 (e.g., a peer-to-peer coupling). The term “transmission medium” and “signal medium” mean the same thing and can be used interchangeably in this disclosure. The terms “transmission medium” and “signal medium” shall be taken to include any intangible medium that is capable of storing, encoding, or carrying the instructions 1116 for execution by the machine 1100, and includes digital or analog communications signals or other intangible media to facilitate communication of such software. Hence, the terms “transmission medium” and “signal medium” shall be taken to include any form of a modulated data signal, carrier wave, and so forth. The term “modulated data signal” means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal.
[0085] Computer-readable medium The terms “machine-readable medium,” “computer-readable medium,” and “device-readable medium” mean the same thing and can be used interchangeably in this disclosure. The terms are defined to include both machine-storage media and transmission media. Thus, the terms include both storage devices / media and carrier waves / modulated data signals.
Claims
1. A system for deriving a digital representation of an online meeting using contextual data inferred from nonverbal communication, the system comprising: Processor (1110); as well as A memory storage device (1130) stores instructions (1116) that, when executed by the processor, cause the system to perform operations including: The first video stream (202-A) is received over the network from the client computing devices of the first conference participants, the first video stream originating from a video camera; One or more object detection algorithms (810) are used to process the first video stream to detect one or more regions of interest, each detected region of interest depicting the first conference participant or a portion of the first conference participant; For each detected region of interest, the pose recognition algorithm (812) is applied to the region of interest to detect the pose and generate a text description of the detected pose and a timestamp indicating the time when the detected pose occurred during the online meeting; The text description of the detected gesture and the timestamp indicating the time when the detected gesture occurred during the online meeting are stored. as well as In response to a query from an end user of the conference analyzer service (900), a response to the query is provided, wherein the response to the query is determined in part based on the text description of the detected gesture and the timestamp indicating the time when the detected gesture occurred during the online conference.
2. The system according to claim 1, wherein, The memory storage device stores additional instructions that, when executed by the processor, cause the system to perform additional operations including: An emotion detection algorithm (814) is used to process the detected region of interest to: i) detect emotions based on the facial expressions of the first meeting participant, and ii) generate a text description of the detected emotions and a timestamp for the detected emotions, the timestamp indicating the time when the facial expressions were made during the online meeting; as well as Store the text description of the detected emotion and the timestamp for the detected emotion; The response to the query is determined in part based on the text description of the detected emotion and the timestamp of the detected emotion.
3. The system according to claim 1, wherein, The memory storage device stores additional instructions that, when executed by the processor, cause the system to perform additional operations including: The second video stream (206) is received via the network from the client computing device of the second meeting participant, the second video stream representing content shared by the second meeting participant with one or more other meeting participants via the online meeting; One or more object detection algorithms are used to process the second video stream to detect one or more regions of interest, each detected region of interest depicting a set of text or a graphic. as well as For each detected region of interest in the set of text being described: One or more layout analysis algorithms (820) are used to process the set of text to identify the structure of the set of text; Based on the structure of the identified set of texts, extract one or more word groups from the set of texts; as well as The one or more word groups are stored and a timestamp (702) is stored with each word group, the timestamp indicating the time when the collection of text from which the word groups are extracted is shared during the online meeting; The response to the query is determined in part based on the word group and the corresponding timestamp for the word group.
4. The system according to claim 3, wherein, The memory storage device stores additional instructions that, when executed by the processor, cause the system to perform additional operations including: For each detected region of interest in the depiction pattern (604): Use a content classification algorithm (822) to generate at least one topic tag for the graphic; Generate an image of the graphic; as well as The image of the graphic is stored together with the at least one hashtag for later retrieval and presentation to conference participants.
5. The system according to claim 3, wherein, The memory storage device stores additional instructions that, when executed by the processor, cause the system to perform additional operations including: For each detected pose, based on the timestamp associated with the detected pose, the text description of the detected pose is inserted into the text-based transcription (700) for the online meeting at a position relative to other texts; as well as The text description of the detected posture is annotated to include: i) the timestamp of the detected posture, and ii) information identifying the first meeting participant.
6. The system according to claim 5, wherein, Adding the one or more word groups to the text-based transcription for the online meeting includes: For each of the one or more word groups extracted from the set of texts, based on the time during which the set of texts from which the word groups were extracted was shared by the second meeting participant during the online meeting, the word group is inserted into the text-based transcription for the online meeting in a position relative to other texts; and The word group is annotated to include: i) the timestamp for the word group, and ii) information identifying the second meeting participant.
7. The system according to claim 1, wherein, Depicting at least one region of interest of the first meeting participant; depicting the head of the first meeting participant or the hands of the first meeting participant; and One or more pose recognition algorithms (812) are used to process each region of interest to detect poses, including hand poses or head poses.
8. The system according to claim 1, wherein, The response to the query is determined by constructing a text-based cue word (908) and providing the text-based cue word as input to a generative language model (902). The text-based cue word includes context and instructions, wherein the text description of the detected gesture and the timestamp indicating the time when the detected gesture occurred during the online meeting are included in the context, and the instructions are determined based on the query.
9. The system according to claim 3, wherein, The response to the query is determined by constructing a text-based prompt (908) and providing the text-based prompt as input to a generative language model (902), the text-based prompt including context and instruction, wherein the one or more word groups and their corresponding timestamps are included in the context, and the instruction is determined based on the query.
10. A computer-implemented method, comprising: The first video stream (202-A) is received over the network from the client computing devices of the first conference participants and originates from a video camera; The first video stream is processed using one or more object detection algorithms to detect one or more regions of interest, each detected region of interest depicting a first conference participant or a portion thereof; For each detected region of interest, the pose recognition algorithm (810) is applied to the region of interest to detect the pose and generate a text description of the detected pose and a timestamp indicating the time when the detected pose occurred during the online meeting; The text description of the detected gesture and the timestamp indicating the time when the detected gesture occurred during the online meeting are stored. as well as In response to a query from an end user of the conference analyzer service, a response to the query is provided, wherein the response to the query is determined in part based on the text description of the detected gesture and the timestamp indicating the time when the detected gesture occurred during the online conference.
11. The computer-implemented method according to claim 10, further comprising: An emotion detection algorithm is used to process the detected region of interest to: i) detect emotions based on the facial expressions of the first meeting participant, and ii) generate a text description of the detected emotions and a timestamp for the detected emotions, the timestamp indicating the time when the facial expressions were made during the online meeting; as well as Store the text description of the detected emotion and the timestamp for the detected emotion; The response to the query is determined in part based on the text description of the detected emotion and the timestamp of the detected emotion.
12. The computer-implemented method according to claim 10, further comprising: The second video stream is received via the network from the client computing device of the second meeting participant, the second video stream representing content shared by the second meeting participant with one or more other meeting participants via the online meeting; One or more object detection algorithms are used to process the second video stream to detect one or more regions of interest, each detected region of interest depicting a set of text or a graphic. as well as For each detected region of interest in the set of text being described: One or more layout analysis algorithms are used to process the set of text to identify the structure of the set of text; Based on the structure of the identified set of texts, extract one or more word groups from the set of texts; as well as The one or more word groups are stored, and a timestamp is stored with each word group, the timestamp indicating the time when the collection of text from which the word groups were extracted is shared during the online meeting; The response to the query is determined in part based on the word group and the corresponding timestamp for the word group.
13. The computer-implemented method according to claim 12, further comprising: For each detected region of interest in the plotted image: Use a content classification algorithm to generate at least one topic tag for the graphic; Generate an image of the graphic; as well as The image of the graphic is stored together with the at least one hashtag for later retrieval and presentation to conference participants.
14. The computer-implemented method according to claim 12, further comprising: For each detected pose, based on the timestamp associated with the detected pose, the textual description of the detected pose is inserted into the text-based transcription for the online meeting at a position relative to other texts; as well as The text description of the detected posture is annotated to include: i) the timestamp of the detected posture, and ii) information identifying the first meeting participant.
15. The computer-implemented method according to claim 12, further comprising: For each of the one or more word groups extracted from the set of texts, based on the time during which the set of texts from which the word groups were extracted was shared by the second meeting participants during the online meeting, the word group is inserted into the text-based transcription for the online meeting in a position relative to other texts; as well as The word group is annotated to include: i) the timestamp for the word group; and ii) information identifying the participants in the second meeting.
16. The computer-implemented method according to claim 10, wherein, Depicting at least one region of interest of the first meeting participant; depicting the head of the first meeting participant or the hands of the first meeting participant; and One or more pose recognition algorithms are used to process each region of interest to detect poses, including hand poses or head poses.
17. The computer-implemented method according to claim 10, wherein, The response to the query is determined by constructing text-based cue words and providing the text-based cue words as input to a generative language model. The text-based cue words include context and instructions, wherein the text description of the detected gesture and the timestamp indicating the time when the detected gesture occurred during the online meeting are included in the context, and the instructions are determined based on the query.
18. The computer-implemented method according to claim 12, wherein, The response to the query is determined by constructing text-based prompts and providing the text-based prompts as input to a generative language model. The text-based prompts include context and instructions, wherein the one or more word groups and their corresponding timestamps are included in the context, and the instructions are determined based on the query.
19. A system for deriving a digital representation of an online meeting using contextual data inferred from nonverbal communication, the system comprising: A unit for receiving a first video stream from a client computing device of a first conference participant via a network, the first video stream originating from a video camera; A unit for processing the first video stream using one or more object detection algorithms to detect one or more regions of interest, each detected region of interest depicting the first conference participant or a portion of the first conference participant; For each detected region of interest, a unit is provided for applying the pose recognition algorithm to the region of interest to detect the pose and generating a text description of the detected pose and a timestamp indicating the time when the detected pose occurred during the online meeting; A unit for storing the text description of the detected gesture and the timestamp indicating the time when the detected gesture occurred during the online meeting; as well as A unit for providing a response to a query from an end user of a conference analyzer service, wherein the response to the query is determined in part based on the text description of the detected gesture and the timestamp indicating the time when the detected gesture occurred during the online conference.
20. The system of claim 19, further comprising: A unit for receiving a second video stream from a client computing device of a second conference participant via the network, the second video stream representing content shared by the second conference participant with one or more other conference participants via the online conference; A unit for processing the second video stream using one or more object detection algorithms to detect one or more regions of interest, each detected region of interest depicting a set of text or a graphic; as well as For each detected region of interest in the set of text being described: A unit for processing the set of text using one or more layout analysis algorithms to identify the structure of the set of text; Units for extracting one or more word groups from a set of texts based on the identified structure of the set of texts; as well as A unit for storing the one or more word groups and storing a timestamp along with each word group, the timestamp indicating the time when the collection of text from which the word groups were extracted was shared during the online meeting; The response to the query is determined in part based on the word group and the corresponding timestamp for the word group.