Voting issues for teleconference discussions

By using machine learning models to identify and automatically handle voting issues in conference calls, this technology solves the problems of discussion interruption and resource consumption caused by voting issues in existing technologies, and achieves a more efficient meeting process.

CN115735357BActive Publication Date: 2026-01-20GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180046380.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-24
Filing Date
2021-06-29
Publication Date
2026-01-20
Estimated Expiration
2041-06-29

AI Technical Summary

Technical Problem

Existing teleconference systems are prone to interrupting the discussion process when identifying and handling voting issues, increasing the length of the discussion and consuming system resources, resulting in reduced overall efficiency and increased latency.

Method used

Machine learning models are used to identify spoken phrases in conference calls, automatically identify and process voting questions, reduce interruptions, and automatically fill in answers through a graphical user interface, thereby reducing system resource consumption.

Benefits of technology

It improved the efficiency of teleconferences, reduced discussion length, and decreased system resource consumption and overall latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115735357B_ABST
    Figure CN115735357B_ABST
Patent Text Reader

Abstract

Systems and methods are provided for identifying polling questions from a conference call discussion. One or more text strings are identified, the one or more text strings comprising a textual form of one or more spoken phrases provided by one or more participants of a conference call. The one or more text strings are provided as input to a trained machine learning model. One or more outputs are obtained from the trained machine learning model. A spoken phrase of the one or more spoken phrases provided by the one or more participants is extracted from the one or more outputs that comprises a confidence level that a question associated with a poll during the conference call. In response to a determination that the confidence level satisfies a confidence criterion, the spoken phrase is designated as a polling question presented during the conference call.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Aspects and implementations of the present disclosure relate to voting questions for teleconference discussions. BACKGROUND

[0002] Video- or audio-based teleconference discussions can be conducted among multiple participants via a conferencing platform. The conferencing platform includes tools that allow multiple client devices to connect over a network and share audio data (e.g., user speech recorded by microphones of the client devices) and / or video data (e.g., video captured by cameras of the client devices or from screen images of the client devices) of one another for effective communication. The conferencing platform can also include tools that allow participants of a teleconference to pose questions to other participants (e.g., via a conferencing platform user interface (UI)) during a teleconference discussion to solicit responses (referred to as votes). The conferencing platform can collect the responses provided by the other participants and generate vote results. SUMMARY

[0003] The following summary is a simplified summary of the disclosure to present some concepts of the disclosure in a simplified form, to provide a basic understanding of certain aspects of the disclosure. This summary is not an extensive overview of the disclosure. It is neither intended to identify important or key elements of the disclosure nor to delineate any scope of any particular implementation of the disclosure or any scope of any claims. Its sole purpose is to present some concepts of the disclosure in a simplified form as a prelude to the more detailed description that is presented later.

[0004] In some implementations, a system and method are disclosed for designating a spoken phrase presented during a teleconference as a voting question. In one implementation, one or more text strings are identified, the one or more text strings comprising a textual form of one or more spoken phrases provided by one or more participants of a teleconference. The one or more text strings are provided as input to a trained machine learning model. One or more outputs are obtained from the trained machine learning model. A confidence level that a spoken phrase of the one or more spoken phrases provided by the one or more participants comprises a question associated with a vote during the teleconference is extracted from the one or more outputs. In response to a determination that the confidence level satisfies a confidence criterion, the spoken phrase is designated as a voting question presented during the teleconference.

[0005] In some implementations, a system and method for training a machine learning model to determine whether a spoken phrase provided by a participant of a teleconference includes a polling question is disclosed. In one implementation, training data for a machine learning model is generated. Generating the training data includes generating training inputs that include data corresponding to phrases provided by users of a platform. Generating the training data also includes generating target outputs for the training inputs, where the target outputs include an indication of whether the phrases correspond to questions that were previously used to poll one or more additional users of the platform. The training data is provided to train the machine learning model on (i) a set of training inputs that includes the training inputs and (ii) a set of target outputs that includes the target outputs.

[0006] In some implementations, a method and system for populating answers to polling questions based on initial responses is disclosed. In one implementation, the method includes providing a graphical user interface (GUI) to a set of participants of a teleconference. The GUI includes a first portion that presents a polling question to be answered by the set of participants. The method also includes, in response to detecting that a first participant of the set of participants has provided a spoken answer to the polling question, modifying the GUI provided to a second participant of the set of participants to present a textual form of the spoken answer provided by the first participant in a second portion. The second portion also includes a GUI element that allows the second participant to provide a non-spoken answer to the polling question by selecting the GUI element associated with the spoken answer provided by the first participant. BRIEF DESCRIPTION OF DRAWINGS

[0007] The aspects and implementations of the present disclosure will become more fully understood from the detailed description given herein and the accompanying drawings, which are meant to be illustrative only and thus are not intended to limit the present disclosure to a particular aspect or implementation, but merely to explain and understand.

[0008] Figure 1 FIGURE 1 illustrates an example system architecture in accordance with implementations of the present disclosure.

[0009] Figure 2 is a block diagram illustrating a conference platform and one example polling engine of a conference platform in accordance with implementations of the present disclosure.

[0010] Figure 3A illustrates identifying a spoken phrase provided during a teleconference as a polling question in accordance with implementations of the present disclosure.

[0011] Figures 3B-3D illustrates using a spoken phrase identified as a polling question to poll participants of a teleconference in accordance with implementations of the present disclosure.

[0012] Figure 4A- B illustrates recognizing a spoken phrase included in a teleconference transcript as a polling question, in accordance with an embodiment of the disclosure.

[0013] Figure 5 is a block diagram illustrating a conferencing platform and another example polling engine of a conferencing platform, in accordance with an embodiment of the disclosure.

[0014] Figure 6A illustrates an example user interface for a video teleconference between multiple participants via a conferencing platform, in accordance with an embodiment of the disclosure.

[0015] Figure 6B illustrates a sample UI for a participant responding to a polling question during a teleconference, in accordance with an embodiment of the disclosure.

[0016] Figures 6C-6D illustrates a sample UI for populating an answer to a polling question based on a prior spoken answer provided by another participant for other participants of a teleconference, in accordance with an embodiment of the disclosure.

[0017] Figure 7 depicts a flowchart of a method for designating a spoken phrase provided during a teleconference as a polling question, in accordance with an embodiment of the disclosure.

[0018] Figure 8 depicts a flowchart of a method for training a machine learning model to recognize a spoken phrase provided during a teleconference corresponding to a polling question, in accordance with an embodiment of the disclosure.

[0019] Figure 9 depicts a flowchart of a method for populating an answer to a polling question based on an initial response, in accordance with an embodiment of the disclosure.

[0020] Figure 10 is a block diagram illustrating an example computer system, in accordance with an embodiment of the disclosure. DETAILED DESCRIPTION

[0021] Aspects of the present disclosure relate to polling questions for teleconference. In particular, aspects of the present disclosure relate to identifying polling questions from teleconference discussions. Additional or alternative aspects of the present disclosure relate to populating answers to polling questions based on initial responses provided by participants of a teleconference. A conferencing platform can enable video or audio based teleconference discussions between a plurality of participants via respective client devices connected over a network, and share audio data (e.g., user speech recorded by microphones of the client devices) and / or video data (e.g., video captured by cameras of the client devices) with each other during the teleconference. In some cases, the conferencing platform can enable a large number of client devices (e.g., up to one hundred or more client devices) to connect through the teleconference.

[0022] Participants of a teleconference can want to ask questions to other participants of the teleconference to solicit responses from the other participants (referred to as polling). A participant can provide a polling question during the teleconference, and the conferencing platform can collect responses provided by other participants of the teleconference. Some existing conferencing platforms can provide a user interface (UI) for each client device connected to the teleconference, where the UI displays video data and / or audio data shared over the network, and can also display messages exchanged between participants during the teleconference. A participant who wants to ask a polling question can ask the polling question to other participants by causing the client devices associated with the other participants to display a message based on the polling question. For example, the participant can select a polling option in the UI of the conferencing platform and type the polling question in a designated area of the conferencing platform UI at the client device. In response to receiving the polling question, the conferencing platform can generate a message based on the polling question and cause the message to be displayed via the conferencing platform UI at the client devices associated with the other participants.

[0023] In some systems, a teleconference participant who wants to pose a voting question to other participants can consider the voting question prior to the teleconference and pose the voting question via the meeting platform UI in the particular instance of the discussion. However, introducing a predetermined voting question during the teleconference interrupts the organization or natural flow of the teleconference discussion and increases the length of the discussion. In some cases, a teleconference participant can consider a voting question based on the discussion between the teleconference participants. To use the voting question to poll other participants, the participant can ask the other participants to pause the discussion, provide the voting question through the meeting platform UI on the client device, and cause the voting question to be displayed to the other participants, as previously described. However, the process of providing the voting question through the meeting platform UI and causing the voting question to be displayed to the other participants can take a significant amount of time (e.g., several minutes). If multiple voting questions are posed to the participants of the teleconference, the length of the teleconference discussion is significantly increased. Furthermore, the organization and / or natural flow of the teleconference discussion is interrupted and the participants are unable to effectively discuss each topic of the teleconference. By increasing the length of the teleconference discussion, additional system resources are used to facilitate connections between the client devices associated with each participant. Thus, fewer system resources are available for other processes to use on the client devices and computing devices hosting the meeting platform, thereby decreasing overall efficiency and increasing overall latency of the client devices and computing devices.

[0024] Some implementations of the present disclosure address the above and other deficiencies by providing a machine learning model that identifies a polling question from a teleconference discussion. The machine learning model can be trained based on phrases previously used by users of a polling platform (e.g., a meeting platform, a collaboration platform, etc.). During a teleconference, the meeting platform can generate audio data associated with spoken phrases provided by participants of the teleconference. In some embodiments, the meeting platform can generate one or more text strings that include the provided spoken phrases and provide the text strings as input to the trained machine learning model. In other or similar embodiments, the audio data itself (i.e., the recording) can be used as input to the trained machine learning model. The machine learning model can be trained to determine whether the text string includes a spoken phrase that represents a question for the participants of the teleconference to vote on. In response to receiving the text string as input, the machine learning model can provide a confidence level as output that indicates a likelihood that the text string includes a spoken phrase that represents a polling question. The meeting platform can determine whether to designate the spoken phrase of the text string as a polling question (e.g., based on a confidence threshold). In response to determining to designate the spoken phrase as a question for polling, the meeting platform can display a message asking the participant whether they would like to pose the question to other participants of the teleconference. The message can be displayed via a UI of a client device associated with the participant that provided the spoken phrase. In response to receiving an indication that the participant would like to pose the question to other participants of the teleconference (e.g., if the participant selects a designated checkbox), the meeting platform can update the meeting platform UI on client devices associated with other participants of the teleconference to include a message associated with the polling question.

[0025] The machine learning model can also be used to identify a polling question and provide a response after the teleconference is completed. For example, the meeting platform can generate a transcript of the teleconference after the teleconference is completed. The meeting platform can identify one or more text strings that include spoken phrases provided by participants of the teleconference based on the generated transcript and provide the text strings as input to the machine learning model. The machine learning model can provide a confidence level as output that indicates a likelihood that the text string includes a spoken phrase that represents a question for polling. In response to designating the spoken phrase as a question for polling, the meeting platform can identify one or more text strings that include spoken phrases that represent a response to the polling question based on the generated transcript. The meeting platform can generate a polling result based on the identified polling question and the identified response and provide the generated polling result to participants of the teleconference (e.g., an organizer of the teleconference).

[0026] Some aspects of the present disclosure identify voting questions based on verbal phrases provided during a teleconference discussion. During a teleconference, a meeting platform can identify a voting question posed by a participant and can automatically display a message asking the participant providing the verbal phrase whether the participant would like to pose the voting question to other participants. Thus, when a participant provides a voting question through a meeting platform UI, a participant who would like to pose a voting question to other participants does not interrupt the organization or natural flow of the teleconference by asking other participants to pause the discussion. Moreover, as described above, the verbal phrases including the voting questions and responses can be identified from a transcript generated after the teleconference. In this way, the meeting platform can identify voting questions and responses provided during a teleconference without requiring participants to interrupt the discussion to pose the voting questions and solicit responses. Thus, participants of a teleconference can effectively conduct a teleconference discussion with a reduced number of interruptions, thereby reducing the length of the teleconference discussion. As a result of reducing the length of the teleconference discussion, the amount of system resources used to facilitate connections between client devices associated with each participant is reduced. Thus, more system resources are available at the client devices and other computing devices for other processes, thereby improving overall efficiency and reducing overall latency.

[0027] In some additional or alternative systems, each teleconference participant who would like to provide an answer to a voting question can consider an answer and provide the answer via the meeting platform UI upon presentation of the voting question. However, a teleconference can include a large number of participants (e.g., tens, hundreds, thousands of participants). Each participant can spend a significant amount of time (e.g., several minutes) to consider an answer and provide the answer through the meeting platform UI. In some cases, this can interrupt the organization or natural flow of the teleconference discussion and can increase the length of the teleconference discussion. If multiple voting questions are posed to the participants of the teleconference, the length of the teleconference discussion can significantly increase. Moreover, the organization and / or natural flow of the teleconference discussion can be interrupted and participants can not effectively discuss each topic of the teleconference. By increasing the length of the teleconference discussion, additional system resources can be utilized to facilitate connections between client devices associated with each participant. Thus, less system resources are available at the client devices and computing devices hosting the meeting platform for other processes, thereby reducing overall efficiency and increasing overall latency of the client devices and computing devices.

[0028] Some additional or alternative implementations of the present disclosure address the above and other deficiencies by providing systems and methods for populating answers to polling questions based on initial responses from participants of a teleconference. A teleconference platform can provide a conference UI to a client device associated with each participant of a teleconference. A participant of the teleconference can provide, through the conference UI, a polling question to be presented to other participants of the teleconference. The teleconference platform can present, through the conference UI, the polling question to each other participant of the teleconference and can also present one or more UI elements to allow each participant to provide a verbal answer to the polling question. In some embodiments, a first participant can provide a verbal answer to the polling question before other participants provide answers. In response to detecting that the first participant provided a verbal answer, the conference platform can modify the UI for one or more other participants to present a textual form of the verbal answer provided by the first participant. The conference platform can also modify the UI to include a UI element that allows the other participants to provide an answer to the polling question that corresponds to the verbal answer provided by the first participant.

[0029] In one illustrative example, a first user can provide a verbal answer of "Wednesday" to a particular polling question. The conference platform can modify the UI of a second participant to include a textual form of the verbal answer "Wednesday" and a UI element that enables the second participant to provide a non-verbal answer that corresponds to the answer "Wednesday." The second participant can interact with the UI element to provide a non-verbal answer or can interact with another UI element to provide a (verbal or non-verbal) answer that is different from the answer "Wednesday" (e.g., "Friday"). In a further example, the conference platform can modify the UI of a third participant to include textual forms of the verbal answer "Wednesday" and the answer "Friday" and one or more UI elements that enable the third participant to provide an answer that corresponds to the presented answers.

[0030] Some aspects of the present disclosure automatically populate answers to polling questions to participants of a teleconference based on initial responses provided by other participants of the teleconference. During the teleconference, the UI presented to the participants of the teleconference can be updated to include the polling question and a UI element that allows the participants to provide a verbal response to the question. By allowing the participant to provide a verbal response to the question, the participant is able to provide an answer to the polling question more quickly than if he or she were providing a non-verbal answer to the question (e.g., by the participant typing a response in a text box UI element). In addition, the conferencing platform can suggest an answer to the participant (populate the answer for the participant) based on answers previously provided by other participants before the participant provides his or her answer. The participant can interact with the UI element to provide a non-verbal answer that corresponds to the previously provided answer, or can provide a verbal answer, as previously described. By allowing the participant to provide a non-verbal answer based on previously provided answers, the participant can consider answers to the polling question and can provide an answer more quickly than if the participant were to consider and / or input his or her own answer. In this way, participants of the teleconference can effectively conduct the teleconference discussion with a reduced number of interruptions, thereby reducing the length of the teleconference discussion. As a result of reducing the length of the teleconference discussion, the amount of system resources used to facilitate connections between client devices associated with each participant is reduced. As a result, more system resources are available at the client devices and other computing devices for other processes, thereby improving overall efficiency and reducing overall latency.

[0031] Figure 1 An example system architecture 100 is illustrated in accordance with an embodiment of the present disclosure. The system architecture 100 (also referred to herein as a “system”) includes client devices 102A-N, a data store 110, and a conferencing platform 120, as well as one or more server machines 130-150, each of which is connected to a network 108.

[0032] In embodiments, the network 108 can include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or wide area network (WAN)), a wired network (e.g., an Ethernet network), a wireless network (e.g., an 802.11 network or a Wi-Fi network), a cellular network (e.g., a Long Term Evolution (LTE) network), routers, hubs, switches, server computers, and / or combinations thereof.

[0033] In some implementations, the data store 110 is a persistent storage that is capable of storing data and data structures for tagging, organizing, and indexing data. According to embodiments described herein, data items can include audio data and / or video data. The data store 110 can be hosted by one or more storage devices, such as main memory, magnetic or optical storage-based disks, tapes, or hard drives, NAS, SAN, and the like. In some implementations, the data store 110 can be a network-attached file server, while in other embodiments, the data store 110 can be some other type of persistent storage, such as an object-oriented database, a relational database, and the like, which can be hosted by the conferencing platform 120 or one or more different machines (e.g., server machines 130-150) coupled to the conferencing platform 120 via the network 108.

[0034] The conferencing platform 120 can enable users of the client devices 102A-N to connect with one another via a teleconference, such as a video teleconference or an audio teleconference. A teleconference refers to an audio-based call and / or a video-based call in which a participant of the call can connect with one or more other participants. The conferencing platform 120 can allow users to join and participate in video teleconferences and / or audio teleconferences with other users of the platform. While embodiments of the present disclosure relate to multiple participants (e.g., 3 or more) connected via a teleconference, it should be noted that embodiments of the present disclosure can be implemented with any number of participants (e.g., 2 or more) connected via a teleconference.

[0035] The client devices 102A-N can each include a computing device, such as a personal computer (PC), a laptop computer, a mobile phone, a smartphone, a tablet computer, a netbook computer, a networked television, and the like. In some implementations, the client devices 102A-N can also be referred to as “user devices.” Each client device 102A-N can include a web browser and / or a client application (e.g., a mobile or desktop application), such as the conferencing applications 104A-N. In some implementations, the web browser and / or the conferencing applications 104 can display a graphical user interface (GUI) provided by the conferencing platform 120 for users to access the conferencing platform 120. For example, a user can join and participate in a video teleconference or an audio teleconference via a GUI provided by the conferencing platform 120 and rendered by the web browser or the conferencing applications 104 or another client application.

[0036] Each client device 102A-N can include an audiovisual assembly that can generate audio and video data for streaming to the conferencing platform 120. In some implementations, the audiovisual assembly can include a device (e.g., a microphone) for capturing audio signals representative of a user’s speech and generating audio data (e.g., an audio file) from the captured audio signals. The audiovisual assembly can include another device (e.g., a speaker) for outputting the audio data to a user associated with a particular client device 102A-N. In some implementations, the audiovisual assembly can also include an image capture device (e.g., a camera) for capturing images and generating video data of the captured images.

[0037] In some implementations, the conferencing platform 120 can include a conference management assembly 122. The conference management assembly 122 is configured to manage a conference call between a plurality of users of the conferencing platform 120. In some implementations, the conference management assembly 122 can provide a GUI to each client device that enables the users to view and listen to one another during the conference call. The conference management assembly 122 can also collect and provide data associated with the conference call to each participant of the conference call. For example, the conference management assembly 122 can detect a particular user that is speaking during the conference call and provide a notification to each client device associated with the conference call that includes an identifier of the particular user. In some cases, the conference management assembly 122 and / or components of each respective client device 102A-N can modify the UI based on the notification.

[0038] As previously described, the audiovisual assembly of each client device 102A-N can capture audio signals representative of a user’s speech and generate audio data based on the captured audio signals. For example, a participant of a conference call can provide a spoken phrase. The audiovisual assembly of the client device associated with the participant can capture audio signals that identify the spoken phrase provided by the participant and generate audio data (e.g., an audio file) based on the captured audio signals. In some implementations, the client device 102A-N can transmit the generated audio data to the conference management assembly 122. In some embodiments, the conference management assembly 122 can generate one or more text strings based on the received audio data that include the spoken phrase provided by the participant. For example, the conference management assembly 122 can convert an audio file received from the client device 102A-N into a file that includes one or more text strings. The conference management assembly 122 can store the one or more text strings or the file that includes the one or more text strings at the data store 110. In additional or alternative embodiments, the conference management assembly 122 can store the audio data (e.g., the received audio file) at the data store 110.

[0039] In some embodiments, the conference management component 122 can receive audio data at multiple instances during a conference call. For example, at each instance in which a participant provides a spoken phrase, an audiovisual component of a client device associated with the participant can generate audio data based on the spoken phrase and transmit the audio data to the conference management component 122. In accordance with previously described embodiments, the conference management component 122 can generate a separate text string that includes each spoken phrase provided by a participant of the conference call as each spoken phrase is recorded on a respective client device. During or after completion of the conference call (e.g., after each participant of the call has ended a connection between a client device and the conference platform), the conference management component can generate a transcript of the conference call based on each separately generated text string and store the generated transcript in the data store 110. In other or similar embodiments, the conference management component 122 can receive audio data generated for each participant after completion of a conference call (e.g., from each client device 102A-N). In such embodiments, the conference management component 122 can generate a text string that includes spoken phrases provided by each participant of the conference call after completion of the conference call. Each text string generated after completion of the conference call can be included in a conference call transcript stored at the data store 110. In some embodiments, the conference call transcript can be generated by a transcript generation module 212 of the conference management component 122, as described with respect to Figure 2

[0040] The conference management component 122 can enable a participant of a conference call to poll other participants during the conference call. A participant can poll other participants of a conference call by posing a question to the other participants to solicit responses to the question. In some implementations, a participant can pose a question to other participants by providing the question through a GUI element of a GUI provided by the conference management component 122. For example, a GUI provided by the conference management component 122 can include a text box to enable a participant to type a question for polling and a GUI element (e.g., a button) configured to enable the participant to submit the question to the conference platform 120. Additionally or alternatively, in accordance with embodiments described herein, a participant can provide a question for polling other participants orally.

[0041] ​The recommendation system can be configured to identify a spoken phrase provided by a participant of a teleconference and recommend a polling question 124 to be posed to other participants based on the identified spoken phrase. The recommended polling question 124 can be an indicator (e.g., an interface component, an electronic message, a recommendation feed, etc.) that provides a suggestion of a polling question that can be posed to other participants of a teleconference to the user. For example, during a teleconference, a participant can provide a spoken phrase, “Does everyone agree we should move the meeting to Tuesday??” The recommendation system can determine whether the provided spoken phrase includes a question for polling other participants of the teleconference. The recommendation system can include at least a training set generator 131, a training engine 141, a machine learning model 160, and a polling engine 151. In some implementations, the recommended polling question 124 can be based on an output of a trained machine learning model, such as machine learning models 160A-N.

[0042] The server machine 130 can include a training set generator 131 that can generate training data (e.g., a training input set and a target output set) to train the ML models 160A-N. The training data can be generated based on phrases previously provided by users of a platform, such as the meeting platform 120 or other platforms (e.g., a collaboration platform) included in the system 100, for polling other users of the platform. The collaboration platform (not shown) can provide tools to enable users to collaborate with one another through messages, documents, etc. For example, the collaboration platform can include a survey component to enable a user to prepare a survey to poll other users of the collaboration platform. The user can provide one or more questions to include in the survey (e.g., using a UI provided by the survey component). In response to receiving the one or more questions, the collaboration platform can cause the phrases including each of the provided questions to be stored in the data store 110. In some embodiments, the collaboration platform can also store an indication for each phrase that the phrase includes a question for polling.

[0043] As described above, the data store 110 can include phrases that include questions previously provided by users of a polling platform (e.g., the meeting platform 120, a collaboration platform, etc.). For example, the data store 110 can include questions previously provided by users of a collaboration platform for a survey, as described previously. The data store 110 can also store an indication that the phrase includes a question for polling (e.g., the question is included in a survey created by a user). In some embodiments, the data store 110 can also store an indication of one or more answers provided by a user of the platform in response to the question (e.g., “yes,” “no,” etc.).

[0044] The training set generator 131 can generate a training dataset by identifying data corresponding to previously provided questions stored in the data store 110. The training dataset can include a subset of training inputs and target outputs based on the identified data. The subset of training inputs can include phrases previously provided by users of the platform (e.g., the meeting platform 120, the collaboration platform, etc.). In some embodiments, the phrases can be included in text strings, as previously described. In other or similar embodiments, the phrases can be included as audio data generated by the client device 102 (e.g., a recording of a statement provided by a user of the platform). In some embodiments, the subset of training inputs can also include one or more attributes associated with the previously provided phrases. The training set generator 131 can include an attribute component 132 configured to obtain one or more attributes associated with each phrase at the data store 110. In some embodiments, the attributes can include an indication of whether the phrase includes a question. In similar embodiments, the attributes can include a question type associated with the previously provided question. The question type can correspond to one or more answers provided in response to the question. For example, in response to determining that the one or more answers provided in response to the question correspond to a “yes” answer or a “no” answer, the training set generator 131 can determine that the question corresponds to a “yes / no” type question. The training set generator 131 can determine that the one or more answers correspond to a “yes” answer or a “no” answer by determining that a context associated with each answer provided in response to the previously provided question. For example, the answer to the previously provided question can be “sure” or “okay.” The training set generator 131 can determine that the previously provided question corresponds to a “yes” answer or a “no” answer by determining that the context of the answer to the question corresponds to a “yes” answer. In another example, in response to determining that the one or more answers provided in response to the question correspond to a day of the week, the training generator 131 can determine that the question corresponds to a “day of the week” type question.

[0045] As described above, the training input set can include a recording of a statement provided by a user of the platform. In such embodiments, the attribute component 132 can identify one or more attributes associated with the recording. For example, the attribute component 132 can identify that a tone of the user in the recording corresponds to a recording portion of the question (e.g., the tone of the user rises to a higher pitch). The attributes associated with the recording can include a timestamp associated with the identified portion of the recording. The attribute component 132 can also identify attributes associated with the recording.

[0046] Each target output of the target output subset of the training data set can include data regarding whether the phrase includes a question that was previously used for voting. In accordance with the previously described embodiments, the training set generator 131 can determine whether the phrase includes a question that was previously used for voting based on an indication of each question stored in the data store 110.

[0047] The server machine 140 can include a training engine 141. The training engine 141 can train the machine learning models 160A-N using the training data from the training set generator 131. The machine learning models 160A-N can refer to model artifacts that are created by the training engine 141 using the training data that includes training inputs and corresponding target outputs (correct answers to the respective training inputs). The training engine 141 can find patterns in the training data that map the training inputs to the target outputs (the answers to be predicted) and provide machine learning models 160A-N that capture these patterns. The machine learning models 160A-N can consist of, for example, a single level of linear or non-linear operations (e.g., a support vector machine (SVM) or can be a deep network, i.e., a machine learning model that consists of multiple levels of non-linear operations). One example of a deep network is a neural network with one or more hidden layers, which can be trained by adjusting the weights of the neural network, for example, according to a backpropagation learning algorithm or the like. For convenience, the remainder of this disclosure will refer to embodiments as neural networks, even though some embodiments can use SVMs or other types of learning machines instead of or in addition to neural networks. In an aspect, the training set generator 131 hosted by the server machine 130 obtains the training set. In some embodiments, the machine learning models 160 can be voting question models, as described with respect to Figure 2

[0048] ​Server 150 includes a voting engine 151 configured to facilitate voting by participants of a teleconference. In some embodiments, voting engine 151 can be configured to provide a spoken phrase provided by a participant of a teleconference as input to a trained machine learning model 160A-N to obtain one or more outputs. In some embodiments, the spoken phrase can be included in a text string, as previously described. In other or similar embodiments, the spoken phrase can be included in an audio recording. The model can be used to determine whether the spoken phrase provided by the participant of the teleconference includes a question associated with a poll. The model can provide one or more outputs indicating a likelihood (e.g., a confidence level) that the spoken phrase provided by the user is associated with a poll question. Voting engine 151 can determine whether to recommend the spoken phrase as a poll question by determining whether the confidence level associated with the spoken phrase satisfies a confidence criterion. In some embodiments, voting engine 151 can determine that the confidence level associated with the spoken phrase satisfies the confidence criterion by determining that the confidence level satisfies or meets a threshold confidence level. In response to determining that the confidence criterion is satisfied, voting engine 151 can designate the spoken phrase as a poll question presented during the teleconference and provide the poll question to meeting platform 120 as a recommended poll question 124. In some embodiments, the model can further provide one or more outputs indicating potential answers related to the spoken phrase provided as input to the model. In response to determining that the confidence criterion is satisfied, voting engine 151 can provide the potential answers to the recommended poll question 124 to meeting platform 120.

[0049] As described above, in some embodiments, the GUI provided by the meeting management component 122 can include a GUI element (e.g., a button) to enable a participant to verbally provide a question for a poll. In response to detecting that a participant has interacted with the GUI element, the client device 102A-N associated with the participant can initiate a recording function that causes the audiovisual components of the client device 102A-N to generate audio data (e.g., an audio file) for the verbal phrases provided by the participant. In some embodiments, the client device 102A-N terminates the recording function in response to detecting a subsequent interaction with the GUI element (e.g., the participant has subsequently interacted with the GUI element, the participant has ceased interacting with the GUI element, etc.). In other or similar embodiments, the client device 102A-N terminates the recording function in response to determining that a particular amount of time has elapsed since the participant initially interacted with the GUI element. The client device 102A-N can transmit the generated audio data to the meeting management component 122, and in some embodiments, the meeting management component 122 can generate one or more text strings based on the received audio data, as previously described. In some embodiments, the client device 102A-N can send a message with the generated audio data that indicates that the audio data was generated in response to the participant interacting with the GUI element. In such embodiments, the meeting management component 122 can determine that the verbal phrases correspond to a poll question based on the message and can present the poll question to other participants of the teleconference in accordance with the embodiments described herein.

[0050] In other or similar embodiments, a participant can verbally provide a question for a poll without interacting with a GUI element, as described above. For example, a participant can provide one or more verbal phrases during a teleconference. In accordance with the previously described embodiments, the audiovisual components of the client device 102A-N associated with the participant can generate audio data based on an audio signal that identifies the verbal phrases and can transmit the generated audio data to the meeting management component 122. In some embodiments, the meeting management component 122 can provide the audio data to the polling engine 151. In additional or alternative embodiments, the meeting management component 122 can generate one or more text strings that include the verbal phrases and can provide the generated text strings to the polling engine 151.

[0051] In response to presenting the polling question to the participant via the GUI, the meeting management component 122 can receive one or more answers to the polling question. For example, the GUI provided by the meeting platform 120 can be modified to present the polling question and one or more GUI elements that allow the participant to provide an answer to the polling question. In some implementations, the GUI elements can allow the participant to provide a verbal answer to the polling question. In accordance with the previously described embodiments, the client device 102A-N associated with the participant can detect that the participant has interacted with the GUI element and can initiate a recording function. In response to detecting that the participant has completed providing the verbal answer (e.g., the participant has subsequently interacted with the GUI element, the participant has ceased interacting with the GUI element, etc.), the client device 102A-N can terminate the recording function and can transmit the audio data associated with the verbal answer to the meeting management component 122.

[0052] The meeting management component 122 can generate one or more text strings based on the received audio data and can transmit one or more of the generated text strings that include the verbal answer to the polling question to the client devices 102A-N of additional participants that have not provided an answer to the polling question. The meeting management component 122 can automatically (without any request from any user) populate the possible answers to the polling question for the participant based on the previously received answers (i.e., as populated polling answers 126). For example, the GUI provided by the meeting platform 120 can be modified to present the polling answers 126 in text form and GUI elements (e.g., buttons) that allow the participant to provide a non-verbal answer to the polling question. For example, the other participants can interact with the GUI elements to provide an answer to the polling question corresponding to the answers 126. Additionally or alternatively, the other participants can interact with another GUI element that allows the participant to provide a verbal answer to the polling question. For example, in accordance with the previous embodiments, if the additional participant disagrees with the answers 126 presented via the GUI, the additional participant can interact with the GUI element that allows the additional participant to provide a verbal response. The client devices 102A-N can transmit data associated with the responses (i.e., verbal answers or non-verbal answers) provided by the additional participants to the meeting management component 122. Further details are provided herein regarding collecting and populating answers to polling questions.

[0053] In some embodiments, the meeting platform 120, collaboration platform, and / or server machines 130-150 can be in one or more computing devices such as rackmount servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, and the like, data storage (e.g., hard disks, memory, databases), networks, software components, and / or hardware components that can be used to enable users to connect with other users over a teleconference. In some embodiments, the functionality of the meeting platform 120 can be provided by more than one machine. For example, in some embodiments, the functionality of the meeting management component 122 and the voting engine 151 can be provided by two or more separate server machines. The meeting platform 120 can also include a website (e.g., webpages) or application backend software that can be used to enable users to connect with other users over a teleconference.

[0054] It should be noted that in some other embodiments, the functionality of the server machines 130, 140, and 150 or the meeting platform 120 can be provided by a fewer number of machines. For example, in some embodiments, the server machines 130 and 140 can be integrated into a single machine, while in other embodiments, the server machines 130, 140, and 150 can be integrated into multiple machines. Further, in some embodiments, one or more of the server machines 130, 140, and 150 can be integrated into the meeting platform 120.

[0055] Generally, the functionality described in the embodiments as being performed by the meeting platform 120 or server machines 130, 140, 150 can also be performed on the client devices 102A-N, if appropriate, in other embodiments. Further, the functionality attributed to a particular component is performed by different or multiple components that can run together. The meeting platform 120 and / or server machines 130, 140, 150 can also be accessed as a service provided to other systems or devices through appropriate application programming interfaces, and thus are not limited to use in a website.

[0056] While the embodiments of the present disclosure are discussed in terms of the meeting platform 120 and users of the meeting platform 120 participating in a video and / or audio teleconference, the embodiments can also be applied generally to any type of telephone call or teleconference between users. The embodiments of the present disclosure are not limited to content sharing platforms that provide teleconference tools to users.

[0057] In implementations of the disclosure, a "user" can be represented as a single individual. However, other implementations of the disclosure encompass a "user" that is an entity controlled by a set of users and / or automated sources. For example, a set of individual users that are federated into a community in a social network can be considered a "user." In another example, an automated consumer can be an automated ingestion pipeline of the meeting platform 120, such as a topic channel.

[0058] In situations in which the systems discussed here collect personal information about users, or can make use of personal information, the users can be provided with an opportunity to control whether programs or features collect user information (e.g., information about a user's social network, social actions or activities, profession, a user's preferences, or a user's current location), or to control whether and / or how to receive content from the content server that can be more relevant to the user. In addition, certain data can be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user's identity can be treated so that no personally identifiable information can be determined for the user, or a user's geographic location can be generalized where location information is obtained (such as to a city, postal code, or state level), so that a particular location of a user cannot be determined. Therefore, the user can have control over how information is collected about the user and used by the meeting platform 120.

[0059] Figure 2 is a block diagram illustrating the meeting platform 120 and one example voting engine 151, in accordance with implementations of the disclosure. As described with respect to Figure 1 The meeting platform 120 can provide tools for users of client devices 102 to join and participate in video and / or audio conference calls, as described with respect to

[0060] The text string identifier module 210 of the meeting management component 122 can identify text strings in textual form that include one or more spoken phrases provided by participants of a conference call. In some embodiments, the text strings can be identified during a conference call or based on a transcript 236 generated by the transcript generation module 212 after the conference call, in accordance with previously described embodiments. Figure 3AA video teleconference between a plurality of participants via the conferencing platform 120 is illustrated. As shown, the conference management component 122 provides a UI 310 to enable participants (e.g., participants A-N) to join and participate in the teleconference. The UI 310 is depicted as a UI displayed via a client device 102 associated with participant A of the teleconference. However, it should be noted that the UI 310 can be displayed on a client device 102 associated with any participant of the teleconference.

[0061] The UI 310 can include a plurality of sections, including a first section 312 and a second section 314. In some embodiments, the first section 312 can include one or more portions for outputting video data captured at the client device associated with each participant. For example, the first section 312 can include at least a first portion 316 and a second portion 318, each portion displaying video data captured by a user device associated with a participant of the video teleconference. In some implementations, the first portion 316 of the section 312 can display video data captured by a user device associated with a participant providing an oral statement during the teleconference (i.e., a participant that is currently speaking). In other words, the first portion 316 can display video data associated with a participant that is currently speaking. As shown, participant A provides the oral phrase “Does everyone agree we should move the meeting to Tuesday?” Accordingly, the first portion 316 of the section 312 displays video data captured by the client device associated with participant A. The second portion 318 of the section 312 can display video data captured by the client devices of participants (e.g., participants B-N) that are not providing an oral statement during the teleconference (i.e., participants that are not currently speaking). In other or similar embodiments, the section 312 can include one or more sections configured to display video data associated with participants according to other orientations. For example, the section 312 can include a single portion that displays video data captured by the client device of a participant that is currently speaking and does not display video data captured by the client devices of participants that are not currently speaking. In another example, the section 312 can include a plurality of portions, each portion displaying video data associated with a participant of the video teleconference regardless of whether the participant is currently speaking. Figure 3A

[0062] As Figure 3A ​As shown, participant A can provide the spoken phrase "Does everyone agree we should move the meeting to Tuesday?" In accordance with previously described embodiments, the audiovisual components of the client device associated with participant A can capture an audio signal based on the spoken phrase and generate audio data (e.g., an audio file) based on the captured audio signal. The client device can transmit the audio data to the meeting management component 122. In some embodiments, the client device can transmit an identifier of the participant that provided the spoken phrase with the audio data. In accordance with previously described embodiments, in response to receiving the audio data, the meeting management component 122 can generate one or more text strings based on the received audio data that include the spoken phrase provided by participant A. The text string identifier module 210 can identify the one or more text strings generated by the meeting management component as text strings to provide to the voting question model 220.

[0063] Referring back to Figure 2 , the text string identifier module 210 can identify one or more text strings to provide to the voting question model 220 based on the transcription 236 generated by the transcription generation module 212, as previously described. Figures 4A-4B FIGURE 6 illustrates portions of a transcription generated by the transcription generation module 212 in accordance with previously described embodiments. In some embodiments, the text string identifier module 210 can identify text strings that include spoken phrases provided by participants of a teleconference by parsing each text string included in the generated transcription. In such embodiments, the text string identifier module 210 can identify each text string of the transcription as a text string to provide as input to the voting question model 220. In other or similar embodiments, the text string identifier module 210 can parse each text string included in the generated transcription and determine whether each text string contains a phrase that corresponds to a question. For example, as shown, the text string identifier module 210 can parse each text string included in the transcription 410 and determine that the first text string 412 includes a phrase that corresponds to a question (e.g., "Did everybody have a good weekend?"). Figure 4A In response to determining that a text string includes a spoken phrase that corresponds to a question, the text string identifier module 210 can identify the text string as a text string to provide as input to the voting question model 220.

[0064] In some embodiments, the text string identifier module 210 can identify a text string including a spoken phrase corresponding to an answer to a question by identifying a portion of the transcription that includes one or more spoken phrases corresponding to an answer to a question (e.g., yes, no, etc.). In some embodiments, the text string identifier module 210 can identify a text string corresponding to an answer to a question based on a previously defined list of phrases (e.g., stored in the data store 110) that correspond to answers to questions. For example, the previously defined list of phrases corresponding to answers can include phrases or words such as "yes," "no," or "maybe." As Figure 4B shown, the first portion 422 of the transcription 420 includes a spoken phrase provided by participant B (i.e., "yes") that is included in the previously defined list of phrases corresponding to answers. In other or similar embodiments, the text string identifier module 210 can determine that a phrase corresponds to an answer to a question based on the context of the phrase. For example, the phrases provided by participants C, D, and E included in the portion 422 of the transcription 420 are not included in the previously defined list of phrases. However, the text string identifier module 210 can determine that the context of each of the phrases provided by C, D, and E (e.g., "yeah," "that's fine," and "okay," respectively) corresponds to the context of the phrase "yes," which is included in the previously defined list of device phrases. According to other embodiments, the text string identifier module 210 can identify a portion of the transcription that includes a phrase corresponding to an answer to a question. In some embodiments, the text string identifier module 210 can identify a portion of the transcription in which multiple participants of a teleconference provided the same or similar spoken phrases. For example, the text string identifier module 210 can identify a portion of the transcription in which multiple participants provided spoken phrases that include words indicating a day of the week.

[0065] In response to identifying a portion of the transcription that includes a spoken phrase corresponding to an answer to a question, the text string identifier module 210 can parse the portion of the transcription surrounding the identified portion and determine whether the surrounding portion includes a spoken phrase corresponding to the question. For example, as Figure 4B shown, the text string identifier module 210 can identify that the first portion 422 of the transcription 420 includes a spoken phrase corresponding to an answer to a question. The text string identifier module 210 can parse the transcription 420 to identify the surrounding portion that includes a spoken phrase corresponding to the question. As Figure 4BAs shown in the middle, portion 424, which is adjacent to portion 422, includes a question provided by participant A (e.g., "Does everyone agree that we should change our meeting times to Tuesday?"). As such, text string recognizer module 210 can recognize the text string included in portion 424 of transcription 420 as a text string to provide to voting question model 220.

[0066] Referring back to Figure 2 , in response to recognizing one or more text strings, conference management component 122 can provide the recognized text strings as input to voting question model 220. In some embodiments, other audio data of the conference call (e.g., a portion of an audio file) can be provided as input to voting question model 220. As previously described, voting question model 220 can receive, as input, one or more text strings that include spoken phrases provided by participants of the conference call and provide, as output, a confidence level associated with the one or more text strings that indicates a likelihood that the one or more text strings include a question for voting by other participants of the conference call. In some embodiments, conference management component 122 can obtain one or more attributes associated with the phrases included in each text string, such as the attributes obtained by attribute component 132 described previously. In such embodiments, conference management component 122 can also provide each obtained attribute as input to voting question model 220. In other or similar embodiments, conference management component 122 can provide, as input to voting question model 220, audio data received from a client device associated with participant A to supplement or replace the recognized one or more text strings. For example, in accordance with previously described embodiments, conference management component 122 can provide, as input to voting question model 220, an audio file that includes a recording of the spoken phrase. Figure 1

[0067] ​In response to providing one or more recognized text strings (or audio files including spoken phrases) as input to the voting question model 220, the conference management component 122 can receive, as output, a confidence level associated with the spoken phrase. The conference management component 122 can determine whether the spoken phrase corresponds to a question for voting by determining whether the confidence level associated with the spoken phrase satisfies a confidence criterion. In some embodiments, the conference management component 122 can determine that the confidence level satisfies the confidence criterion in response to determining that the confidence level associated with the spoken phrase satisfies or exceeds a threshold confidence level. In response to determining that the confidence level of the spoken phrase satisfies the confidence criterion, the conference management component 122 can designate the spoken phrase as a question for voting by other participants of the teleconference. In response to the conference component 122 designating the spoken phrase as a question for voting, the conference management component 122 and / or the voting engine 151 can store the designated voting question 232 in the data store 110.

[0068] In some embodiments, the voting question model 220 can also provide, as output, one or more potential responses to the spoken phrase, as previously described. In response to designating the spoken phrase as a voting question, the conference management component 122 can designate the one or more potential responses as potential answers to the voting question and can store the potential answers at the data store 110.

[0069] As previously described, the conference management component 122 can designate a spoken phrase included in one or more text strings as a question for voting during a teleconference. In such embodiments, in response to designating the spoken phrase as a question for voting, the conference management component 122 can generate and transmit a message to a client device associated with a participant who provided the designated spoken phrase, the message asking the participant whether the participant would like to pose the question to other teleconference participants. Figure 3B FIG. 4 illustrates a message provided to participant A after the conference management component 122 designates the spoken phrase provided by participant A (e.g., “Does everyone agree we should move the meeting to Tuesday?”) as a question for voting. As shown, the second section 314 of the UI 310 can display a message 320 to participant A asking participant A whether participant A would like to pose the designated question to other participants of the teleconference (e.g., participants B-N).

[0070] In some embodiments, message 320 can include one or more elements to enable participant A to either pose or not pose the specified question to other participants. For example, message 320 can include a first element 322 that enables participant A to ignore message 320. In response to receiving a notification that participant A interacted with element 322 (e.g., clicked, selected, etc.), the client device associated with participant A can remove message 320 from the second section 314 of UI 310. In some embodiments, the client device can generate and transmit a notification to the meeting management component 122 indicating that participant A ignored message 320. In response to receiving the notification, the meeting management component 122 can determine that the previously specified question does not correspond to a question for enabling conference call participants to vote and can provide feedback to the voting engine 151 and / or the training set generator 131 based on the determination to further train the voting question model 220.

[0071] Message 320 can also include a second element 324 configured to enable participant A to decline to pose the question to other participants of the conference call. In response to receiving a notification that participant A interacted with element 324, the client device associated with participant A can generate and transmit a notification to the meeting management component 122 indicating that participant A does not want to pose the specified question to other participants of the conference call. In accordance with the previously described embodiments, the meeting management component 122 can provide feedback to the voting engine 151 and / or the training set generator 131 in response to receiving the notification.

[0072] Message 320 can also include a third element 326 configured to enable participant A to pose the question to other participants of the conference call. In response to receiving a notification that participant A interacted with element 326, the client device associated with participant A can generate and transmit a notification to the meeting management component 122 indicating that participant A wants to pose the specified question to other participants. In accordance with the previously described embodiments, the meeting management component 122 can pose the question to other participants of the conference call. Figure 3D

[0073] ​Message 320 can also include a fourth element 328 configured to enable participant A to edit the specified polling question included in message 320. In response to receiving a notification that participant A interacts with element 328, the client device associated with participant A can provide, through second section 314 of UI 310, an additional element (not shown) configured to enable participant A to edit one or more portions of the specified polling question. For example, the client device can provide a text box in or around second section 314 of UI 310 to enable participant A to modify one or more words or phrases included in the specified polling question. In response to receiving the modification to the specified polling question, the client device can generate and transmit a notification including the modification to the polling question to meeting management component 122. In accordance with previously described embodiments, meeting management component 122 can provide feedback to polling engine 151 and / or training set generator 131 based on the received notification.

[0074] In some embodiments, in response to receiving an indication that participant A interacts with element 326 (i.e., to indicate that he or she would like to pose the specified polling question to other participants of the teleconference), the client device associated with participant A can modify one or more portions of UI 310 to include an additional message 330. In some embodiments, additional message 330 can enable the participant to customize one or more settings associated with posing the question to other participants of the teleconference. For example, as shown, message 330 provides participant A with answer options that can be associated with the question posed to other participants. As shown, message 330 includes one or more first elements 332 configured to enable participant A to specify one or more answer options associated with the polling question. For example, participant A can specify that the answer options associated with the polling question "Does everyone agree we should move the meeting to Tuesday?" include "yes" or "no." In some embodiments, participant A can specify the one or more answer options by interacting with one or more first elements 332 of message 330. For example, participant A can type or otherwise provide the one or more answer options through one or more first elements 332 of message 330. Figure 3C

[0075] ​In other or similar embodiments, message 330 can include one or more additional elements (not shown) that enable participant A to verbally provide one or more answer options for the polling question. For example, participant A can interact with (i.e., click on) one or more of the additional elements and verbally provide one or more answer options for the polling question. In accordance with the previously described embodiments, in response to determining that participant A has interacted with the one or more additional elements, the client device associated with participant A can generate audio data that includes one or more verbal phrases provided by participant A. In some embodiments, the client device associated with participant A can convert the audio data into one or more text strings that include the verbal phrases. The client device can parse the one or more text strings and identify one or more answers to the question contained in the verbal phrases. In response to identifying one or more answers to the question, the client device associated with participant A can modify message 330 to include each of the identified answers provided by participant A. As described above, the client device can identify one or more answers to the polling question. In other or similar embodiments, the client device associated with participant A can transmit a message that includes the audio data to the meeting management component 122. In accordance with the previously described embodiments, the meeting management component 122 can identify one or more answers to the polling question and transmit the identified answers to the client device associated with participant A.

[0076] In other or similar embodiments, in accordance with the previously described embodiments, message 330 can include answer options determined (e.g., by the meeting management component 122) to be associated with the polling question. It should be noted that while embodiments of the present disclosure are directed to enabling participant A to specify one or more answer options associated with the polling question, message 330 can include additional settings associated with the polling question that participant A can specify. For example, message 330 can include an element that enables participant A to specify a particular participant to whom the question is to be posed.

[0077] Message 330 can include a second element 334 configured to enable participant A to pose a question to one or more participants of the conference call. In response to receiving an indication that participant A has interacted with element 334, the client device can generate and transmit a notification to the meeting management component 122 that indicates that participant A would like to pose a question to other participants of the conference call. In some embodiments, the notification can include an indication of any modifications to the question provided and / or an indication of any settings (e.g., answer options) associated with the polling question.

[0078] In response to receiving the notification, the meeting management component 122 can cause a UI on each client device associated with other participants of the conference call (e.g., participants B-N) to display a message that includes the posed question.Figure 3D UI 350 for participant B of the conference call is illustrated. As shown, the client device associated with participant B updates section 314 of UI 350 to include message 340. Message 340 can include one or more UI elements configured to enable participant B to provide a response to the polling question. In accordance with the previously described embodiments, each element of message 340 can correspond to a specified answer option provided by participant A in some embodiments. For example, as shown, message 340 can include a first element 342 associated with a response corresponding to the answer “yes” and a second element 344 associated with a response corresponding to the answer “no.” Participant B can provide a response of “yes” or “no” by interacting with the respective element. In other or similar embodiments, message 340 does not include one or more elements associated with the specified answer options provided by participant B. In such embodiments, message 340 can include one or more elements (not shown) configured to enable participant B to provide a custom answer to the polling question. For example, message 340 can include a text box configured to enable participant B to enter a custom answer to the polling question. In response to detecting that participant B has provided a response to the polling question (e.g., has interacted with element 342 or element 344, etc.), the client device associated with participant B can generate and transmit a notification indicating the response to the polling question provided by participant B. Figure 3D

[0079] Referring back to Figure 2 , in response to receiving a response to the polling question of the conference call from a participant, the meeting management component 122 can provide the received response to the answer collection component 222. The answer collection component 222 can store the received response (i.e., the answer to the polling question) as a collected answer 234 at the data store 110. In some embodiments, the answer collection component 222 can generate a mapping between the specified polling question and the collected answer 234 and store the mapping at the data store 110.

[0080] ​During the conference call, the poll question model 220 can identify one or more poll questions 232 and the answer collection component 222 can collect answers to the poll questions provided by participants of the conference call. During or after the conference call, the poll results component 224 can generate results for each poll question stored in the data store 110. For example, the poll results component 224 can identify each poll question 232 posed during the conference call at the data store 110 and each collected answer associated with each poll question 232. The poll results component 224 can analyze each collected answer for each poll question 232 and generate data associated with each poll question 232 based on the analysis. The generated data can be stored at the data store 110 as poll results 238. In some embodiments, the generated data can include a number of participants that provided a particular answer to a poll question. According to the previously provided example, the poll results component 224 can determine a number of participants that answered "yes" and "no" to the poll question "Does everyone agree we should move the meeting to Tuesday?" In other or similar embodiments, the poll results 238 can include data associated with participants that provided a particular answer. For example, the poll results component 224 can determine that a particular portion of participants that answered "yes" to the poll question are associated with a particular characteristic (e.g., identified via a profile associated with each participant).

[0081] As previously described, the text string identifier module 210 can identify (e.g., from the transcription 236) text strings to provide to the poll question model 220 after the conference call is completed. The conference management component 122 can provide each identified text string as input to the poll question model 220 and receive as output a confidence level indicating a likelihood that the identified text string includes a spoken statement corresponding to a poll question. According to the previously described embodiments, the conference management component 122 can determine whether a spoken phrase corresponds to a poll question based on the confidence level. With reference to Figure 4A and 4BBased on the confidence levels provided by the voting question model 220, the meeting management component 122 can determine that the text string 412 (i.e., including the question "Did everybody have a good weekend?") does not include a spoken statement corresponding to a voting question and that the text string (i.e., including the question "Does everyone agree that we should change our meeting time to Tuesday?") does include a spoken statement corresponding to a voting question. In accordance with the previously described embodiments, the voting engine 151 can store the questions included in the text strings at the data store 110.

[0082] The answer collection component 222 can identify an answer corresponding to each specified voting question 232 included in the transcription 236. The answer collection component 222 can identify one or more text strings that are within a particular proximity to a text string of the transcription 236 that includes a specified voting question. In some embodiments, the particular proximity can correspond to a distance between a text string that includes a specified voting question and an additional text string of the transcription, where the distance corresponds to a number of participants of the teleconference. The answer collection component 222 can determine whether each identified text string includes a spoken phrase corresponding to an answer to the specified voting question. In response to determining that a text string includes a spoken phrase corresponding to an answer to a specified voting question, the answer collection component 222 can generate a mapping between the answer and the voting question and store the mapping and / or the answer in the data store 110. As shown in FIG. 4B, the answer collection component 222 can determine that each text string that is within the particular proximity to the text string (i.e., the text strings included in the portion 422) includes a spoken phrase corresponding to an answer to the voting question "Does everyone agree that we should change our meeting time to Tuesday?" In this way, the answer collection component 222 can generate a mapping between each answer included in a text string and a specified voting question and store the mapping and / or each answer in the data store 110. Figure 4B

[0083] In accordance with the above-described embodiments, the voting results component 224 can generate voting results 238 for the voting questions and answers identified from the transcription 236. In response to the voting results component 224 generating the voting results 238, the meeting management component 122 can provide the voting results 238 to a client device 102 associated with one or more participants of the teleconference. For example, the meeting management component 122 can provide the voting results 238 to a client device associated with an organizer of the teleconference.

[0084] Figure 5 ​is a block diagram illustrating a meeting platform 120 and another example voting engine 151 in accordance with implementations of the present disclosure. In some embodiments, Figure 5 The voting engine 151 illustrated in FIG. 1 can be the same as or similar to the voting engine 151 described with respect to Figure 2 The voting engine 151 described with respect to Figure 5 The voting engine 151 illustrated in FIG. 1 can be a different voting engine than the voting engine 151 described with respect to Figure 2 The voting engine 151 described with respect to

[0085] As described with respect to Figure 1 The meeting platform 120 can provide users of client devices 102 with tools to join and participate in video and / or audio conference calls, as described above with respect to the meeting platform 120. The meeting platform 120 can include a meeting management component 122. The meeting management component 122 can include a text string identifier module 510. In some embodiments, the text string identifier module 510 can correspond to the text string identifier module 210, described above. As described above. The voting engine 151 can facilitate voting by participants of a conference call. The voting engine 151 can include a voting question component 520 and an answer collection component 522. In some embodiments, the answer collection component 522 can correspond to the answer collection component 222 described above. One or more operations performed by the meeting management component 122 and / or the voting engine 151 are described with respect to Figures 6A-6D The meeting management component 122 and / or the voting engine 151 perform one or more operations.

[0086] Figure 6A An example user interface (UI) for a video conference call between multiple participants via a meeting platform 120 is illustrated in accordance with some implementations of the present disclosure. As shown, the meeting management component 122 provides a UI 610 to enable participants (e.g., participants A-N) to join and participate in a conference call. The UI 610 is described as a GUI displayed via a client device 102 associated with participant A of the conference call. However, it should be noted that the UI 610 can be displayed on a client device 102 associated with any participant of the conference call. In some embodiments, the UI 610 can correspond to the UI 310 described above.

[0087] UI 610 may have multiple sections, including a first section 612 and a second section 614. In some embodiments, the first section 614 may include one or more portions for outputting video data captured at a client device associated with each participant. For example, the first section 612 may include at least a first portion 616 and a second portion 618, each portion displaying video data captured by a user device associated with a participant in the video conference call. In some embodiments, the first portion 616 of section 612 may display video data captured by a user device associated with a participant providing oral presentations during the conference call (i.e., the participant currently speaking). In other words, the first portion 616 may display video data associated with the participant currently speaking. Figure 6A As shown, participant A is providing the spoken phrase, “What day should we move the meeting to next week?”. Therefore, the first portion 616 of section 612 displays video data captured by the client device associated with participant A. The second portion 618 of section 612 may display video data captured by the client device of a participant (e.g., participant BN) who is not currently providing a spoken statement during the conference call (i.e., not currently speaking). In other or similar embodiments, section 612 may include one or more sections configured to display video data associated with a participant based on other orientations. For example, section 612 may include a single section that displays video data captured by the client device of a participant currently speaking and does not display video data captured by the client devices of participants not currently speaking. In another example, section 612 may include multiple sections, each displaying video data associated with a participant in the video conference call, regardless of whether the participant is currently speaking.

[0088] As previously described, participants in a conference call can provide voting questions to be presented to other participants. For example, participant A can provide the spoken phrase “What day should we move the meeting to next week?”. In some embodiments, participant A can provide the spoken phrase in response to interaction with a specific GUI element (not shown) of UI 610, which initiates a recording operation via the audiovisual component of a client device associated with participant A. According to the previously described embodiments, the audiovisual component of the client device can capture an audio signal based on the provided spoken phrase and generate audio data (e.g., an audio file) based on the captured audio signal. In other or similar embodiments, the audiovisual component can capture an audio signal without participant A interacting with a specific GUI element to initiate a recording operation. The client device can transmit the audio data to the conference management component 122. In some embodiments, the client device can transmit an identifier of the participant who provided the spoken phrase with audio data. According to the previously described embodiments, in response to receiving audio data, the conference management component 122 can generate one or more text strings based on the received audio data, which include the spoken phrase provided by participant A.

[0089] Return to reference Figure 5 The text string recognizer module 510 can determine to provide one or more text strings generated by the meeting management component 122 to the voting question component 520. In some embodiments, the text string recognizer module 510 can determine to provide one or more text strings to the voting question component 520 in response to determining that a spoken phrase provided by participant A corresponds to a question phrase. For example, the text string recognizer module 510 can parse each received text string and determine that the context of the text string corresponds to the context of the question phrase. In some embodiments, the voting question component 520 can provide the text string received from the text string recognizer module 510 as input to a trained machine learning model. The machine learning model can be trained to receive text strings in text form including one or more spoken phrases as input and provide the confidence level of the spoken phrases in the received text strings including the question associated with the vote as output. The voting question component 520 can obtain one or more outputs from the trained machine learning model and determine whether the confidence level of the spoken phrases meets a confidence criterion (e.g., whether the confidence level of the spoken phrases exceeds a threshold confidence level) based on the received outputs. In response to determining that the confidence level meets the confidence criteria, the voting engine 151 can designate the spoken phrase as the voting question 532 and store the voting question 532 at the data storage 110.

[0090] In other or similar embodiments, the voting question component 520 may determine whether a spoken phrase associated with a text string is voting question 532 based on a message received by the meeting management component 122 containing audio data for a spoken phrase. For example, participant A may interact with a GUI element to initiate a recording operation at a client device. In such an example, the client device may transmit a message containing audio data associated with a recorded spoken phrase, indicating that the spoken phrase corresponds to a voting question. Given the received message, the voting question component 520 may designate the spoken phrase as voting question 532 and store voting question 532 at the data storage 110.

[0091] In some implementations, the voting question component 520 may allow the UI to present the voting question 532 only to participant A (or alternatively to some participants specified by participant A) and request participant A to confirm that participant A wants to present the voting question 532 to other participants (all other participants or some other participants specified by participant A).

[0092] As described above, in some embodiments, participant A can provide a voting question to be presented to other participants (e.g., participant BN) without providing a spoken phrase. For example, participant A can type the voting question into a text box GUI element. In such an embodiment, meeting management component 122 can receive the non-spoken phrase along with a message indicating that the non-spoken phrase corresponds to a voting question from a client device associated with participant A. Meeting management component 122 can transmit the text string including the non-spoken phrase to voting question component 520. According to the previously described embodiments, voting question component 520 can designate the non-spoken phrase as the voting question based on the message received from the client device.

[0093] In response to the voting question component 520 designating a phrase (oral or non-oral) as the voting question, the conference management component 122 can transmit the text form of the voting question 532 for presentation on client devices associated with other participants in the conference call. For example, in response to the voting question component 520 designating an oral phrase provided by participant A as the voting question 532, the conference management component 122 can transmit the text form of the voting question 532 for presentation on client devices associated with participant BN.

[0094] Figure 6BThe illustration shows an example UI 620 of participant B responding to a voting question during a conference call. As shown, in response to the voting question component 620 designating the question “What day should we move the meeting to next week?” as voting question 532, the meeting management component 122 causes UI 620 to present a message 622 including the question in section 614. In some embodiments, section 614 also includes a GUI element 624 that allows participant B to provide a verbal answer to the voting question. According to the previously described embodiments, in response to participant B's interaction with GUI element 624, a client device associated with participant B can initiate a recording operation to generate audio data for the provided verbal answer. The client device can transmit the audio data associated with the provided verbal answer to the meeting management component 122, and the meeting management component 122 can generate one or more text strings including the provided verbal answer, as previously described. In some embodiments, the client device can also transmit a message to the meeting management component 122 indicating that the provided verbal answer corresponds to an answer phrase, since the verbal phrase is provided in response to UI 620 presenting voting question 232.

[0095] Return to reference Figure 2 The text string recognizer module 510 can determine whether a provided verbal response included in one or more text strings corresponds to an answer phrase. For example, the text string recognizer module 510 can parse each text string and determine that the context of the text string corresponds to the context of the answer phrase. In another example, according to the previously described embodiment, the text string recognizer module 510 can determine that the provided verbal response corresponds to an answer phrase based on a message received from a client device. In response to determining that the provided verbal response corresponds to an answer phrase, the text string recognizer module 510 can provide one or more text strings including the verbal response to the voting response component 522.

[0096] The voting response component 522 can store the provided verbal responses as collected responses 524 at data storage 110. In some embodiments, the voting response component 522 can determine whether the text form of the provided verbal responses should be presented to one or more additional participants in the conference call as a candidate response to the voting question 532. The voting response component 522 can determine that the provided verbal responses will be presented to additional participants in response to determining that the provided verbal responses meet the response criteria. The provided responses meet the response criteria in response to the voting response component 522 determining that the number of conference call participants providing the corresponding responses meets a threshold number. For example, in response to the voting question “What day should we move the meeting to next week?”, participant B can provide the verbal response “Wednesday”. The voting response component 522 can determine, based on the collected responses 534, the number of additional participants in the conference call who also provided the response (verbal or nonverbal) “Wednesday”. In response to determining that the total number of responses corresponding to the response “Wednesday” meets a threshold number, the voting response component 522 can determine that the text form of the provided verbal responses will be presented to one or more participants in the conference call (only additional participants or alternatively one or more other participants and participant B). In some embodiments, the voting response component 522 may determine that the provided verbal response will be presented to the participant in response to receiving one or more text strings including a verbal response, without determining whether the verbal response meets the response criteria.

[0097] In some embodiments, the voting response component 522 may determine one or more additional response candidates to present to additional participants along with the provided verbal response. For example, the voting response component 522 may determine that the context of the verbal response (“Wednesday”) provided by participant B corresponds to a day of the week. The voting response component 522 may determine, based on the response context, that the response to the voting question corresponds to a specific day of the week and may specify the voting question as a question of the “day of the week” type. In response to specifying the voting question as a question of the “day of the week” type, the voting response component 522 may determine one or more response candidates (e.g., Monday, Tuesday, Thursday, Friday, etc.) corresponding to a question of the “day of the week” type. The voting response component 522 may determine that one or more response candidates will be presented to participants along with the verbal response and may store the response candidates as filler responses 536 at the data storage 110.

[0098] Figures 6C-6DThe illustration shows a sample UI according to an embodiment of the present disclosure, which fills in the answers to voting questions for other participants in a teleconference based on previously provided verbal responses from other participants.

[0099] Figure 6C The illustration shows a sample UI 630 of participant C responding to a voting question during a conference call. As shown, in response to the voting response component 522 determining that a verbal response provided by participant B will be presented to additional participants in the conference call, the conference management component 122 causes UI 630 to display a message 632 in text form, including the verbal response 634, in section 614. In some embodiments, section 614 also includes a GUI element 636 that allows participant C to provide a nonverbal response to the voting question. It should be noted that message 632 and / or GUI element 636 may be provided in a section different from section 614 of UI 630. Participant C may interact with GUI element 636 to provide a nonverbal response to the voting question corresponding to the verbal response. For example, GUI element 636 may be a button element that participant C can select to provide a nonverbal response corresponding to the answer “Wednesday”. Alternatively, GUI element 636 may be a checkbox, button, or another UI element that allows participants to confirm that the generated text (e.g., “Wednesday”) is their answer to the voting question or, in some embodiments, to reject the generated text (e.g., “Wednesday”) as their answer to the voting question. In some embodiments, GUI element 636 is an element separate from or adjacent to the UI element displaying the generated text (e.g., “Wednesday”). In some embodiments, according to the previously described embodiments, message 632 may also include a textual form of each answer candidate identified by the voting question component 220 (not shown), and one or more additional GUI elements that allow participant C to provide a nonverbal answer corresponding to the answer candidate.

[0100] Additionally or alternatively, participant C may interact with GUI element 624 to provide a verbal answer to the voting question. For example, participant C may want to provide an answer different from the answer options contained in message 632 (i.e., verbal answer 634, answer candidate (not shown)). Thus, according to the previously described embodiment, participant C can provide a verbal answer to the voting question by interacting with GUI element 624. In other or similar embodiments, participant C may interact with additional GUI elements (not shown) to provide a nonverbal answer to the voting question (e.g., typing a nonverbal answer in a text box element). As previously described, the client device associated with participant C can transmit data associated with the provided answer (verbal or nonverbal) to meeting management component 122.

[0101] According to the previously described embodiments, the conference management component 122 may receive data associated with a response (verbal or nonverbal) provided by participant C. In some embodiments, the voting response component 122 may determine that the response provided by participant C will be presented to other participants in the conference call. For example, participant C may provide the response "Friday" to a voting question. According to the previously described embodiments, in some embodiments, the voting response component 122 may determine that the provided response meets the response criteria and will be presented to additional participants in the conference call.

[0102] Figure 6D The illustration depicts a sample UI 640 of participant D responding to a voting question during a teleconference, according to some implementations. As shown, in response to a voting response component 122 determining that the answer provided by participant C meets the response criteria, the conference management component 122 causes a section of UI 640 to include a textual form of answer 642 (oral or nonverbal). For example, answer 642 may be provided in message 632 of section 614. In some embodiments, according to the previously described embodiments, section 614 (or another section of UI 640) also includes a GUI element that allows participant D to provide a nonverbal answer to the voting question. Additionally or alternatively, participant D may interact with additional GUI elements (e.g., GUI element 624) to provide different answers to the voting question, as previously described.

[0103] In some implementations, GUI elements 636 and 644 may be provided as checkboxes, buttons, or other types of UI elements that allow participants to confirm that the generated text (e.g., “Wednesday” or “Friday”) is their answer to the voting question (or selectively reject the generated text (e.g., “Wednesday” and / or “Friday”) as their answer to the voting question). In some implementations, GUI elements 636 and 644 are elements separate from or adjacent to the corresponding UI elements displaying the generated text (e.g., “Wednesday” and “Friday”).

[0104] In some embodiments, a voting results component (not shown) may generate voting results 538 for voting on question 532 and answer 534, stored in data storage 110. In response to the voting results component generating voting results 538, a conference management component 122 may provide voting results 538 to client devices 102 associated with one or more participants in the conference call. For example, the conference management component 122 may provide voting results 538 to a client device associated with the organizer of the conference call.

[0105] Figure 7A flowchart is depicted for a method 700 for designating spoken phrases provided during a teleconference as voting questions, according to an embodiment of the present disclosure. Figure 8 A flowchart is depicted for a method 800 for training a machine learning model to identify spoken phrases corresponding to voting questions provided during a teleconference, according to an embodiment of the present disclosure. Figure 9 A flowchart depicts a method 900 for filling in answers to a voting question based on an initial response, according to an embodiment of the present disclosure. Methods 700, 800, and 900 may be executed by processing logic that may include hardware (circuit, dedicated logic, etc.), software (e.g., instructions running on a processing device), or a combination thereof. In one embodiment, it may be performed by… Figure 1 One or more components of system 100 perform some or all of the operations of methods 700, 800, and 900.

[0106] Now for reference Figure 7 In block 710, processing logic identifies one or more text strings comprising textual forms of one or more spoken statements provided by one or more participants in a teleconference. The processing logic can identify one or more text strings by generating an audio file during the teleconference that includes one or more spoken phrases provided by one or more participants in the teleconference. The processing logic can convert the content of the audio file into a set of text strings comprising one or more text strings. In some embodiments, the processing logic can convert the content of the audio file into a set of text strings by generating a transcript of the teleconference that includes the set of text strings during or after the teleconference. The processing logic can identify one or more text strings based on the generated transcript. For example, the processing logic can identify a specific text string comprising textual forms of additional spoken phrases provided by participants in the teleconference, where the additional spoken phrases correspond to answers to potential questions. The processing logic can determine that the distance between the text string of the transcript including the spoken phrases and the specific text string including the additional spoken phrases meets a distance criterion (e.g., meets or exceeds a distance threshold).

[0107] In block 720, the processing logic feeds one or more text strings as input to the trained machine learning model. In block 730, the processing logic obtains one or more outputs from the trained machine learning model. In block 740, the processing logic extracts the confidence level of the first oral statement in one or more oral statements from the one or more outputs. In block 750, the processing logic determines whether the confidence criterion is met. In response to the processing logic determining that the confidence criterion is met, method 700 continues to block 760. In response to the processing logic determining that the confidence criterion is not met, method 700 terminates. In block 760, the processing logic designates the first oral statement as the voting problem.

[0108] In some embodiments, in response to designating spoken phrases as voting questions presented during a conference call, the processing logic may generate a file including the results of the voting questions. A first portion of the file includes one or more text strings containing spoken phrases, and a second portion includes data corresponding to one or more additional spoken phrases associated with the answers to the voting questions.

[0109] In some embodiments, in response to designating a spoken phrase as a voting question to be presented during a conference call, processing logic may identify a specific participant who provides one or more spoken phrases. The processing logic may cause the spoken phrase to be displayed in a first UI of a first client device associated with the specific participant. The first UI may include one or more first UI elements configured to enable the specific participant to present the spoken phrase as a voting question to one or more other participants in the conference call. In response to receiving an indication that the specific participant has interacted with one or more first UI elements of the first UI, the processing logic may cause a notification to be displayed through a second UI of a second client device associated with one or more additional participants in the conference call. The notification may include a voting question and one or more second UI elements configured to enable the additional participant to respond to the voting question.

[0110] As mentioned above, Figure 8 A flowchart of a method 800 for training a machine learning model to identify spoken phrases corresponding to voting questions provided during a teleconference, according to an embodiment of this disclosure, is depicted. In block 810, processing logic initializes a training set T to {}. In block 820, processing logic identifies data corresponding to phrases provided by platform users. In block 830, processing logic generates an input / output map, with input based on the identified data and output identifying whether the phrase corresponds to a question previously used to enable additional user voting on the platform. In block 840, processing logic adds the input / output map to the training set T. In block 850, processing logic determines whether set T is sufficient for training. In response to processing logic determining that set T is sufficient for training, method 800 continues to block 860. In response to processing logic determining that set T is insufficient for training, method 800 returns to block 820. In block 860, processing logic provides the training set T to train the machine learning model.

[0111] As mentioned above, Figure 9A flowchart of a method 900 for filling in answers to a voting question based on an initial response, according to an embodiment of the present disclosure, is depicted. At block 910, a processing device provides a graphical user interface (GUI) to a set of participants in a teleconference. The GUI includes a first portion that presents the voting question to be answered by the set of participants. At block 920, the processing device detects that a first participant has provided a verbal answer to the voting question. At block 930, the processing device determines whether the answer criteria are met. In response to determining that the answer criteria are not met, method 900 terminates. In response to determining that the answer criteria are met, method 900 continues to block 940. At block 940, the processing device modifies the GUI provided to a second participant to present the textual form of the verbal answer provided by the first participant and GUI elements that allow the second user to provide a nonverbal answer to the voting question by selecting GUI elements. At block 950, the processing device detects that the second participant has provided a verbal and / or nonverbal answer to the voting question. At block 960, the processing device modifies the GUI provided to a third participant to present the textual form of the verbal answer provided by the first participant and the verbal and / or nonverbal answers provided by the second participant.

[0112] Figure 10 This is a block diagram illustrating an exemplary computer system according to an embodiment of the present disclosure. The computer system 1000 may be... Figure 1 The conferencing platform 120 or client devices 102A-N are used. Machines can operate as servers or endpoint machines in an endpoint-server network environment, or as peer machines in a peer-to-peer (or distributed) network environment. A machine can be a television, personal computer (PC), tablet, set-top box (STB), personal digital assistant (PDA), cellular phone, network device, server, network router, switch, or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) specifying the actions the machine should take. Furthermore, although only a single machine is illustrated, the term "machine" should also be considered to include any collection of machines that individually or jointly execute one or more sets of instructions to perform any one or more methods discussed herein.

[0113] Example computer system 1000 includes a processing device (processor) 1002, main memory 1004 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) (e.g., synchronous DRAM (SDRAM), dual data rate (DDRSDRAM) or DRAM (RDRAM) etc.), static memory 1006 (e.g., flash memory, static random access memory (SRAM) etc.) and data storage device 1018, which communicate with each other via bus 1040.

[0114] Processor (processing device) 1002 represents one or more general-purpose processing devices, such as microprocessors, central processing units, etc. More specifically, processor 1002 may be a Complex Instruction Set Computing (CISC) microprocessor, a Reduced Instruction Set Computing (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, or a processor implementing other instruction sets or combinations thereof. Processor 1002 may also be one or more special-purpose processing devices, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), network processors, etc. Processor 1002 is configured to execute instruction 1005 (e.g., for predicting channel queuing ratings) to perform the operations discussed herein.

[0115] The computer system 1000 may also include a network interface device 1008. The computer system 1000 may also include a video display unit 1010 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an input device 1012 (e.g., a keyboard and alphanumeric keypad, a motion-sensing input device, a touch screen), a cursor control device 1014 (e.g., a mouse), and a signal generation device 1020 (e.g., a speaker).

[0116] Data storage device 1018 may include a non-transitory machine-readable storage medium 1024 (also a computer-readable storage medium) storing one or more instruction sets 1005 (e.g., for predicting channel queuing ratings), embodying any one or more methods or functions described herein. Instructions may also reside wholly or at least partially within main memory 1004 and / or processor 1002 during execution by computer system 1000, which also constitute machine-readable storage media. These instructions may also be sent or received on network 1030 via network interface device 1008.

[0117] In one embodiment, instruction 1005 includes instructions for designating a verbal statement as a voting question. In additional or alternative embodiments, instruction 1005 includes instructions for populating the answer to the voting question based on the initial response. Although computer-readable storage medium 1024 (machine-readable storage medium) is shown as a single medium in the exemplary embodiment, the terms "computer-readable storage medium" and "machine-readable storage medium" should be considered to include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) storing one or more sets of instructions. The terms "computer-readable storage medium" and "machine-readable storage medium" should also be considered to include any medium capable of storing, encoding, or carrying a set of instructions for machine execution and enabling the machine to perform any one or more methods of this disclosure. The terms "computer-readable storage medium" and "machine-readable storage medium" should accordingly be considered to include, but are not limited to, solid-state memory, optical media, and magnetic media.

[0118] Throughout this specification, the reference to "one embodiment" or "implementation" means that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. Therefore, depending on the context, the appearance of the phrase "in one embodiment" or "in an embodiment" in different places throughout this specification may, but does not necessarily, refer to the same embodiment. Furthermore, in one or more embodiments, particular features, structures, or characteristics may be combined in any suitable manner.

[0119] With regard to the use of the terms “comprising,” “containing,” “having,” “including,” variations thereof, and other similar words in the detailed description or claims, these terms are intended to be inclusive in a manner similar to the term “comprising” as an open transition word, without excluding any additional or other elements.

[0120] As used herein, the terms “component,” “module,” “system,” etc., are generally intended to refer to a computer-related entity (which is hardware (e.g., circuitry), software, or a combination of hardware and software) or an entity associated with an operating machine having one or more specific functions. For example, a component can be, but is not limited to, a process, processor, object, executable, execution thread, program, and / or computer running on a processor (e.g., a digital signal processor). For illustration, an application running on a controller and the controller itself can both be components. One or more components may reside within an execution process and / or thread, and components may reside on a single computer and / or be distributed among two or more computers. Furthermore, “device” can appear in the following forms: specially designed hardware; general-purpose hardware specialized by executing software thereon to perform a specific function (e.g., generating points of interest and / or descriptors); software on a computer-readable medium; or a combination thereof.

[0121] The aforementioned systems, circuits, modules, etc., have already been described with regard to the interactions between several components and / or blocks. It is understood that such systems, circuits, components, blocks, etc., may include those components or designated sub-components, some designated components or sub-components, and / or additional components, according to the various permutations and combinations described above. Sub-components may also be implemented as components communicatively coupled to other components, rather than being contained within a parent component (hierarchical structure). Furthermore, it should be noted that one or more components may be combined into a single component providing aggregate functionality or divided into several separate sub-components, and any one or more intermediate layers, such as a management layer, may be provided to communicatively couple to such sub-components to provide integrated functionality. Any component described herein may also interact with one or more other components not specifically described herein but known to those skilled in the art.

[0122] Furthermore, the terms “example” or “exemplary” are used herein to indicate that something is used as an example, instance, or illustration. Any aspect or design described herein as “exemplary” is not necessarily to be construed as preferred or superior to other aspects or designs. Rather, the use of the terms “example” or “exemplary” is intended to present concepts in a specific manner. As used in this application, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless otherwise stated or clear from the context, “X uses A or B” is intended to mean any naturally inclusive arrangement. That is, if X uses A; X uses B; or X uses A and B, then “X uses A or B” is satisfied in any of the foregoing cases. In addition, the articles “a” and “an” used in this application and the appended claims should generally be construed as meaning “one or more” unless otherwise specified or clearly pointed to from the context in the singular form.

[0123] Finally, the implementations described herein include collecting data describing users and / or user activities. In one implementation, such data is collected only if the user consents to its collection. In some implementations, the user is prompted to explicitly allow data collection. Furthermore, the user can choose to join or opt out of such data collection activities. In one implementation, the collected data is anonymized before any analysis is performed to obtain any statistical patterns, making it impossible to determine the user's identity from the collected data.

Claims

1. A method for designating a spoken phrase as a polling question, comprising: identifying one or more text strings comprising a textual form of one or more spoken phrases provided by one or more participants of a teleconference during the teleconference; providing the one or more text strings as input to a trained machine learning model; obtaining one or more outputs from the trained machine learning model; extracting, from the one or more outputs, a confidence level that a spoken phrase of the one or more spoken phrases provided by the one or more participants during the teleconference comprises a question associated with a poll during the teleconference; in response to determining that the confidence level satisfies a confidence criterion, designating the spoken phrase as a polling question presented during the teleconference; identifying, among the one or more participants of the teleconference, a particular participant that provided the spoken phrase of the one or more spoken phrases; causing the spoken phrase to be displayed in a user interface (UI) of a particular client device associated with the particular participant of the teleconference; and in response to receiving an indication that the particular participant has interacted with one or more UI elements of the UI of the particular client device, displaying a notification for the polling question via a UI of an additional client device associated with an additional participant of the one or more additional participants of the teleconference.

2. The method of claim 1, wherein, identifying the one or more text strings comprises: generating an audio file during the teleconference, the audio file comprising the one or more spoken phrases provided by the one or more participants of the teleconference; and converting content of the audio file to a set of text strings comprising the one or more text strings.

3. The method of claim 2, wherein, converting the content of the audio file to the set of text strings comprises: generating a transcription of the teleconference comprising the set of text strings, wherein the one or more text strings are identified based on the generated transcription.

4. The method of claim 3, wherein, identifying the one or more text strings further comprises: identifying, in the transcription, a particular text string comprising a textual form of an additional spoken phrase provided by a participant of the teleconference, the additional spoken phrase corresponding to an answer to a potential question; and determining that a distance between a text string of the transcription comprising the spoken phrase and the particular text string comprising the additional spoken phrase satisfies a distance criterion.

5. The method of claim 3, further comprising: in response to designating the spoken phrase as the polling question presented during the teleconference, generating a file comprising a result of the polling question, wherein a portion of the file comprises the one or more text strings containing the spoken phrase, and wherein an additional portion of the file data comprises data corresponding to one or more additional spoken phrases associated with answers to the polling question.

6. The method of any one of claims 1-5, wherein, The UI of the particular client device includes one or more UI elements configured to enable the particular participant to pose the spoken phrase as the polling question to one or more additional participants of the teleconference.

7. The method of claim 6, wherein, The notification includes the polling question and one or more additional UI elements configured to enable the additional participant to respond to the polling question.

8. A system for training a machine learning model to determine a spoken phrase as a polling question, comprising: a storage device; and a processing device coupled to the storage device, the processing device performing operations comprising: generating training data for a machine learning model, wherein generating the training data comprises: generating training inputs containing data corresponding to prior phrases provided by users of a platform; and generating target outputs for the training inputs, wherein the target outputs include an indication of whether the phrases correspond to questions that were previously used to poll one or more additional users of the platform; and providing the training data to train the machine learning model on (i) a set of training inputs comprising the training inputs and (ii) a set of target outputs comprising the target outputs, wherein, based on a given one or more text strings comprising a textual form of one or more spoken phrases provided by one or more participants of a teleconference, one or more outputs of the trained machine learning model include a confidence that a spoken phrase of the one or more spoken phrases includes a question associated with a poll during the teleconference, and wherein a spoken phrase that satisfies a meeting criterion is displayed via a UI of a client device associated with at least one participant of the one or more participants of the teleconference in response to an indication that an additional participant of the one or more participants that provided the spoken phrase has interacted with one or more UI elements of a UI displayed via an additional client device associated with the additional participant.

9. The system of claim 8, wherein: to generate the training inputs comprising data corresponding to the phrases provided by the users of the platform, the operations further comprise identifying one or more text strings comprising provided phrases from a database comprising a plurality of spoken phrases provided by the users of the platform.

10. The system of claim 9, wherein, to determine whether the question was used to poll the one or more additional users of the platform, the operations further comprise: determining whether the phrase corresponding to the question was previously provided to solicit a response to the question from the one or more additional users.

11. The system of any one of claims 8 to 10, wherein, the machine learning model is configured to determine whether one or more spoken phrases provided by one or more participants of a teleconference include a question associated with a poll during the teleconference.

12. A non-transitory computer-readable storage medium comprising instructions for a server that, when executed by a processing device, cause the processing device to perform the method of any one of claims 1-7.

13. A method for populating answers to a polling question, comprising: providing a graphical user interface (GUI) comprising a first portion to a plurality of participants of a teleconference, the first portion presenting a polling question to be answered by the plurality of participants; determining, from audio data received from a client device associated with a proponent of the teleconference, that the proponent has provided a polling question to the plurality of participants during the teleconference; updating the first portion of the GUI provided to each of the plurality of participants to present a textual form of the polling question and a first GUI element that allows each of the plurality of participants to provide a spoken answer to the polling question during the teleconference; and in response to detecting that a first participant of the plurality of participants has provided a spoken answer to the polling question via selection of the first GUI element, modifying the GUI provided to a second participant of the plurality of participants to present, in a second portion, a textual form of the spoken answer provided by the first participant and a second GUI element that allows the second participant to provide a non-spoken answer to the polling question via selection of the second GUI element associated with the spoken answer provided by the first participant.

14. The method of claim 13, further comprising: detecting that the second participant has provided at least one of a spoken answer or a non-spoken answer to the polling question; and modifying the GUI provided to a third participant of the plurality of participants to present, via the second portion of the GUI, the textual form of the spoken answer provided by the second participant and an additional GUI element that allows the third participant to provide an additional non-spoken answer to the polling question via selection of the additional GUI element associated with the at least one of the spoken answer or the non-spoken answer provided by the second participant.

15. The method of claim 13 or 14, wherein, in response to determining that the spoken answer provided by the first participant satisfies an answer criterion, modifying the GUI provided to the second participant of the plurality of participants to present the second portion.

16. The method of claim 15, wherein, determining that the spoken answer provided by the first participant satisfies the answer criterion comprises: determining that the spoken answer provided by the first participant corresponds to a particular set of answers to the polling question, wherein each particular answer of the particular set of answers is provided by a respective participant of the plurality of participants; and computing a total number of the particular answers provided to the polling question based on the particular set of answers and the spoken answer provided by the first participant, wherein the answer criterion is determined to be satisfied in response to the total number of the particular answers exceeding a threshold number of answers.

17. The method of claim 13, wherein, detecting that the first participant has provided the spoken answer to the poll question comprises: generating an audio file comprising one or more spoken phrases provided by the first participant; converting content of the audio file to one or more text strings; and identifying the text form of the spoken answer from the one or more text strings.

18. The method of claim 13, further comprising: determining one or more answer candidates for the poll question based on the spoken answer provided by the first participant, wherein the second portion of the GUI provided to the second participant is to further include a text form of each of the determined one or more answer candidates and one or more additional GUI elements that allow the second participant to provide a non-spoken answer to the poll question via selection of the additional GUI elements associated with the respective answer candidate.

19. The method of claim 13, further comprising: responsive to detecting that a third participant of the plurality of participants has provided at least one of an additional spoken answer or an additional non-spoken answer to the poll question, updating the modified GUI provided to the second participant to present, in a third portion, a text form of the at least one of the additional spoken answer or the non-spoken answer provided by the third participant and one or more additional GUI elements that allow the second participant to provide the non-spoken answer to the poll question via selection of the respective additional GUI elements associated with the at least one of the additional spoken answer or the additional non-spoken answer provided by the third participant.

20. A system for populating answers to a poll question, comprising: a memory; and a processing device coupled to the memory, the processing device to perform the method of any one of claims 13-19.

21. A non-transitory computer-readable storage medium comprising instructions for a server that, when executed by a processing device, cause the processing device to perform the method of any one of claims 13-19.

Citation Information

Patent Citations

  • Graphical user interface for an interactive collaboration system

    US20020085030A1

  • Administering a digital survey over voice-capable devices

    US20180232752A1

  • Systems and methods for content curation in video based communications

    US20180359530A1