Biasing meeting audio and / or generating meeting documents based on meeting content and / or related data
Biased ASR and dynamic meeting document generation address the limitations of current videoconferencing by enhancing transcription accuracy and action item identification, ensuring relevant content is captured and documented accurately.
Patent Information
- Application Number
- JP2025207212
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-02-23
- Filing Date
- 2025-11-27
- Publication Date
- 2026-02-24
AI Technical Summary
Current videoconferencing technologies struggle with accurately transcribing diverse meeting topics and generating meeting documents due to limitations in speech recognition and the inability to prioritize important discussions, leading to incomplete or inaccurate action items.
Implementing biased automatic speech recognition (ASR) based on relevant meeting documents and participant interactions, using term frequency and inverse document frequency to enhance transcription accuracy, and generating meeting summaries and action items dynamically.
Improves the accuracy of meeting document creation and action item identification by prioritizing relevant content, reducing manual input, and enabling real-time document editing and reminder generation.
Smart Images

Figure 2026031594000001_ABST
Abstract
Description
[Background technology]
[0001] As videoconferencing software becomes more available, many users are beginning to recognize some shortcomings in the current state of this technology. For example, like in-person meetings, videoconferencing can involve participants discussing a variety of topics. Some participants may manually take notes during the meeting, and these notes may later be referenced by specific participants who may be tasked with completing action item(s) identified during the meeting. When participants take notes during a meeting in this manner, they may miss specific topics discussed during the meeting, and as a result, some action items may not be accurately addressed. While some applications may function to provide a meeting transcript (e.g., a text version of most of the spoken input provided during the meeting), such a transcript may not reflect that certain topics are more important than others (e.g., a discussion of "lunch orders" versus a discussion of an experiment conducted by meeting participants). Thus, relying on some available transcription applications may not provide any additional efficiency in streamlining the creation of related meeting documents and / or action items.
[0002] For applications that facilitate audio transcription of meetings, speech recognition may be limited in situations where a variety of subjects are discussed and / or otherwise referenced during the meeting. For example, the spoken words and phrases occurring during a meeting may be specific to an organization and / or may have been relatively recently developed by the organization. As a result, typical speech recognition applications may not accurately transcribe words and / or phrases that have recently been adopted into the vocabulary of a particular industry and / or organization. Therefore, meeting participants who rely on such transcripts may not be working with an accurate transcription when attempting to complete specific action items from the meeting. Summary of the Invention
[0003] Embodiments described herein relate to techniques for automating certain aspects of group meetings involving multiple users. Such aspects may include, for example, note-taking and / or biasing automatic speech recognition (ASR) based on relevant meeting documents and / or other content to generate meeting notes. Such aspects may additionally or alternatively include, for example, generating a meeting summary based on a meeting transcript, notes taken by at least one participant during the meeting, visual and / or audio cues captured during the meeting (with prior permission from the participant(s)), and / or explicit requests from participants in the meeting regarding content to include in the summary. Such aspects may further additionally or alternatively include generating action items from the meeting so that those action items can be linked to specific portions of the summary and / or used to create reminders that can be rendered for specific participants based on conditions that can be determined from the content of the meeting.
[0004] In some embodiments, the ASR may be biased toward transcribing content spoken by participants during a meeting to facilitate more accurate creation of meeting notes and / or other content. The ASR may be biased using content from documents and / or other files that may be selected by the conferencing application and / or other applications (e.g., automated assistant applications) based on the relevance of each particular document to the meeting. For example, an upcoming meeting may have M total participants (e.g., M=10), and a subset N of the M total participants (e.g., N=3) may have accessed and / or shared one or more documents before the meeting. Document(s) may be determined to be associated with the upcoming meeting based on at least a threshold number or percentage of participants (or invitees) accessing the document(s) before and / or during the meeting and / or based on a determination that the content of the document(s) is associated with other content in the meeting invitation for the upcoming meeting (e.g., a determination that the content of the document is similar to the title of the meeting invitation).
[0005] Once a document is determined to be relevant to the upcoming meeting, the ASR used to transcribe audio during the upcoming meeting can be biased according to the document's content. For example, a document accessed by some of the participants in the upcoming meeting may contain multiple instances of the term "cardinal," which may refer to a product being discussed at the upcoming meeting. During the upcoming meeting, participants may frequently speak the term "cardinal" in connection with the product, and using the ASR to process audio data generated during the meeting, candidate interpretations of the audio embodying the spoken term "cardinal" can be identified. For example, the candidate interpretations may include "garden hole," "card in a," "guard the ball," and "cardinal." Each candidate interpretation can be assigned a score based on a variety of different factors (e.g., the interpretation's relevance to other recent audio, context, location, etc.), and the candidate interpretation with the highest score can be incorporated into the meeting document generated by the conferencing application and / or other applications. However, in some implementations, one or more respective scores may be weighted according to whether their corresponding candidate interpretations are associated with a document(s) accessed by some of the participants. For example, the score of a candidate interpretation "cardinal" may be increased and / or the scores of other candidate interpretations may be decreased based on the explicit appearance of the term "cardinal" in a document(s) accessed by some of the meeting participants.
[0006] In some implementations, the score may be based on the term frequency (TF) of the term “cardinal” appearing in multiple different documents determined to be associated with the conference and / or one or more participants, and / or the inverse document frequency (IDF) of the term “cardinal” in another corpus, such as a global corpus of Internet documents and / or a corpus of training instances used in the training model(s) used in the ASR. In other words, a term's score may be generated such that it is more heavily biased against a term if it has a higher TF and / or lower IDF, as opposed to a lower TF and / or higher IDF. For example, a given term may be more heavily biased if it appears frequently in documents determined to be associated with the conference and / or if it is not included in the training example(s) used to train any (or very few) of the ASR model(s). These and other methods can improve ASR of speech during a conference and further enable automated note-taking and / or other functions that rely on ASR results to perform more accurately. This may conserve computational resources by reducing the number of inputs participants need to manually provide to each device to edit incorrect ASR entries. This may also encourage more participants to actively rely on ASR-based functions over other manual functions, which may distract participants from engaging with other conference participants during the conference.
[0007] In some embodiments, the conferencing application and / or other applications may additionally or alternatively generate a summary, action item(s), and / or other content for the meeting based on various features of the meeting and / or interactions that may occur during the meeting. Such features may include participant note-taking, audio from one or more participants, direct and / or indirect requests from participants, visual content from the meeting, gestures from one or more participants, and / or any other feature that may indicate content for incorporation into the summary. In some embodiments, portions of the content included in the summary may be generated in response to utterances from multiple different participants during a portion of the meeting related to a particular topic. For example, after a participant has spoken for a period of time, multiple other participants may provide feedback regarding a particular topic that the one participant mentioned. As a result, topic-based summary items may be generated for the summary document to facilitate creating a summary that includes content that is important to more meeting participants.
[0008] In some embodiments, summary items for a meeting summary may be automatically generated based on associations between the content of the meeting (e.g., audio from one or more participants) and other content related to the meeting (e.g., the title of the meeting invitation, the content of attachments to the meeting invitation, and the content of files accessed by meeting participants before, during, and / or after the meeting). For example, the title of the meeting invitation may be "Meeting on Phase II Cellular Trials," and attachments provided with the meeting invitation may include a spreadsheet of clinical trial data. At the start of the meeting, a participant may ask, "How was everyone's weekend?" and other participants may respond by briefly describing their own weekend (e.g., "It was great. We went to a concert on the waterfront."). However, because terms such as "weekend," "concert," and "waterfront" do not appear in the meeting title or the meeting attachments, a summary may be generated that does not mention all content from this portion of the meeting.
[0009] Following the foregoing example, somewhere during the "Meeting on Phase II Cellular Trials," while a second participant is speaking about "Batch T Results," a first participant may raise their hand (actually and / or virtually via a "Raise Hand" interface element) and make a request such as, "Bill, I don't think the Batch T results are complete. Can you check them after the meeting?" Image data embodying the first participant's raising of their hand may be captured by a video camera (with prior permission from the participant(s)) and processed using one or more trained machine learning models and / or one or more heuristic processes. Alternatively, or additionally, the content of the spoken request from the first participant may be captured as audio data (with prior permission from the first participant) and processed using one or more trained machine learning models and / or one or more heuristic processes. Based on these processes, the conference application and / or other applications may generate summary items to be incorporated into a summary being generated for the conference. For example, language processing may be used to determine that a term in a conference title (e.g., "...Phase II Cellular Trial...") may often be associated with a term such as "results." Based on this determination, content from a request from a first participant to a second participant may be ranked for inclusion in the summary above other conference content (e.g., "How was everyone's weekend?") that may not be deemed relevant enough to incorporate into the summary.
[0010] In some embodiments, a summary item may be incorporated into a meeting summary document based on a threshold number (N) of individuals determined (with prior permission from participants) to be taking notes on a topic during the meeting. A summary item may then be generated to address the discussion of that particular topic. Alternatively, or additionally, the participant(s)' attention level(s) regarding a particular topic may be determined (with prior permission from participants), and / or changes in the participant(s)' attention level(s) during the discussion of the particular topic may be determined. Based on an increase or change in attention level during the discussion of the particular topic, the particular topic may be the subject of a summary item to be included in the summary document or other document related to the meeting. In some embodiments, determining participants' attention levels may be performed using one or more cameras, with prior permission from participants, during a meeting that is an in-person meeting, a virtual video conference (e.g., when all participants connect to the meeting via an internet or other network connection), and / or any meeting with a combination of remote and in-person participants.
[0011] In some embodiments, a summary item generated by a conference application and / or an automated assistant may be designated an “action item” based on at least the corresponding conference content (e.g., the content that served as the basis for the action item) and / or the context in which the corresponding conference content was presented. Following the example above, content provided by a first participant to a second participant (e.g., “Bill”) (e.g., “…I don’t think the Batch T results are complete. Can you check after the meeting?”) may be incorporated into the conference summary as an action item for the second participant. The action item may be included in the conference summary along with embedded links to any files the action item may reference (e.g., a “Batch T Results” document) and / or along with reminders for the second and / or first participants. In some embodiments, a reminder may be rendered to the first participant, the second participant, and / or any other party in response to one or more conditions being met. The conditions may be selected, for example, based on the conference content and / or context. For example, the second participant may receive notification about the action item after the meeting in response to the second participant accessing the "Batch T Results" document. Alternatively or additionally, the first participant and / or the second participant may receive notification about the action item after and / or during another meeting to which the first and second participants are invitees. Alternatively or additionally, the first participant and / or the second participant may receive notification about the action item in response to receiving and / or sending a message to another attendee (e.g., a third participant) of the meeting from which the action item was derived.
[0012] In some embodiments, the meeting summary is non-static and / or generated in real time as the meeting is progressing, thereby allowing participants to verify that particular items are included in the summary and / or further modify the summary before the meeting concludes. For example, an action item automatically included in the meeting summary may be modified by a participant to target one or more additional participants not originally targeted by the action item. Alternatively, or additionally, a summary of discussed topics automatically included in the meeting summary before the meeting ends (e.g., "Batch B Results") may be editable to add, remove, modify, and / or otherwise change the topic summary. In some embodiments, portions of the automatically generated summary may be automatically edited, for example, when a particular topic is revisited during the meeting, when additional contextual data becomes available, when additional content (e.g., meeting attachments, documents, files, etc.) becomes available and / or is otherwise accessed by participants, and / or when additional meeting information becomes otherwise available.
[0013] The above description is provided as a summary of some embodiments of the present disclosure. Details of these and other embodiments are described in more detail below.
[0014] Other embodiments may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., central processing unit (CPU)(s), graphics processing unit (GPU)(s), and / or tensor processing unit (TPU)(s)) to perform methods such as one or more of the methods described above and / or elsewhere herein. Still other embodiments may include systems of one or more computers including one or more processors operable to execute stored instructions to perform methods such as one or more of the methods described above and / or elsewhere herein.
[0015] It should be understood that all combinations of the foregoing concepts, and additional concepts described in more detail herein, are contemplated as being part of the presently disclosed subject matter, for example, all combinations of claimed subject matter listed at the end of this disclosure are contemplated as being part of the presently disclosed subject matter. [Brief explanation of the drawings]
[0016] [Figure 1A] 1 shows a view of the audio biasing and document generation being performed for a conference based on data created before and / or during the conference. [Figure 1B] 1 shows a view of the audio biasing and document generation being performed for a conference based on data created before and / or during the conference. [Figure 1C] 1 shows a view of the audio biasing and document generation being performed for a conference based on data created before and / or during the conference. [Figure 1D] 1 shows a view of the audio biasing and document generation being performed for a conference based on data created before and / or during the conference. [Figure 1E] 1 shows a view of the audio biasing and document generation being performed for a conference based on data created before and / or during the conference. [Figure 2] A system is shown that provides applications, such as automated assistants and / or conferencing applications, that can bias ASR and / or generate meeting documents based on data generated before and / or during the meeting. [Figure 3] A method is shown for biasing automatic speech recognition according to instances of data determined to be relevant to the meeting and / or meeting participants before and / or during the meeting. [Figure 4]We present a method for automatically incorporating specific content into a meeting document to facilitate generating a meeting summary and / or other types of documents based on the content of the meeting. [Figure 5] We show how to generate action items based on natural language content provided during a participant meeting, remind specific participants of action items, and / or designate action items as completed based on specific conditions. [Figure 6] 1 illustrates an exemplary computer system. DETAILED DESCRIPTION OF THE INVENTION
[0017] 1A, 1B, 1C, 1D, and 1E illustrate views 100, 120, 140, 160, and 180, respectively, of audio biasing and document generation being performed for a meeting based on data created before and / or during the meeting. Such operations can be performed to minimize the number of inputs a user must manually enter during a meeting, thereby streamlining certain types of meetings and conserving the computational resources of meeting-related devices. Additionally, the accuracy of certain meeting documents can be improved using certain processes for generating meeting summary documents and / or meeting action items.
[0018] Prior to the meeting, as shown in FIG. 1A , a first user 102 (e.g., invitee and / or participant) may access, via their respective computing device 104, a meeting invitation 114 that was communicated to the first user 102 by a conferencing application 106. The meeting invitation 114 may have a title that is rendered in the application interface 108 of the conferencing application 106, and the title may include terms that can serve as a basis for biasing ASR during the meeting. Alternatively, or additionally, the content of the meeting invitation 114 (i.e., the first document) may include words and / or phrases (i.e., terms) that can indicate whether a particular topic to be discussed during the meeting is sufficiently relevant to be included in the meeting document and / or designated as an action item. Other users who receive the meeting invitation 114, such as a second user 110 operating an additional computing device 112, can also influence whether particular data is considered relevant to the meeting.
[0019] For example, second user 110 may be viewing collaborative spreadsheet 122 via their respective computing device 112, as shown in view 120 of FIG. 1B . Collaborative spreadsheet 122 (i.e., the second document) may be accessible to one or more meeting participants and, therefore, may be considered relevant to the meeting by the meeting application and / or other supporting applications. For example, collaborative spreadsheet 122 may be a cloud-based document created / owned by one of the meeting invitees and shared with all other meeting invitees, thereby making it accessible to all of the meeting invitees. Collaborative spreadsheet 122 may be determined to be relevant to the meeting based on being shared with at least a threshold number or a threshold percentage of the meeting invitees (e.g., all) and / or based on whether the document was presented to one or more participants during the meeting. Optionally, determining that collaborative spreadsheet 122 is relevant may also be based on determining that collaborative spreadsheet 122 has been shared with less than a threshold number or percentage of individuals who are not meeting invitees. For example, a first document shared only with all meeting invitees may be determined to be relevant to the meeting, while a second document shared with all meeting invitees but also with N additional individuals (e.g., 50 additional individuals) who are not meeting invitees may be determined to be not relevant to the meeting. Alternatively, or additionally, a collaborative spreadsheet 122 may be deemed relevant to the meeting by the conferencing application based on the content of the collaborative spreadsheet 122. For example, the collaborative spreadsheet 122 may include content related to the content of the meeting invitation 114, thereby indicating that the collaborative spreadsheet 122 is relevant to the meeting. The content of the collaborative spreadsheet 122 may include, for example, the price of ingredients 124 for "hummus," and "hummus" may be a term mentioned in the meeting invitation (e.g., in the title of the meeting invitation and / or the description section of the meeting invitation).Based on this correspondence, various portions of the document content may be processed (e.g., using inverse document frequency and / or other document review processes) to identify portions that may be relevant to ASR bias and / or other process(es) to be performed during the meeting. These portions may then be utilized as a basis for biasing the ASR during the meeting, as a basis for identifying relevant content (e.g., input from participants) during the meeting, and / or as a basis for identifying conditions for action item reminders and / or conditions for action item fulfillment.
[0020] In some embodiments, the conferencing application, computing device, and / or server device 142 may determine that a conference has started (e.g., as shown in FIG. 1C ) based on data available from one or more devices and / or applications. For example, calendar data and / or data from the conferencing application may be utilized to determine that the conference has started and / or that one or more people have joined the conference. Once the conference has started, data 146 from various devices may be processed to facilitate biasing the ASR and generating a conference document. This conference document may include the meeting content 144 and may not include any data unrelated to the conference (e.g., other topics such as “small talk” during conference breaks). In some cases, the data 146 may include audio data embodying voices from various participants in the conference.
[0021] For example, the data 146 may characterize spoken words 148 provided by a third user 150, such as "Let's think of some peppers to add to the hummus." The data 146 may be processed using ASR biased according to instances of data identified as relevant to the meeting. For example, a meeting invitation 114 having a title including the specific term "hummus" may have one or more candidate terms, such as "hummus" and / or "recipe," assigned a higher probability and / or weight value to a transcription of the spoken words 148 than other words and / or phrases that may be pronounced similarly (e.g., "honeys" for "hummus" and "rest in peace" for "recipe"). Alternatively, or additionally, the resulting transcription of the spoken words 148 may be processed to determine whether the content of the transcription is sufficiently relevant to the meeting to be included in the meeting content 144 for the meeting document. For example, if the content of the transcription includes the terms "hummus" and "recipe," and data accessed before and / or during the meeting includes the terms "hummus" and "recipe," the content of the transcription may be deemed relevant enough to be incorporated into the meeting document. Following the example above, because the second user 110 was viewing hummus ingredients before the meeting and the meeting invitation 114 includes "hummus recipe" in the title, the content of the transcription of the spoken words 148 may be deemed relevant enough to be included in the meeting content 144 and / or the meeting document.
[0022] In some implementations, non-verbal gestures and / or other non-verbal cue(s) captured by one or more sensors during the conference may be utilized to determine relevance of participant input to the conference. For example, in response to spoken words 148 from the third user 150, the first user 102 may provide a separate spoken word 162, such as, "Yes, I'll create a list and send it to Jeff for pricing." The first user 102 may also perform a non-verbal gesture 164 while providing the spoken words 162 and / or within a threshold duration of providing the spoken words 162, that may indicate the importance of what they are saying. With prior permission from the participant(s), audio and image data captured by the camera 156 and computing device 154 may be processed by the local computing device and / or server device 142 to determine whether to incorporate a response from the first user 102 into the conference document. Additionally or alternatively, the data may be processed to determine whether to generate an action item 166 based on one or more conditions of the spoken words 162 and / or the action item 166 .
[0023] For example, the audio and / or video data may be processed using one or more heuristic processes and / or one or more trained machine learning models to determine whether a text entry should be included in the conference document. In some embodiments, this determination may be based on whether the spoken words 162 were provided within a threshold duration during which the third user 150 provided the spoken words 148. Alternatively, or additionally, the determination of whether a text entry should be included in the conference document may be based on whether the spoken words 162 are responsive to the conference-related input (e.g., the spoken words 148) and / or whether the spoken words 162 are directed to the person who provided the conference-related input. In some embodiments, a score may be assigned to the text entry according to one or more of these and / or other factors, and the score may be compared to a score threshold. If the score meets the score threshold, the text entry may be incorporated into the conference document (e.g., a conference “summary” document).
[0024] Once a text entry is determined to be incorporated into a meeting document, a determination can be made as to whether the text entry is an action item, and if so, whether the action item should have a condition. For example, a text entry corresponding to spoken words 162 can be designated as an action item based on the first user 102 expressing an intention to take an action (e.g., "make a list"). Alternatively, or additionally, the action item can be assigned one or more conditions based on the content of the text entry and / or the context in which the spoken words 162 were provided. For example, the action item can be stored with conditional data characterizing a reminder that can be rendered the next time the first user 102 communicates with the second user 110 (e.g., "Jeff"). Alternatively, or additionally, the action item can be stored with conditional data indicating that the action item will be fulfilled when the first user 102 communicates "Pepper's list" to the second user 110. In this way, not only can action items be incorporated into meeting documents to accurately track action items, but reminders can also be established and / or action items can be automatically updated based on user actions (with prior permission from the user(s)).
[0025] FIG. 1E shows a view 180 of a summary document 182 that may be automatically created by the conferencing application and / or other applications based on the content of the meeting and / or data related to the meeting. For example, the summary document 182 may include a list of summary items summarizing various topics discussed during the meeting and action items identified during the meeting. In some embodiments, the summary items may embody terms that may or may not have been explicitly stated during the meeting, either verbally or in writing. For example, a summary item may include "The group agreed that ingredients in hummus should include pepper," which may be a statement that was not explicitly stated in those terms during the meeting. Alternatively, or additionally, the summary document may be generated using a list of action items identified during the meeting. In some embodiments, the action items may be generated to include reminders (e.g., reminders before the next meeting) that may be rendered to particular participants when certain conditions are met. Alternatively, or in addition, the action item may include embedded links to specific data (e.g., documents, websites, images, contact information, and / or any other data), such as a participant's electronic address (e.g., "@Jeff") and / or a specific reminder (e.g., a reminder before the next meeting). In some embodiments, the summary document 182 may be viewed during the meeting, thereby allowing participants to edit the meeting document 182 as it is being created. Alternatively, or in addition, the summary document 182 may have editable embedded data, such as allowing participants to edit when specific reminders are rendered and / or whether an action item is still outstanding.
[0026] FIG. 2 illustrates a system 200 that provides an application, such as an automated assistant and / or a conferencing application, that can bias ASR and / or generate conference documents based on data created before and / or during a meeting. Automated assistant 204 may operate as part of an assistant application provided on one or more computing devices, such as computing device 202 and / or a server device. A user may interact with automated assistant 204 through assistant interface(s) 220, which may be a microphone, camera, touchscreen display, user interface, and / or any other device capable of providing an interface between one or more users and an application. For example, a user may initialize automated assistant 204 by providing verbal, text, and / or graphical input to assistant interface 220 to cause automated assistant 204 to initiate one or more actions (e.g., providing data, controlling peripheral devices, accessing agents, generating input and / or output, etc.). Alternatively, automated assistant 204 may be initialized based on processing of contextual data 236 using one or more trained machine learning models.
[0027] The context data 236 may characterize one or more features of an environment accessible to the automated assistant 204 and / or one or more features of a user who is predicted to intend to interact with the automated assistant 204. The computing device 202 may include a display device, which may be a display panel including a touch interface for receiving touch input and / or gestures to enable a user to control the application 234 of the computing device 202 via the touch interface. In some embodiments, the computing device 202 may lack a display device and thus provide an audible user interface output without providing a graphical user interface output. Additionally, the computing device 202 may provide a user interface, such as a microphone, to receive spoken natural language content from a user. In some embodiments, the computing device 202 may include a touch interface and may not have a camera, but may optionally include one or more other sensors.
[0028] The computing device 202 and / or other third-party client devices may communicate with the server device over a network such as the Internet. Additionally, the computing device 202 and any other computing devices may communicate with each other over a local area network (LAN) such as a Wi-Fi network. The computing device 202 may offload computing tasks to a server device to conserve computing resources at the computing device 202. For example, the server device may host the automated assistant 204, and / or the computing device 202 may send inputs received at one or more assistant interfaces 220 to the server device. However, in some embodiments, the automated assistant 204 (e.g., a conferencing application) may be hosted on the computing device 202, and various processes that may be associated with automated assistant operation may be executed on the computing device 202.
[0029] In various embodiments, all or less than all aspects of automated assistant 204 may be implemented on computing device 202. In some of these embodiments, aspects of automated assistant 204 are implemented via computing device 202 and may interface with a server device that may implement other aspects of automated assistant 204. The server device may optionally provide services to multiple users and their associated assistant applications via multiple threads. In embodiments in which all or less than all aspects of automated assistant 204 are implemented via computing device 202, automated assistant 204 may be an application separate from (e.g., installed "on top of") the operating system of computing device 202, or alternatively, may be implemented directly by (e.g., considered an operating system application but integral to) the operating system of computing device 202.
[0030] In some embodiments, the automated assistant 204 can include an input processing engine 206 that can use multiple different modules to process input and / or output of the computing device 202 and / or the server device. For example, the input processing engine 206 can include a speech processing engine 208 that can process audio data received at the assistant interface 220 to identify text embodied in the audio data. The audio data can be transmitted from the computing device 202 to a server device, for example, to conserve computational resources at the computing device 202. Additionally or alternatively, the audio data can be processed exclusively at the computing device 202.
[0031] The process for converting audio data to text may include speech recognition algorithms that may use neural networks and / or statistical models to identify groups of audio data that correspond to words or phrases. The text converted from the audio data may be analyzed by data analysis engine 210 and made available to automated assistant 204 as text data that may be used to generate and / or identify command phrase(s), intent(s), action(s), slot value(s), and / or any other content specified by the user. In some embodiments, output data provided by data analysis engine 210 may be provided to parameter engine 212 to determine whether the user has provided input corresponding to a particular intent, action, and / or routine executable by automated assistant 204 and / or an application or agent accessible via automated assistant 204. For example, assistant data 238 may be stored on server device and / or computing device 202 and may include data defining one or more actions that automated assistant 204 can perform, as well as parameters required to execute the action. The parameter engine 212 may generate one or more parameters for the intent, action, and / or slot value and provide the one or more parameters to the output generation engine 214. The output generation engine 214 may use the one or more parameters to communicate with the assistant interface 220 to provide output to the user and / or to communicate with one or more applications 234 to provide output to the one or more applications 234.
[0032] In some embodiments, the automated assistant 204 may be an application that can be installed “on top of” the operating system of the computing device 202 and / or that itself may form part of (or the entire) the operating system of the computing device 202. The automated assistant application may include and / or access on-device speech recognition, on-device natural language understanding, and on-device fulfillment. For example, on-device speech recognition may be performed using an on-device speech recognition module that processes audio data (detected by the microphone(s)) using an end-to-end speech recognition machine learning model stored locally on the computing device 202. The on-device speech recognition generates recognized text for spoken words present in the audio data (if any). In some embodiments, the speech recognition may be biased according to the operation of the ASR bias engine 218, which may actively bias particular instances of audio according to available data prior to and / or during capture of the audio. Also, for example, on-device natural language understanding (NLU) may be performed using an on-device NLU module that processes recognized text generated using on-device speech recognition and optionally processes contextual data to generate NLU data.
[0033] The NLU data may include intent(s) corresponding to the spoken words and, optionally, parameters (e.g., slot values) of the intent(s). On-device fulfillment may be performed using an on-device fulfillment module that utilizes the NLU data (from the on-device NLU) and, optionally, other local data, to determine action(s) to take to resolve the intent(s) of the spoken words (and, optionally, parameter(s) for the intent(s)). This may include determining local and / or remote responses (e.g., replies) to the spoken words, interaction(s) with locally installed application(s) to perform based on the spoken words, command(s) to send to Internet of Things (IoT) device(s) (directly or via corresponding remote system(s)) based on the spoken words, and / or other resolution action(s) to perform based on the spoken words. The on-device fulfillment may then initiate local and / or remote implementation / execution of the determined action(s) to resolve the spoken words.
[0034] In various implementations, remote speech processing, remote NLU, and / or remote fulfillment may be utilized, at least selectively. For example, recognized text may be at least selectively sent to a remote automated assistant component(s) for remote NLU and / or remote fulfillment. For example, recognized text may optionally be sent for remote performance in parallel with on-device performance or in response to a failure of on-device NLU and / or on-device fulfillment. However, on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device execution may be prioritized due to the reduced latency they offer, at least when resolving spoken words (due to the absence of client-server round-trip(s) to resolve spoken words). Furthermore, on-device functionality may be the only functionality available in situations where network connectivity is absent or limited.
[0035] In some embodiments, computing device 202 may include one or more applications 234 that may be provided by a third-party entity different from the entity that provided computing device 202 and / or automated assistant 204. An application state engine of automated assistant 204 and / or computing device 202 may access application data 230 to determine one or more actions that may be performed by one or more applications 234, as well as the state of each application of one or more applications 234 and / or the state of each device associated with computing device 202. A device state engine of automated assistant 204 and / or computing device 202 may access device data 232 to determine one or more actions that computing device 202 and / or one or more devices associated with computing device 202 may perform. Additionally, application data 230 and / or any other data (e.g., device data 232) can be accessed by automated assistant 204 to generate context data 236 that can characterize the context in which a particular application 234 and / or device is running, and / or the context in which a particular user is accessing computing device 202, application 234, and / or any other device or module.
[0036] While one or more applications 234 are executing on the computing device 202, the device data 232 may characterize the current operational state of each application 234 executing on the computing device 202. Additionally, the application data 230 may characterize one or more features of the executing applications 234, such as the content of one or more graphical user interfaces rendered at the direction of the one or more applications 234. Alternatively, or additionally, the application data 230 may characterize action schemas, which may be updated by the respective applications and / or by the automated assistant 204 based on the respective applications' current operational state. Alternatively, or additionally, one or more action schemas for one or more applications 234 may remain static but may be accessed by the application state engine to determine suitable actions to initialize via the automated assistant 204.
[0037] Computing device 202 may further include assistant invocation engine 222, which may use one or more trained machine learning models to process application data 230, device data 232, context data 236, and / or any other data accessible to computing device 202. Assistant invocation engine 222 may process this data to determine whether to wait for the user to explicitly speak an invocation phrase to invoke automated assistant 204, or whether to consider the data as indicative of the user's intent to invoke the automated assistant, instead of requiring the user to explicitly speak an invocation phrase. For example, one or more trained machine learning models may be trained using instances of training data based on scenarios in which a user is in an environment where multiple devices and / or applications exhibit various operating states. The instances of training data may be generated to capture training data that characterize contexts in which the user invokes the automated assistant and other contexts in which the user does not invoke the automated assistant. When one or more trained machine learning models are trained according to these instances of training data, assistant invocation engine 222 may cause automated assistant 204 to detect, or be limited in detecting, a voice input from the user based on features of the context and / or environment.
[0038] In some embodiments, the system 200 may include a relevance data engine 216 that can process data from various sources to determine whether the data is relevant to an upcoming meeting and / or other gathering of one or more people. For example, the relevance data engine 216 may utilize one or more heuristic processes and / or one or more trained machine learning models to process data to determine whether a meeting is expected to occur or is occurring. Based on this determination, the relevance data engine 216 may process data from various sources (e.g., various devices, applications, servers, and / or any other source that can provide data related to meetings) to determine whether the data is relevant to a particular meeting. In some embodiments, the relevance of the data may be characterized by a metric (i.e., a score) that can be compared to a relevance threshold. If the metric meets the relevance threshold, the data may be considered relevant to the meeting. For example, one or more trained machine learning models may be utilized to generate embeddings from data that may be relevant to a meeting. For example, the trained machine learning model(s) may include Word2Vec, BERT, and / or other model(s) that can be used to process data (e.g., text data) and generate semantically meaningful, reduced-dimensional embeddings in a latent space. The embeddings may be in a latent space, and the distance between the embeddings and the meeting embeddings (also mapped to the latent space) may be characterized by a metric. The metric may be compared to a relevance threshold in determining whether the data is relevant to the meeting (e.g., if the distance is closer than a threshold, the data may be determined to be relevant). The meeting embedding may be generated based on processing one or more features of the meeting with one or more trained machine learning models (e.g., models used to generate the data embeddings).For example, the meeting characteristic(s) may include the title of the meeting, a description or notes included in the meeting invitation, the time of the meeting, the time the meeting invitation was scheduled, the number of participants in the meeting, and / or any other characteristics associated with the meeting.
[0039] In some embodiments, once one or more instances of data related to the conference are identified, the ASR bias engine 218 may process the one or more instances of data to determine whether to bias the ASR based on the content of the data. For example, terms within the content of the data may be identified using one or more heuristic processes and / or one or more trained machine learning models. In some embodiments, an inverse document frequency (IDF) measure of a term within the data may be identified and utilized to determine whether the ASR should be biased toward that particular term. For example, the IDF measure may be based on the frequency of the term in the spoken language utilized to train the ASR model(s) utilized in the ASR. For example, a first term within the data, such as "garbanzo," may be selected for use in biasing the ASR based on having a high IDF measure (e.g., very few utterances containing "garbanzo" are used in training the ASR model(s)). Meanwhile, a second term within the data, such as "oil," may not be selected for use in biasing the ASR based on having a high IDF measure. Furthermore, in some embodiments, the degree of bias for a term may be a function of its IDF measure and / or term frequency (TF) measure (e.g., a function of how frequently the term appears in the data). Additional and / or alternative techniques may be utilized in determining whether a term is important to the data and / or meeting instance. For example, where the term is used in the document (e.g., its use in the title, first sentence, and / or conclusion may be more important than its use “in the middle” of the document), and / or whether the term is also used in the meeting invitation. Once a particular term is determined to be important to the data and / or meeting instance, the particular term may be utilized in biasing the ASR during the meeting. For example, the particular term may be weighted and / or assigned a higher value score or probability during the ASR of input spoken during the meeting.
[0040] In some embodiments, the system 200 can include a document input engine 226, which can automatically generate meeting document(s) using data generated by the ASR bias engine 218 and / or the input processing engine 206. The meeting document can be generated, for example, to represent a summary of the meeting and thus can describe points relevant to the meeting while omitting content from the meeting discussion that may not be relevant to the meeting. In some embodiments, embeddings can be generated for text and / or audio data transcribed during the meeting, and the embeddings can be mapped to a latent space that can also include an embedding of the meeting. If the embedding distance between the embedding and the meeting embedding is determined to meet a threshold for incorporating the text entry into the meeting document, the text entry can be incorporated into the meeting document. In some embodiments, candidate text entries can be rendered in an interface during the meeting, and participants and / or others can choose to incorporate the text entry into the meeting document regardless of whether the embedding distance meets the threshold.
[0041] In some embodiments, the system 200 can include an action item engine 224, which can determine whether a text entry should be designated as an action item in a meeting document and / or other data and / or whether the action item should have certain conditions. For example, the action item engine 224 can utilize one or more heuristic processes and / or one or more trained machine learning models to determine whether input from meeting participants should be considered an action item. For example, an action item can refer to input that describes a task to be completed by at least one participant, or others, following input provided during the meeting. Thus, an input that embodies a request to another participant and may optionally have a deadline (e.g., "Let's follow up on the budget at the next meeting") can be assigned a higher action item score than other input that does not embody a request or deadline (e.g., "Did you enjoy your lunch?").
[0042] In some embodiments, the content of the input underlying the action item, along with any other relevant data, may be processed to identify conditions to be stored in association with the action item input. For example, reminder conditions and / or fulfillment conditions may be generated based on the content of text entries and / or other inputs, and / or any data associated with the input during the meeting. In some embodiments, a condition explicitly provided in the input (e.g., "Please send the report after the meeting") may be processed to generate a condition that can be stored with the action item (e.g., actionItem("Send report", nextMeetingTime(), reminderEmail())). Alternatively, or additionally, conditions inferred from data associated with the input (e.g., documents accessed during the meeting) can be used to generate the action item condition (e.g., actionItem("Send report", nextMeetingTime(), fulfillmentCondition(email, "report," "budget," jeff@email.com))). In this way, action items can be generated automatically during a meeting without requiring manual user input, which can interfere with a user's participation in the meeting and waste computational resources on specific computing devices and each user's interface.
[0043] FIG. 3 illustrates a method 300 for biasing automatic speech recognition according to instances of data determined to be related to a conference and / or conference participants before and / or during a conference. Method 300 may be performed by one or more computing devices, applications, and / or any other device or module that may be associated with an automated assistant. Method 300 may include operation 302 of determining whether a conference is occurring or is expected to occur. The determination at operation 302 may be performed by an application, such as a conferencing application and / or an automated assistant application, accessible via a computing device (e.g., a server device, a portable computing device, etc.). The determination at operation 302 may be performed to facilitate a user associated with the application determining whether to participate in a conference in which one or more participants may communicate information to one or more other participants. In some embodiments, the determination may be based on data accessible to the application, such as contextual data (e.g., a schedule stored by the application) and / or other application data (e.g., a conference invitation provided to multiple invitees).
[0044] Method 300 may proceed from operation 302 to operation 304, which may include determining whether any instances of data related to the conference are available. The instances of data may be determined to be associated with the conference using one or more heuristic processes and / or one or more trained machine learning models. For example, data associated with one or more invitees and / or participants of a conference may be processed (with prior permission from the user) to determine whether the data is relevant to the conference. The data may include files (e.g., documents) that one or more invitees to the conference have permission to access and / or that were accessed within a threshold duration before the conference. In some embodiments, the duration may be based at least in part on the time at which at least one invitee received a conference invitation for the conference. For example, the threshold duration may be directly proportional to the amount of time between the time an invitation to the conference was initially sent or received by at least one invitee and the scheduled time of the conference. In this manner, the threshold duration before the conference during which relevant files may have been accessed may be longer for conferences that are more pre-planned. Alternatively, the threshold duration may be based on other factors, such as the duration of the meeting, the number of invitees to the meeting, the location of the meeting, and / or any other characteristic that may be specified for the meeting.
[0045] In some implementations, an instance of data may be determined to be relevant to a meeting based on the content of the instance of data compared to the content of data provided with a meeting invitation (e.g., the title of the meeting invitation, a description within the meeting invitation, attachments to the meeting invitation, etc.). For example, the content of a file may be determined to be relevant to a meeting if terms within the file are also present in and / or synonymous with terms in the meeting invitation. In some implementations, specific terms determined to be relevant to characterizing a particular file and comparing it to a meeting invitation may be identified using inverse document frequency metrics for those specific terms. Alternatively, or additionally, specific terms determined to be relevant to characterizing a particular file and comparing it to a meeting invitation may be identified using contextual data associated with the particular file.
[0046] In some embodiments, an instance of data may be determined to be relevant to a meeting based on the number of participants with whom the particular instance of data is shared and / or whether the particular instance of data includes terms or terms corresponding to content related to the meeting (e.g., a meeting invitation, audio captured during the meeting with prior permission from the participants, documents created and / or shared by the participants, etc.). For example, a document that does not include terms associated with a meeting invitation (e.g., no terms considered relevant according to IDF) but is shared with 80% of the meeting participants may not be considered relevant for purposes of ASR biasing. However, a document that includes one or more terms associated with a meeting invitation (e.g., terms considered relevant according to IDF) and is shared with 60% of the meeting participants may be considered relevant for purposes of ASR biasing. In some embodiments, the extent to which terms embodied in an instance of data are considered relevant may be based on a variety of different characteristics of the meeting. Alternatively, or additionally, the threshold for the number of participants with whom a document is shared before the document is considered relevant can be based on the number of relevant terms in the document (e.g., the percentage threshold can be inversely proportional to the degree of relevance of particular document terms).
[0047] Alternatively, or additionally, an instance of data may be deemed relevant or irrelevant based on whether a threshold percentage of participants accessed the data during the meeting. For example, a document may be deemed less relevant based on only one participant accessing the document during a large portion of the meeting (e.g., the one participant may not be paying full attention and may be distracted by content unrelated to the meeting). Rather, a document may be deemed relevant if at least a threshold percentage of participants accessed the data (e.g., document) during the meeting and / or if a threshold percentage of participants accessed the data for a threshold duration (e.g., at least a threshold percentage of the total scheduled time of the meeting). In this way, the ASR may be biased according to the terms of documents deemed relevant to the meeting during the meeting, without considering specific data that may only be relevant to individuals during the meeting.
[0048] If it is determined that the instance of data is associated with a meeting, method 300 may proceed from operation 304 to operation 306, which may include determining whether the content of the data satisfies a condition(s) for using the content as a basis for automating speech recognition biasing. In some embodiments, the content of the data may satisfy a condition for using the content as a basis for biasing automatic speech recognition when it is determined that the content embedding is a threshold distance in latent space from the meeting embedding. In other words, meeting data (e.g., meeting invitations, meeting attachments, etc.) may be processed using one or more trained machine learning models to generate meeting embeddings. Furthermore, the content of the data may also be processed using one or more trained machine learning models to generate content embeddings. Each embedding may be mapped to a latent space, and their distance in the latent space may be determined. When the distance between the embeddings meets the distance threshold, a condition for biasing automatic speech recognition based on the content of the data and / or one or more terms in the content of the data may be satisfied.
[0049] Alternatively, or in addition, the data content may satisfy a condition for biasing automatic speech recognition based on terms in the data content if characteristics of the terms in both the data content and the conference data satisfy one or more conditions. For example, if a term shared by both the data content and the conference data is determined to have a particular inverse document frequency, the condition for biasing automatic speech recognition may be satisfied. Alternatively, or in addition, the condition for biasing automatic speech recognition may be deemed satisfied if the shared term appears in a similar section of each source (e.g., title, first sentence, abstract section, etc.).
[0050] If the content of the data satisfies one or more conditions for biasing the automatic speech recognition, method 300 may proceed from operation 306 to operation 308, which may include biasing the automatic speech recognition based on the content of the instance(s) of data. Alternatively, if the content of the data does not satisfy the conditions for biasing the automatic speech recognition, method 300 may proceed from operation 306 to operation 310. Operation 308 of biasing the automatic speech recognition may be performed according to one or more different processes. For example, in some embodiments, automatic speech recognition may be performed by assigning probabilities to various hypotheses for parts of speech (e.g., words, phonemes, and / or other hypothetical parts of speech). The probabilities may then be adjusted according to whether any of the parts of speech correspond to any of the content of the data related to the meeting. For example, the probability assigned to a phoneme of a spoken term such as "Quadratic" may be increased when the term appears in an instance of the data related to the meeting. Alternatively, or additionally, the phonemes of a spoken term such as "insurance" may be assigned a higher probability than the phonemes of the term "assurance" if one or more participants write the term "insurance" into a meeting notes document during the meeting. In this manner, automatic speech recognition biasing may be performed in real time during the meeting as additional content related to the meeting is created and / or discovered.
[0051] As shown in FIG. 4 , method 300 may proceed from operation 308 to operation 310 and, optionally, via continuation element “B,” to operation 402 of method 400. Operation 310 may include determining whether conference participants and / or invitees are gathering for the conference. If it is determined that conference participants and / or invitees are gathering for the conference (e.g., based on schedule data, geolocation data, conference application data, video data, etc.), method 300 may proceed from operation 310 to operation 312. Otherwise, if participants and / or invitees have not yet gathered for the conference, method 300 may proceed from operation 310 to operation 302. Operation 312 may include determining whether any conference participants (or others associated with the conference) are accessing any instances of data during the conference. For example, the data may include memo documents being accessed by participants, portions of a conference transcript, one or more different types of media files (e.g., images, videos, etc.), and / or any other data accessible to a person. If it is determined that at least one participant is accessing an instance of the data during the conference, method 300 may return to operation 306 to further bias the automatic speech recognition according to the content of the data being accessed. Otherwise, method 300 may proceed from operation 312 to operation 302 to determine whether the conference is still in progress and / or whether another conference is expected to occur.
[0052] FIG. 4 illustrates a method 400 for automatically incorporating specific content into a meeting document to facilitate generating a meeting summary and / or other types of documents based on the content of the meeting. Method 400 may be performed by one or more applications, devices, and / or any other apparatus or module capable of interacting with meeting participants. Method 400 may include operation 402, which may optionally be a continuation of method 300, as indicated by continuation element “B” shown in FIGS. 3 and 4 . Operation 402 may include determining whether natural language content was provided by a participant (or others associated with the meeting) during the meeting. The natural language content may be, for example, spoken words from participants in a meeting (e.g., a lunch meeting, a college class, a family meal, and / or any other gathering) regarding a particular topic of the meeting. The spoken words may be, for example, “That’s a good idea. We should each think about how we can implement it in our own projects.”
[0053] If natural language content is provided by a participant in the conference, method 400 may proceed from operation 402 to operation 404. Otherwise, method 400 may optionally proceed from operation 402 via continuation element "A" to operation 302 of method 300, as shown in FIGS. 3 and 4. Operation 404 may include determining a degree of relevance of the natural language content to the conference. In some embodiments, if the natural language content is written input to an application, the degree of relevance may be based on whether one or more participants provided text entries and / or voice inputs similar to the written input. Alternatively, or additionally, if the natural language content is audio input captured by one or more audio interfaces present during the conference (e.g., a video conference in which participants each use their own laptop in a home office), the degree of relevance may be based on whether one or more other participants provided similar voice inputs and / or written inputs. For example, spoken input reflected by another participant in a written notes application (e.g., "See if you can implement Keith's idea in your project.") may be assigned a higher degree of relevance than if the other participant does not reflect the spoken input in their written notes (e.g., if a meeting participant's notes do not mention "Keith's idea").
[0054] Alternatively, or in addition, the degree of relevance assigned to natural language content provided by a participant may be based on whether the natural language content is associated with any conference documents and / or other instances of data related to the conference. For example, natural language content embodying terms contained in the title and / or other portions of a conference invitation may be assigned a higher degree of relevance than other natural language content that does not specifically have any other terms related to the conference. Alternatively, or in addition, natural language content embodying terms contained in other data related to the conference (e.g., messages between invitees, media accessed by invitees, places visited by invitees, and / or any other relevant data with prior permission from participant(s)) may be assigned a higher degree of relevance than other natural language content that does not embody such terms.
[0055] Method 400 may proceed from operation 404 to operation 406, which may include determining whether the degree of relevance satisfies a threshold for incorporating text entries characterizing the natural language content into a conference document (e.g., an automatically generated conference summary document). If the degree of relevance assigned to the natural language content meets the threshold, method 400 may proceed to operation 408. Otherwise, method 400 may return to operation 402. In some embodiments, the threshold for incorporating text entries may be based on one or more inputs from one or more participants. Alternatively, or additionally, the threshold may be based on the number of people participating in the conference, the frequency of inputs from users during the conference, the volume of content provided during the conference (e.g., number of words, phrases, pages, etc.), the location of the conference, the conference format (e.g., video, in-person, audio-only, etc.), and / or any other characteristic of the conference.
[0056] Act 408 may include incorporating and / or modifying the text entry into the meeting document. In some cases, the text entry may be incorporated into the meeting document to summarize spoken input and / or gestures from one or more participants for later reference. In this manner, participants may avoid manually providing typed input into the meeting document to summarize portions of the meeting during and / or after the meeting. This may conserve resources on each computing device that would normally be utilized to process such input. In some embodiments, method 408 may optionally include act 410 of determining whether the text entry corresponds to a meeting action item. For example, a meeting action item may be a task created by one or more participants during a meeting that requires one or more people to take action (e.g., gather specific information before a follow-up meeting). This determination may be based on manual input from participants and / or others to explicitly designate the text entry as an action item. Alternatively, or additionally, the determination may be based on terms contained in the text entry, the tone of the detected text entry (e.g., an inquisitive tone), the context in which the text entry was entered into the meeting document (e.g., the moment in the meeting when a particular participant describes what needs to be done before the next meeting). Once it is determined that the text entry corresponds to an action item, method 400 may proceed from operation 410, optionally via continuation element "C," to operation 502 of method 500, as shown in Figures 4 and 5.
[0057] If it is determined that the text entry does not correspond to an action item and / or operation 410 is optionally bypassed, method 400 may proceed to operation 412. Operation 412 may include determining whether other meeting content indicates a change in relevance to the text entry. For example, if additional natural language content from another participant and / or other contextual data indicates that the text entry is less relevant, the text entry may be deemed less relevant. For example, emails received by one or more participants during and / or after the meeting may be processed, with prior permission from one or more participants, to determine whether a particular text entry is more or less relevant. If other meeting content indicates a change in the relevance of the text entry, method 400 may return to operation 406 to determine the degree of relevance of the text entry and / or the natural language content that formed the basis of the text entry. Otherwise, if other meeting content does not indicate a change in the relevance of the text entry, method 400 proceeds from operation 412 to operation 402, and may optionally proceed to operation 302 via continuation element "A" if no additional natural language content has been provided by any of the participants (e.g., when the meeting has ended).
[0058] FIG. 5 illustrates a method 500 for generating action items based on natural language content provided during a participant's conference, reminding a particular participant of the action item, and / or designating the action item as completed based on a particular condition. Method 500 may be performed by one or more applications, devices, and / or any other devices or modules capable of interacting with conference participants. Method 500 may include an operation 502 of generating data characterizing an action item for one or more conference participants and / or other persons. In some embodiments, the generated data may be based on natural language content and / or other data provided by one or more participants in the conference, one or more applications associated with the conference, and / or one or more other persons and / or devices associated with the conference. For example, a participant in a video conference may provide a spoken phrase such as "Yes, let's follow up next month" in response to another participant providing another spoken phrase such as "Can we talk about maintenance fees soon?" Audio corresponding to each spoken phrase may be processed to generate a text entry, which may be further processed to generate data underlying the action item. For example, one or more trained machine learning models may be utilized to process the text entries and generate summary entries from the text entries. The summary entries may be designated as "action items," which may then be incorporated into meeting documents being generated and / or generated by the conferencing application and / or other applications (e.g., assistant applications).
[0059] Method 500 may proceed from operation 502 to operation 504, which may include determining whether data associated with the meeting indicates that the action item should have certain conditions. The conditions may be utilized to render one or more reminders to one or more participants to fulfill the action item and / or to determine whether the action item has been fulfilled (i.e., completed). For example, a spoken statement such as "Let's follow up on that next month" may provide an indication that the action item should have one or more certain conditions. Alternatively, or additionally, a spoken statement during the meeting such as "Send me that attachment and I'll get to work on this" may provide an indication that receipt of the "attachment" should trigger a reminder for a participant to work on the action item. In other words, a conditional statement occurring within a context that also identifies a particular action item may indicate that the action item should be stored in association with a conditional reminder and / or fulfillment condition.
[0060] If the data indicates that the action item should have certain conditions, method 500 may proceed from operation 504 to operation 506. Otherwise, if the data does not indicate that the action item should have certain conditions, method 500 may proceed from operation 504 to operation 510, which incorporates the action item into the meeting document. Operation 506 may include processing data related to the meeting to facilitate identifying the action item conditions. For example, one or more spoken words from one or more participants may provide the basis for establishing the conditions for the particular action item. Alternatively, or additionally, contextual data related to the meeting may provide the basis for establishing the conditions for the particular action item. For example, calendar data correlating a series of meetings and / or reminders for the series of meetings may serve as the basis for the “due date” of the action item and / or the time to remind participants about the action item (e.g., 24 hours before the next meeting in the related series). Alternatively, or additionally, the gathering of participants and / or communication between participants after the meeting may trigger a reminder for the action item generated based on the meeting. For example, a first participant sending an email to a second participant may trigger a reminder to the first participant to complete an action item that may have been generated during a previous meeting in which the second participant was present.
[0061] From operation 506, method 500 may proceed to operation 508, where action item data that conditionally characterizes the action item is generated. The action item data may then be stored in association with one or more participants who may be tasked with completing the action item and / or who may otherwise be associated with the action item. For example, a conferencing application may communicate the action item data to another application (e.g., an automated assistant application if the conferencing application is separate from the automated assistant), which may utilize the action item data to generate reminders for participants (with prior permission from the participants) and / or to determine whether the action item has been completed. From operation 508, method 500 may proceed to operation 510, where the action item is incorporated into a conference document. In this manner, the conference document may provide a summary of relevant topics discussed during the conference and / or a comprehensive list of action items created during the conference. Each action item may optionally serve as an embedded link to other data that may be useful in completing the respective action item.
[0062] From operation 510, method 500 may optionally proceed to optional operation 512, which determines whether one or more conditions and / or action items have been met. If it is determined that one or more conditions have been met, method 500 may proceed to operation 514 to indicate that the action item has been fulfilled and / or to render a reminder of the action item to one or more associated participants and / or others. For example, in response to two participants subsequently meeting in person and / or via conference call may satisfy a condition, the conferencing application may cause a reminder of the action item to be rendered on a device associated with each of the two participants. Method 500 may then optionally proceed from operation 514 via continuation element "A" to operation 302 of method 300, as shown in FIG. 3 .
[0063] 6 is a block diagram 600 of an exemplary computer system 610. The computer system 610 typically includes at least one processor 614 that communicates with a number of peripheral devices via a bus subsystem 612. These peripheral devices may include, for example, a storage subsystem 624 including memory 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices enable user interaction with the computer system 610. The network interface subsystem 616 provides an interface to an external network and is coupled to corresponding interface devices in other computer systems.
[0064] The user interface input devices 622 may include pointing devices such as keyboards, mice, trackballs, touchpads, graphics tablets, scanners, touchscreens integrated into displays, audio input devices such as voice recognition systems, microphones, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computer system 610 or communications network.
[0065] The user interface output devices 620 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or other mechanism for producing a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to encompass all possible types of devices and methods for outputting information from the computer system 610 to a user or to another machine or computer system.
[0066] Storage subsystem 624 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 624 may include logic for performing selected aspects of method 300, method 400, method 500, and / or for implementing one or more of system 200, computing device 104, computing device 112, computing device 152, server device 142, and / or any other applications, devices, apparatuses, and / or modules discussed herein.
[0067] These software modules typically execute on the processor 614 alone or in combination with other processors. The memory 625 used by the storage subsystem 624 may include a main random access memory (RAM) 630 for storing instructions and data during program execution and multiple memories, such as a read-only memory (ROM) 632 in which fixed instructions are stored. The file storage subsystem 626 may provide persistent storage of program files and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of an embodiment may be stored by the file storage subsystem 626 in the storage subsystem 624 or on other machines accessible by the processor(s) 614.
[0068] Bus subsystem 612 provides a mechanism that allows the various components and subsystems of computer system 610 to communicate with each other as intended. Although bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0069] Computer system 610 can be of various types, such as a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computer system 610 shown in Figure 6 is intended only as a specific example to illustrate some implementations. Many other configurations of computer system 610 can have more or fewer components than the computer system shown in Figure 6.
[0070] In situations where the systems described herein collect or may utilize personal information about users (or "participants," as often referred to herein), users may be provided with the opportunity to control whether a program or feature collects user information (e.g., information about the user's social networks, social actions or activities, occupation, user preferences, or the user's current geographic location) or whether and / or how it receives content from content servers that may be more relevant to the user. Also, certain data may be processed in one or more ways to remove personally identifiable information before it is stored or used. For example, a user's identity may be processed so that personally identifiable information about the user cannot be determined, or if geographic location information is obtained (e.g., to the city, zip code, or state level), the user's geographic location may be generalized so that the user's specific geographic location cannot be determined. Thus, users may control how information about them is collected and / or used.
[0071] While several embodiments have been described and illustrated herein, various other means and / or structures can be utilized to perform the functions and / or obtain the results and / or one or more advantages described herein, and each such variation and / or modification is deemed to be within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be illustrative, and the actual parameters, dimensions, materials, and / or configurations will depend on the specific application(s) for which the teachings are used. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. Accordingly, it should be understood that the foregoing embodiments have been presented by way of example only, and that, within the scope of the appended claims and their equivalents, embodiments may be practiced other than as specifically described and claimed. Embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is within the scope of the present disclosure, provided that such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
[0072] In some embodiments, a method implemented by one or more processors is described as including an operation such as determining, by an application, that a conference of a plurality of different participants is occurring or scheduled to occur. The conference provides an opportunity for one or more participants of the plurality of different participants to communicate information to other participants of the plurality of different participants. The method may further include determining, by the application, that one or more instances of data are associated with the conference based at least on one or more instances of data including content determined to be associated with at least one participant of the plurality of different participants. The method may further include biasing automatic speech recognition performed on audio data during the conference of the plurality of different participants according to content of the one or more instances of data. The audio data embodies speech from one or more participants communicating information to other participants. The method may further include generating, by the application, an entry for a conference document based on speech recognition results from the automatic speech recognition biased according to content of the one or more instances of data. The entry characterizes at least a portion of the information communicated from one or more participants to other participants.
[0073] These and other implementations of the techniques disclosed herein may optionally include one or more of the following features.
[0074] In some embodiments, determining that one or more instances of data are associated with a conference includes determining that the one or more instances of data include documents accessed and / or edited by at least one participant of the plurality of different participants prior to the conference. In some embodiments, determining that one or more instances of data are associated with a conference additionally or alternatively includes determining that the one or more instances of data include documents accessed and / or edited by at least one participant of the plurality of different participants within a threshold duration prior to the conference. In some embodiments, determining that one or more instances of data are associated with a conference additionally or alternatively includes determining that the one or more instances of data include documents accessed and / or edited by at least one participant of the plurality of different participants during the conference. In some embodiments, determining that one or more instances of data are associated with a conference additionally or alternatively includes determining that the one or more instances of data include documents that embody one or more terms identified in a conference invitation for the conference and that are accessible to at least one participant of the plurality of different participants.
[0075] In some embodiments, determining that one or more instances of data are related to the conference includes determining that the one or more instances of data include a document that embodies one or more terms identified in a title of a conference invitation for the conference.
[0076] In some embodiments, biasing the automatic speech recognition according to the content of the one or more instances of data includes generating one or more candidate terms for inclusion in the entry for the conference document based on a portion of the audio data, and assigning a weight value to each term of the one or more candidate terms, each weight value being based at least in part on whether a particular term of the one or more candidate terms is included in the content of the one or more instances of data.
[0077] In some embodiments, determining that one or more instances of data are related to the meeting includes determining that one or more documents containing the content were accessed by at least one participant of the plurality of different participants within a threshold duration prior to the meeting. In some of these embodiments, the threshold duration is based on when a meeting invitation was received and / or when the document was accessed by at least one participant who accessed the document. In some embodiments, determining that one or more instances of data are related to the meeting additionally or alternatively includes selecting one or more terms from the one or more documents as content that provides a basis for biasing the automatic speech recognition. The one or more terms may be selected based on an inverse document frequency of the one or more terms appearing in the one or more documents.
[0078] In some embodiments, a method implemented by one or more processors is described as including operations such as causing an application on a computing device to process audio data corresponding to spoken natural language content to facilitate generating a text entry for a conference document. The spoken natural language content is provided by a conference participant to one or more other participants in the conference. The method may further include determining, based on the text entry, a degree of relevance of the text entry to one or more instances of data related to the conference. The one or more instances of data include documents accessed by at least one participant of the conference prior to and / or during the conference. The method may further include determining whether to incorporate the text entry into the conference document based on the degree of relevance. The method may further include, when the application determines to incorporate the text entry into the conference document, causing the application to incorporate the text entry into the conference document, wherein the conference document is rendered on a display interface of a computing device or an additional computing device accessed by one or more other participants of the conference during the conference.
[0079] These and other implementations of the techniques disclosed herein may optionally include one or more of the following features.
[0080] In some embodiments, the method may further include, when the application determines to incorporate the text entry into the conference document, determining that during the conference, a particular participant of the conference has selected, via a computing device or other computing device interface, to generate an action item based on the text entry in the conference document. The action item is generated to provide a conditional reminder to at least one participant of the conference. In some embodiments, the conditional reminder is rendered for at least one participant of the conference if one or more conditions are met. The one or more conditions may be determined to be met using contextual data accessible to at least the application. For example, the contextual data may include a location of at least one participant of the conference, and the one or more conditions may be met when at least one participant of the conference is within a threshold distance of a particular location.
[0081] In some embodiments, determining a degree of relevance of the text entry to one or more instances of data related to the conference includes determining that a first participant in the conference provided text input to a first document during the conference and that a second participant in the conference provided additional text input to a second document during the conference. In some of these embodiments, the degree of relevance is based on whether the text input and the additional text input correlate with a text entry generated from spoken natural language input. In some embodiments, determining a degree of relevance of the text entry to one or more instances of data related to the conference additionally or alternatively includes determining that a first participant in the conference provided spoken input during the conference and that a second participant in the conference provided additional spoken input within a threshold duration after the participant provided spoken natural language content. In some of these embodiments, the degree of relevance is based on whether the spoken input and the additional spoken input correlate with the text entry. In some embodiments, determining a degree of relevance of the text entry to one or more instances of data related to the conference additionally or alternatively includes determining that at least one participant in the conference made a non-verbal gesture when the participant in the conference provided spoken natural language content, and in some of these embodiments, the degree of relevance is based on an interpretation of the non-verbal gesture by the application or another application.
[0082] In some embodiments, a method implemented by one or more processors is described as including an operation such as determining, by an application, that a conference participant has provided natural language content to an accessible computing device during the conference. The application is accessible via the computing device, and the conference includes one or more other participants. The method may further include causing input data to be processed to facilitate generating a text entry in the conference document in response to determining that the participant has provided the natural language content. The input data is captured by an interface of the computing device and characterizes the natural language content provided by the participant. The method may further include determining, based on processing the input data, whether to incorporate the text entry into the conference document as an action item to be completed by at least one participant of the conference. Determining whether to incorporate the text entry as an action item is based at least in part on whether the natural language content embodies a request for at least one participant and / or application. The method may further include, if the application determines to incorporate the text entry into the conference document as an action item, causing the application to incorporate the action item into the conference document. The conference document is accessible during the conference via a display interface on the computing device being accessed by one or more other participants in the conference or on another computing device.
[0083] These and other implementations of the techniques disclosed herein may optionally include one or more of the following features.
[0084] In some embodiments, the method may further include, if the application determines not to incorporate the text entry into a meeting document action item, causing the text entry to be incorporated into the meeting document as a transcription of natural language content provided by a participant of the meeting. In some embodiments, the method may further include, if the application determines to incorporate the text entry into the meeting document as an action item, causing the application to render a conditional reminder for at least one participant when one or more conditions are met, the one or more conditions being determined to be met using at least contextual data accessible to the application.
Claims
1. 1. A method implemented by one or more processors, comprising: Determining, at a computing device and by an application, that a conference of a plurality of different participants is occurring or scheduled to occur; determining that the conference provides an opportunity for one or more participants of the plurality of different participants to communicate information to other participants of the plurality of different participants; determining, by the application, that one or more instances of data are associated with the conference based at least on the one or more instances of data including content determined to be associated with at least one participant of the plurality of different participants; biasing automatic speech recognition performed on audio data during the conference of the plurality of different participants according to the content of one or more instances of the data; biasing the audio data to embody voices from the one or more participants communicating the information to the other participants; generating, by the application, a meeting document entry based on speech recognition results from the automatic speech recognition biased according to the content of the one or more instances of the data; generating the entries, wherein the entries characterize at least a portion of the information to be communicated from the one or more participants to the other participants; A method comprising:
2. Determining that one or more instances of the data are associated with the conference includes:
2. The method of claim 1, comprising determining that one or more instances of the data include documents that were accessed and / or edited by at least one participant of the plurality of different participants prior to the conference.
3. Determining that one or more instances of the data are associated with the conference includes:
2. The method of claim 1, comprising determining that one or more instances of the data include a document that was accessed and / or edited by at least one participant of the plurality of different participants within a threshold duration prior to the conference.
4. Determining that one or more instances of the data are associated with the conference includes:
10. A method according to any preceding claim, comprising determining that one or more instances of said data include a document being accessed and / or edited by at least one participant of said plurality of different participants during said conference.
5. Determining that one or more instances of the data are associated with the conference includes:
10. The method of any preceding claim, comprising determining that the one or more instances of the data include documents that embody one or more terms identified in a meeting invitation for the meeting and that are accessible to at least one participant of the plurality of different participants.
6. Determining that one or more instances of the data are associated with the conference includes:
10. A method according to any preceding claim, comprising determining that one or more instances of said data include documents that embody one or more terms identified in a title of a conference invitation for said conference.
7. Biasing the automatic speech recognition according to the content of the one or more instances of the data comprises: generating one or more candidate terms for inclusion in the entry of the meeting document based on a portion of the audio data; assigning a weight value to each term of the one or more candidate terms; assigning each weight value based at least in part on whether a particular term of the one or more candidate terms is included in the content of one or more instances of the data; 10. A method according to any preceding claim, comprising:
8. Determining that one or more instances of the data are associated with the conference includes:
2. The method of claim 1, comprising determining that one or more documents containing the content were accessed by at least one participant of the plurality of different participants within a threshold duration prior to the conference.
9. The method of claim 8 , wherein the threshold duration is based on when a meeting invitation is received and / or when the document is accessed by at least one participant who accessed the document.
10. Determining that one or more instances of the data are associated with the conference includes: selecting one or more terms from the one or more documents as content that provides a basis for biasing the automatic speech recognition; selecting, wherein the one or more terms are selected based on an inverse document frequency of the one or more terms appearing in the one or more documents; The method of claim 8, comprising:
11. 1. A method implemented by one or more processors, comprising: causing an application on a computing device to process audio data corresponding to spoken natural language content to facilitate generating a text entry of a meeting document; the spoken natural language content being provided by a conference participant to one or more other conference participants for processing; determining, based on the text entry, a degree of relevance of the text entry to one or more instances of data related to the conference; determining that the one or more instances of data include documents accessed by at least one participant of the conference prior to and / or during the conference; determining whether to incorporate the text entry into the meeting document based on the degree of relevance; If the application determines to incorporate the text entry into the meeting document, causing the application to incorporate the text entry into the meeting document; the conference document is rendered on a display interface of the computing device or an additional computing device accessed by the one or more other participants of the conference during the conference; A method comprising:
12. If the application determines to incorporate the text entry into the meeting document, determining, during the conference, that a particular participant of the conference has selected, via an interface of the computing device or the other computing device, to generate an action item based on the text entry in the conference document; determining that the action item is generated to provide a conditional reminder to at least one participant of the meeting; The method of claim 11 further comprising:
13. the conditional reminder is rendered to the at least one participant of the conference if one or more conditions are met; The method of claim 12 , wherein the one or more conditions are determined to be satisfied using at least contextual data accessible to the application.
14. the contextual data includes a location of the at least one participant in the conference; The method of claim 13 , wherein the one or more conditions are met when the at least one participant of the conference is within a threshold distance of a particular location.
15. Determining the degree of relevance of the text entry to one or more instances of the data related to the conference includes: determining that a first participant in the conference has provided text input to a first document during the conference and that a second participant in the conference has provided additional text input to a second document during the conference; determining the degree of relevance based on whether the text input and the additional text input correlate with the text entry generated from the spoken natural language input; The method of any one of claims 11 to 14, further comprising:
16. Determining the degree of relevance of the text entry to one or more instances of the data related to the conference includes: determining that a first participant in the conference has provided spoken input during the conference and that a second participant in the conference has provided additional spoken input within a threshold duration after the participant provided the spoken natural language content; determining whether the degree of relevance is based on whether the spoken input and the additional spoken input are correlated to the text entry; The method according to any one of claims 11 to 14, comprising:
17. Determining the degree of relevance of the text entry to one or more instances of the data related to the conference includes: determining that the at least one participant in the conference made a non-verbal gesture when the participant in the conference provided the spoken natural language content; determining the degree of relevance based on an interpretation of the non-verbal gesture by the application or another application; The method according to any one of claims 11 to 16, comprising:
18. 1. A method implemented by one or more processors, comprising: determining, by an application, that a participant of a conference has provided natural language content to an accessible computing device during said conference; determining that the application is accessible via the computing device and that the conference includes one or more other participants; In response to determining that the participant has provided the natural language content, processing input data to facilitate generating a text entry for a conference document; the input data is captured by an interface of the computing device and processed to characterize the natural language content provided by the participant; determining whether to incorporate the text entry into the meeting document as an action item to be completed by at least one participant of the meeting based on processing the input data; determining whether to incorporate the text entry as the action item is based at least in part on whether the natural language content embodies a request for the at least one participant and / or the application; If the application determines to incorporate the text entry as an action item into the meeting document, causing the application to incorporate the action item into the meeting document; the conference document is accessible during the conference via a display interface of the computing device or another computing device accessed by the one or more other participants of the conference; A method comprising:
19. If the application determines not to incorporate the text entry into the meeting document or the action item, incorporating the text entry into the meeting document as a transcription of the natural language content provided by the participants of the meeting; 20. The method of claim 18, further comprising:
20. If the application determines to incorporate the text entry as an action item into the meeting document, causing the application to render a conditional reminder to the at least one participant when one or more conditions are met; the causing of rendering, wherein the one or more conditions are determined to be satisfied using at least contextual data accessible to the application; 20. The method of claim 18, further comprising:
21. one or more processors; a memory storing instructions that, when executed, cause the one or more processors to perform the operations of any one of claims 1 to 20; A system comprising:
22. A non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform the operations of any one of claims 1-20.