Real-time speech-to-text interaction enhancements and integration

By integrating the transcription pane into productivity applications through a real-time speech-to-text transcription system, the problem of users having difficulty taking notes while listening to lectures is solved. It provides real-time transcription, translation, definition and search functions, improving users' recording efficiency and experience.

CN114787916BActive Publication Date: 2025-09-16MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080084891.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-09
Filing Date
2020-11-05
Publication Date
2025-09-16
Estimated Expiration
2040-11-05

AI Technical Summary

Technical Problem

Users can struggle to take notes while listening to a lecture, especially when faced with hearing or speech impairments and when subtitles make it difficult to interact with other tasks.

Method used

Through a real-time speech-to-text transcription system, a joining code is generated, lecture content is transcribed into text in real time or near real time, and a transcription pane is presented in productivity applications, supporting highlighting, annotation, translation, definition, search, and pause functions to enhance the user interaction experience.

Benefits of technology

It enables real-time note-taking while listening to lectures, improves user comprehension and note-taking efficiency, reduces the need for manual searching, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114787916B_ABST
    Figure CN114787916B_ABST
Patent Text Reader

Abstract

In a non-limiting example of the present disclosure, systems, methods, and devices for integrating speech-to-text transcription in a productivity application are presented. A request is sent by a first device to access real-time speech-to-text transcription of an audio signal being received by a second device. The real-time speech-to-text transcription may be presented in a transcription pane of the productivity application on the first device. A request to translate the transcription into a different language may be received. The transcription may be translated in real time and presented in the transcription pane. A selection of a word in the presented transcription may be received. A request to drag a word from the transcription pane and drop the word into a window outside the transcription pane in the productivity application may be received. The word may be presented in a window outside the transcription pane in the productivity application.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] It is often difficult for users to listen to a presentation (e.g., a lecture) and simultaneously take notes related to the presentation. This can be due to a variety of reasons. For example, the user may be unfamiliar with the presentation topic, have auditory learning problems, have hearing problems, and / or language problems (e.g., the presentation is not in the user's first language). Subtitles are an excellent mechanism for improving a user's ability to understand the content. However, even when subtitles are available during a live presentation, they can be difficult to follow or interact with while performing one or more additional tasks (e.g., taking notes).

[0002] It is against this general technical environment that the various aspects of the technology disclosed herein are contemplated.Furthermore, while a general environment is discussed, it should be understood that the examples described herein should not be limited to the general environment identified in the background. Summary of the Invention

[0003] This summary is provided to introduce a set of concepts in a simplified form that will be further described in the detailed description below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. Additional aspects, features, and / or advantages of examples will be set forth in part in the description that follows and in part will be apparent from the description of the disclosure or may be learned through practice of the disclosure.

[0004] Non-limiting examples of the present disclosure describe systems, methods, and devices for integrating speech-to-text transcription in productivity applications. A join code generation request may be received from a computing device associated with a speaking user. The request may be received by a real-time speech-to-text service. The real-time speech-to-text service may generate a join code and send it to the computing device associated with the speaking user. An audio signal including speech may be received by the computing device associated with the speaking user. The audio signal may be sent to the real-time speech-to-text service, where it may be transcribed.

[0005] A computing device associated with a joining user can request access to the transcription when it is generated (e.g., a transcription instance). The request may include a joining code generated by the real-time speech-to-text service. After being authenticated, the transcription can be presented in real time or near real time in a transcription pane in a productivity application associated with the joining user. Various actions can be performed in association with the transcription, the productivity application, other applications, and / or a combination thereof. In some examples, content in the transcription pane can be highlighted and / or annotated. Content from the transcription pane can be moved (e.g., via dragging or dropping) from the transcription pane to another window of the productivity application (e.g., a notebook window, a note-taking window). Definitions for words and phrases can be presented in the transcription pane. Web searches associated with words and phrases in the transcription pane can be automatically performed. In some examples, a pause function of the transcription pane can be utilized to pause incoming subtitles for the transcription instance. The subtitles maintained during the pause can then be presented after resuming the transcription instance. In an additional example, the transcription pane may include a selectable option for translating the transcription from a first language into one or more additional languages. The real-time speech-to-text service and / or translation service may process such a request, translate the transcription and / or audio signal as it is being received, and send the translation to the joining user's computing device, where it may be presented in a transcription pane. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Non-limiting and non-exhaustive examples are described with reference to the following figures:

[0007] Figure 1 is a schematic diagram illustrating an example distributed computing environment for integrating speech-to-text transcription in productivity applications.

[0008] Figure 2 Shown are exemplary elements of three cloud-based services that can be utilized in integrating speech-to-text transcription in productivity applications.

[0009] Figure 3 Shown is interaction with text in a transcription pane integrated into a productivity application.

[0010] Figure 4 A detached window is shown transferring text in the transcription pane to a productivity application.

[0011] Figure 5 Annotation of text in a transcription pane of a productivity application is shown.

[0012] Figure 6 Shown is a selectable element for changing the language in which real-time speech-to-text transcription is presented in the transcription pane.

[0013] Figure 7AShown are selectable elements for presenting definitions of words and / or phrases included in a transcription pane of a productivity application.

[0014] Figure 7B Shown are selectable elements for causing a web search to be performed in association with words and / or phrases included in a transcription pane of a productivity application.

[0015] Figure 8 Shown are selectable elements and related actions associated with pausing and resuming real-time speech-to-text captioning in a transcription pane of a productivity application.

[0016] Figure 9A is an exemplary method for integrating speech-to-text transcription in productivity applications.

[0017] Figure 9B is an exemplary method for presenting definitions of words and / or phrases from a custom dictionary that are included in a transcription pane of a productivity application.

[0018] Figure 9C An exemplary method for pausing and resuming real-time speech-to-text captioning in a transcription pane of a productivity application.

[0019] Figure 10 and Figure 11 is a simplified diagram of a mobile computing device that can be used to implement aspects of the present disclosure.

[0020] Figure 12 is a block diagram illustrating example physical components of a computing device that can be used to implement aspects of the present disclosure.

[0021] Figure 13 is a simplified block diagram of a distributed computing system in which aspects of the present disclosure may be implemented. DETAILED DESCRIPTION

[0022] Various embodiments will be described in detail with reference to the accompanying drawings, wherein like reference numerals represent like parts and assemblies throughout the several views. Reference to various embodiments does not limit the scope of the claims appended hereto. In addition, any examples set forth in this specification are not intended to be limiting and merely set forth some of the many possible embodiments of the appended claims.

[0023] Non-limiting examples of the present disclosure describe systems, methods, and devices for integrating speech-to-text transcription in productivity applications. According to the example, a first user (i.e., a speaking user) who wants to start a transcription instance that can be accessed by one or more other users can initiate the instance on a computing device. The request to initiate the transcription instance can be received by a real-time text-to-speech service, which can generate a join code that can be used by one or more other users and associated computing devices to join the transcription instance.

[0024] A join code can be sent to users and / or user accounts that the speaking user wants to give access to the transcription instance. In some examples, the join code can be sent electronically back to a computing device associated with the first user. The first user can then provide the join code to other users via various means (e.g., writing it on a whiteboard, sending it via email, posting it on a sharing website, etc.). In other examples, the join code can be automatically sent electronically to one or more user accounts (e.g., a user account associated with a class list service provided via an SMS message, etc.). In additional examples, a device or the first user account can be authorized to receive transcriptions associated with different devices or user accounts (e.g., via a cached token), and a selectable option can be presented on the device associated with the first user account to join the transcription instance that was authorized when the transcription instance was initiated. Thus, the joining user does not need to manually enter a new join code each time they want to join a new transcription instance.

[0025] When the joining user enters the join code on the joining user's computing device, a productivity application with a transcription pane including a real-time transcription of the speaking user's transcription instance may be presented. That is, the join code may be sent to the real-time speech-to-text transcription service whose code is authenticated, and the real-time speech-to-text transcription service may then begin sending the transcription information from the transcription instance to the joining user's computing device. The join code may be entered into the productivity application or in a separate interface on the computing device. The productivity application may include one or more of the following: for example, a note-taking application, a notebook application, a word processing application, a presentation application, a task completion application, a spreadsheet application, and / or a messaging application.

[0026] The transcription pane may include multiple selectable elements for performing multiple actions. A first element may be selected to highlight content (e.g., subtitles, annotations) in the transcription pane and / or move the content from the transcription pane to a second window of a productivity application that includes the transcription pane. The second window may include, for example, a note-taking window, a diary window, or a presentation window. A second element may be selected to change the language in which the transcription is presented. A third element may be selected to add annotations to the transcription and / or to specific content in the transcription. A fourth element may be selected to present definitions associated with words or phrases in the transcription. A fifth element may be selected to perform a web search related to the words or phrases in the transcription. A sixth element may be selected to add a link or pin that will be associated with one or more words in the transcription. A seventh element may be selected to pause and resume the presentation of the content in the transcription. That is, the seventh element may pause the presentation of the subtitles of the current transcription instance, and when resumed, the backlog of subtitles may be presented in the transcription pane.

[0027] According to an example, transcriptions presented in a transcription pane of a productivity application may be automatically saved to the transcription section of the productivity application. Thus, in an example where the productivity application is a notebook application or a note-taking application with multiple sections, each new transcription may be saved by default to the transcription section of the notebook application or the note-taking application. In this way, all transcriptions associated with a user account may be accessed in a single location. In some examples, each transcription may be saved to a section of the corresponding productivity application along with the date and / or time at which the transcription was generated and / or completed. In an additional example, one or more natural language processing models may be applied to the transcriptions. Those one or more natural language processing models may be trained to identify one or more subject types associated with the transcriptions. Thus, the transcriptions may be saved to a location in the productivity application corresponding to one or more identified subject types (e.g., in a "biology" transcription section of a notebook application, in a "chemistry" transcription section of a notebook application, in a class type and / or number of a notebook application). In an additional example, the user may customize the location where the transcriptions are saved.

[0028] The systems, methods, and devices described herein provide technical advantages for integrating real-time speech-to-text transcription in productivity applications. Providing a mechanism for automatically presenting a real-time transcription of a speaking user associated with a productivity application and utilizing a note-taking feature to enhance this presentation also provides an enhanced user experience. For example, a user can take notes related to a speech (e.g., a lecture) in the first window of a productivity application while a real-time transcription of the speech is presented next to the window. The user can then highlight the transcribed text, drag and drop content from the transcription in the user's notes, link notes to the transcription, annotate the transcription, surface standard and custom definitions for words in the transcription, and pause and resume the transcription at will. Automatic web searches related to words and phrases in the transcription and the ability to link the most relevant content from those web searches to words and phrases in the transcription also enhance the user experience and reduce manual searching.

[0029] Figure 1 1 is a schematic diagram illustrating an example distributed computing environment 100 for integrating speech-to-text transcription in productivity applications. Computing environment 100 includes a transcription subenvironment 102, a network and processing subenvironment 114, and a computing device 104B. Network and processing subenvironment 114 may include and / or communicate with productivity application services 120, SST services 122, and / or translation services 124. Any and all devices described herein may communicate with each other via a network, such as network 116 in network and processing subenvironment 114.

[0030] The transcription sub-environment 102 includes a speaking environment 106 and a computing device 104A. In the speaking environment 106, a computing device 110 communicates with a real-time speech-to-text service (e.g., an STT service 122) in the cloud. In this example, the computing device 110 is a smartphone. However, the computing device 110 can be any computing device that includes a microphone or can receive a signal from a microphone (e.g., a laptop, desktop, tablet, smartwatch). The computing device 110 can communicate with the STT service 122 via a specific STT application, via an application that includes a plug-in associated with the STT application, via a web browser, or other communication means (e.g., a speech translation service application, via an auxiliary device and / or application, etc.). The computing device 110 can also communicate with the STT service 122 using an API.

[0031] In this example, computing device 110 receives a request to generate a join code for voice transcription. For example, user 108 may utilize an application executing on computing device 110 to input a generate code request, and the generate code request may be processed by one or both of computing device 110 and / or STT service 122. Processing of the request may include generating a join code that may be used by other devices and / or applications to join a speech-to-text instance of active, real-time speech from computing device 110 (e.g., audio received by computing device 110 and transcription of the audio performed in the cloud). The join code may include one or more characters, a QR code, a barcode, or a different code type that provides access to an active instance of the speech-to-text instance. In this example, the generated join code is join code 112 [JC123].

[0032] The speaking user 108 speaks, and the audio signal is received by the computing device 110. The computing device 110 sends the audio signal to the STT service 122. The STT service 122 analyzes the audio signal and generates a text transcription based on the analysis. Figure 2 The analysis that may be performed when generating a text transcription is described in more detail. The transcription may be performed in the language in which the audio was originally received (e.g., if the speaking user 108 speaks English, the audio may initially be transcribed in English by the STT service 122). In an example, the translation service 124 may translate the transcription into one or more other languages ​​than the language in which the audio was originally received. In some examples, the transcription of the audio from the original language may be translated by the translation service 124. In other examples, the original audio may be directly transcribed into one or more additional languages. Figure 2 Additional details are provided regarding the processing performed by the translation service 124 .

[0033] The information included in productivity application service 120 can be utilized to process audio received from computing device 110, enhance the transcription or translation of the audio, and / or enhance or otherwise supplement the transcription of the audio. As an example, productivity application service 120 can include materials associated with a lecture being given by speaking user 108 (e.g., lecture notes, presentation documents, quizzes, tests, etc.), and this information can be utilized to generate a custom dictionary and / or corpus that is used to generate a transcription of the audio received by computing device 110. In another example, productivity application service 120 can include transcription settings and / or translation settings associated with a user account for computing device 104B and can provide subtitles and / or translations to computing device 126 based on these settings.

[0034] In this example, a productivity application is displayed on computing device 104A. Specifically, the note-taking productivity application is displayed, and within the application, a captioning window is presented for joining an ongoing lecture associated with a speech / lecture by speaking user 108 and a join code 112. Join code 112 is entered into the "Join Conversation" area of ​​the captioning window, and the user associated with computing device 104A has selected English as her preferred language for receiving a transcription of the transcription instance. Join code 112 is sent from computing device 104A to a real-time speech-to-text service, which authenticates the code and authorizes the speech-to-text of the transcription instance from speaking user 108 to be provided to computing device 104A. In this example, the speech-to-text is sent to computing device 104B, which is the same computing device as computing device 104A, as shown by captions 128 in transcription pane 129. Transcription pane 129 is included in the note-taking productivity application, next to note window 126 for "Lecture #1." For example, the speaking user 108 may be an organic chemistry professor delivering her first lecture in class, a transcript of which may be automatically generated via the STT service 122 and presented in a transcript pane 129 in a note-taking application where a student user is taking notes related to the first lecture. Additional details regarding various interactions that may be performed with respect to the captions 128 are provided below.

[0035] Figure 2 2 shows exemplary elements of three cloud-based services 200 that can be utilized to integrate speech-to-text transcription in productivity applications. The cloud-based services include productivity application service 221, speech-to-text (STT) service 222, and translation service 224. One or more of these services can be delivered via a web interface such as Figure 1 The networks in network 116 communicate with each other.

[0036] The productivity application service 221 includes a service repository 220 that may include stored data associated with one or more user accounts related to one or more productivity applications hosted by the productivity application service 221. These user accounts may additionally or alternatively be associated with an STT service 222 and / or a translation service 224. In the illustrated example, the service repository 220 includes document data 216, which may include one or more stored productivity documents and / or associated metadata; email data 212 and associated email metadata; calendar data 214 and associated calendar metadata; and user settings 218, which may include, for example, privacy settings, language settings, location preferences, and dictionary preferences. In some examples, the document data 216 may include lecture materials 232, which will be discussed below with respect to the STT service 222.

[0037] STT service 222 includes one or more speech-to-text language processing models. These language processing models are illustrated by neural network 228, supervised machine learning model 224, and language processing model 226. In some examples, when a computing device (e.g., Figure 1 When a computing device 110 in the example of FIG. 1 initiates a transcription instance, the audio signal received from the device can be sent to an STT service 222 where it is processed for transcription. The audio signal is represented by speech 230. Speech 230 can be provided to one or more speech-to-text language processing models. As shown, when processing speech 230, one or more speech-to-text language processing models can be trained and / or can utilize documents such as lecture materials 232. For example, if the user providing speech 230 is presenting a lecture using one or more corresponding documents related to the lecture (e.g., electronic slides from a presentation application on organic chemistry, lecture notes, etc.), the language processing model used to transcribe speech 230 can utilize the material (e.g., via analysis of those electronic documents) to develop a custom corpus and / or dictionary that can be utilized in the language processing model to determine the correct output for speech 230. This is illustrated by a domain-specific dictionary / corpus 234.

[0038] According to some examples, vocabulary (e.g., words, phrases) determined to be specific and / or unique to a particular language processing model, custom corpus, and / or custom dictionary can be automatically highlighted and / or otherwise distinguished from other subtitles in the transcription pane of the productivity application. For example, if there are terms used in a particular discipline provided as subtitles from the transcribed audio in the transcription pane (e.g., organic chemistry, evolutionary biology, mechanical engineering, etc.), these terms can be highlighted, underlined, bolded, or otherwise indicated as being associated with the particular discipline.

[0039] In some examples, the documents / materials used to generate, enhance and / or be used for audio / speech processing in the language processing model can be associated with multiple users. For example, electronic documents / materials from a first group of users (e.g., professors) in a first science department of a university can be utilized in a language processing model for speech received by users in that department, and electronic documents / materials from a second group of users (e.g., professors) in a second science department of the university can be utilized in a language processing model for speech received from users in that department. Other electronic documents / materials from other groups can be utilized to process speech from users with similar vocabularies. The language processing model for transcribing speech 230 can utilize a standard dictionary and / or one or more standard corpora to determine the correct output of speech 230. This is illustrated by standard dictionary / corpora 236.

[0040] The translation service 224 can receive output (e.g., a transcription of speech 230) from the STT service 222 and translate the output into one or more additional languages. The translation service 224 includes one or more language processing models that can be utilized in translating the output received from the STT service 222. These models are illustrated by the supervised machine learning model 204, the neural network 206, and the language processing model 208.

[0041] Figure 3 Shown is interaction with text in a transcription pane integrated into a productivity application. Figure 3 Computing device 302 is included that displays productivity application 304. Productivity application 304 includes application notes 310 and a transcription pane 306. Transcription pane 306 is integrated into productivity application 304 and includes subtitles 308. Subtitles 308 are rendered in real time or near real time relative to the receipt of speech (e.g., via an audio signal) and subsequent processing of the audio signal into text. The text is then displayed as subtitles 308 in transcription pane 306.

[0042] In this example, one or more words in the subtitles 308 are selected. The selection is shown as being performed by clicking and dragging the mouse from one side of the selected word to the other side of the selected word. However, it should be understood that other mechanisms for selecting subtitles in the transcription pane 306 (e.g., verbal commands, touch input, etc.) can be utilized. Interaction with selected subtitles will be further described below.

[0043] In some examples, a highlight element 307 in the transcription pane 306 can be selected. This selection can cause the highlight element to be presented, which can be utilized to highlight text, such as the selected text of interest shown here. In some examples, the user can select a color from a variety of colors in which the text can be highlighted. In other examples, the highlighted and / or selected text can be interactive as described more fully below.

[0044] Figure 4 A detached window is shown transferring text in the transcription pane to a productivity application. Figure 4 A computing device 402 is included that displays a productivity application 404. The productivity application 404 includes an application note 410 and a transcription pane 406. The transcription pane 406 is integrated into the productivity application 404 and includes subtitles 408. The subtitles 408 are rendered in real time or near real time relative to the receipt of speech (e.g., via an audio signal) and subsequent processing of the audio signal into text. The text is then displayed as subtitles 408 in the transcription pane 406.

[0045] In this example, one or more words in the subtitles 408 are selected. The one or more words are shown as selected text 414. An indication is made to interact with the selected text 414. Specifically, a click and drag is performed on the selected text 414, thereby receiving a click relative to the selected text 414 in the subtitles 408. Then, a drag and drop mechanism is performed relative to the subtitles 408 of the application note 410. Thus, the selected text 414 can be inserted into the application note 410 at the location where it was placed. In some examples, the selected text 414 can be copied and pasted into the application note 410 via the drag and drop mechanism. In other examples, the selected text 414 can be transferred via a cut and paste type mechanism. In some examples, the selected text 414 can be copied and stored in a temporary storage device on the computing device 402 while it is moved (e.g., via drag and drop) from the transcription pane 406 to the application note 410. In addition, in this example, when the selected text 414 is inserted into the application note 410, it is associated with the link 412. If selected, link 412 can cause the location of the selected text 414 in the subtitles 408 to be presented in the transcription pane 406. In some examples, link 412 can be an embedded link. As an example of how a link can be used, if the user does not currently have lecture notes corresponding to the subtitles 408 displayed in the transcription pane 406, and the user interacts with link 412, those lecture notes corresponding to the selected text 414 and / or a specific location within those lecture notes can be caused to be presented in the transcription pane 406.

[0046] Figure 5 Annotation of text in a transcription pane of a productivity application is shown. Figure 5 A computing device 502 is included that displays a productivity application 504. The productivity application 504 includes application notes 510 and a transcription pane 506. The transcription pane 506 is integrated into the productivity application 504 and includes subtitles 508. The subtitles 508 are rendered in real time or near real time relative to the receipt of speech (e.g., via an audio signal) and subsequent processing of the audio signal into text. The text is then displayed as subtitles 508 in the transcription pane 506.

[0047] In this example, a selection is made to one or more words in the subtitle 508. Those one or more words are shown as selected text 514. Subsequent selections associated with the annotation element 512 are then received in the transcription pane 506. In this example, the selection is made by clicking the mouse on the annotation element 512. However, other selection mechanisms (e.g., touch input, voice input, etc.) may also be considered. After the selection of the annotation element 512, an annotation window 516 is displayed in the transcription pane 506. The annotation window 516 provides a mechanism for the user to leave an annotation that will be associated with the selected text. In this example, the user adds the text "The professor said this concept will be on the exam" in the annotation window 516 with the selected text 514. In some examples, after the annotation is associated with the selected text, the annotation can be automatically presented (e.g., in the annotation window 516 or in a separate window or pane) when input is received next to the corresponding subtitle / selected text in the transcription pane. In an additional example, after associating an annotation with the selected text, if the selected text is then inserted into application note 510, the user can interact with the inserted text, which can cause the annotation to be automatically presented relative to the inserted text in application note 510.

[0048] Figure 6 Shown is a selectable element for changing the language in which real-time speech-to-text transcription is presented in the transcription pane. Figure 6 Computing device 602 is included that displays productivity application 604. Productivity application 604 includes application notes 610 and a transcription pane 606. Transcription pane 606 is integrated into productivity application 604 and includes subtitles 608. Subtitles 608 are rendered in real time or near real time relative to the receipt of speech (e.g., via an audio signal) and subsequent processing of the audio signal into text. The text is then displayed as subtitles 608 in transcription pane 606.

[0049] In this example, in the transcription pane 606, the translation language element 612 is selected. In this example, the selection is made via a mouse click on the translation language element 612. However, other selection mechanisms are also contemplated (e.g., touch input, voice input, etc.). After the translation language element 612 is selected, a plurality of selectable elements for modifying the presentation language of the subtitles 608 are displayed. In this example, the plurality of selectable elements are presented in a language flyout window 613, however, other user interface elements are contemplated (e.g., pop-up windows, drop-down lists, etc.). Any language included in the language flyout window 613 can be selected, which can cause the subtitles 608 to be presented in the transcription pane 606 in the selected language.

[0050] Figure 7AShown are selectable elements for presenting definitions of words and / or phrases included in a transcription pane of a productivity application. Figure 7A Computing device 702A is included that displays productivity application 704A. Productivity application 704A includes application notes 710A and a transcription pane 706A. Transcription pane 706A is integrated into productivity application 704A and includes subtitles 708A. Subtitles 708A are rendered in real time or near real time relative to the receipt of speech (e.g., via an audio signal) and subsequent processing of the audio signal into text. The text is then displayed as subtitles 708A in transcription pane 706A.

[0051] In this example, a word in subtitle 708A is selected. The word is the selected word 716A. The subsequent selection related to the dictionary search element 714A is then received in the transcription pane 706A. In this example, the selection is made via clicking the mouse on the dictionary search element 714A. However, other selection mechanisms (e.g., touch input, voice input, etc.) can be imagined. After the selection of the dictionary search element 714A, the definition window 712A is displayed in the transcription pane 706A. After the selection of the dictionary search element 714A, the definition of the selected word 716A can be automatically displayed in the definition window 712A. In some examples, the definition can be obtained from the standard dictionary of the computing device 702A local or the standard dictionary accessed via the web. In other examples, if it is determined that the selected word is in a custom dictionary associated with the language processing model for transcription, the definition can be obtained from the custom dictionary. For example, some words (especially those related to science) may not be included in the standard dictionary, so those words can be included in the custom dictionary generated for, for example, lectures, lecture collections and / or academic disciplines of a university. In an additional example, if the subtitle is determined to be related to a particular field (e.g., computer science, chemistry, biology), the definition presented in definition window 712A may be obtained from a technical dictionary in that field obtained via the web. In an additional example, a first definition of the selected word may be obtained from a standard dictionary, a second definition of the selected word may be obtained from a technical and / or custom dictionary, and both definitions may be presented in definition window 712A.

[0052] In some examples, a selection can be made to associate one or more definitions of the future customization window 712A with the selected word 716A. If such a selection is made, the one or more definitions can be displayed upon receiving an interaction with the word (e.g., if an interaction is received with the selected word 716A in the subtitles 708A, the definition can be presented in the transcription pane 706A; if the selected word 716A is inserted into the application note 710A and an interaction is received regarding the word in the application note 710A, the definition can be presented in the application note 710A).

[0053] Figure 7B Shown are selectable elements for causing a web search to be performed in association with words and / or phrases included in a transcription pane of a productivity application. Figure 7B Computing device 702B is included that displays productivity application 704B. Productivity application 704B includes application notes 710b and a transcription pane 706B. Transcription pane 706B is integrated into productivity application 704B and includes subtitles 708B. Subtitles 708B are rendered in real time or near real time relative to the receipt of speech (e.g., via an audio signal) and subsequent processing of the audio signal into text. The text is then displayed as subtitles 708B in transcription pane 706B.

[0054] In this example, a selection is made for a word in subtitle 708B. This word is selected word 716B. Then, as shown in FIG. Figure 7B As described, receive the subsequent selection relevant to the dictionary lookup element in transcription pane 706A.Therefore, the definition of selected word 716B is presented in definition window 712B. However, in this example, web search element 718B is further selected. By selecting web search element 718B, a web search relevant to the selected word 716B and / or its surrounding text in subtitles 708B can be performed, and information from one or more online sources identified as relevant to the search can be presented relative to the selected word 716B in subtitles 708B and / or relative to definition window 712B. In some examples, the content obtained from the web can be associated with one or more words in subtitles 708B. In such an example, when interacting with one or more words (e.g., via mouse hovering, via left mouse click, etc.), web content can be automatically presented.

[0055] Figure 8 Shown are selectable elements and related actions associated with pausing and resuming real-time speech-to-text captioning in a transcription pane of a productivity application. Figure 8Three transcription panes are included (transcription pane 802A, transcription pane 802B, transcription pane 802C), all of which are identical transcription panes at different stages of the pause / resume operation. The transcription panes are shown outside of the productivity application. However, it should be understood that the transcription panes shown herein can be integrated into the productivity application (e.g., in a pane adjacent to the note-taking window).

[0056] Subtitles (subtitles 806A, subtitles 806B, subtitles 806C) are rendered in real time or near real time relative to the receipt of speech (e.g., via an audio signal) and the subsequent processing of the audio signal into text. The text is then displayed in subtitles 806A. However, transcription pane 802A includes a plurality of selectable user interface elements in its upper portion, and pause / resume element 804A is selected. In this example, the selection is made via a mouse click on pause / resume element 804A. However, other selection mechanisms are contemplated (e.g., touch input, voice input, etc.).

[0057] Upon selection of pause / resume element 804A, subtitles may cease to be presented in real time in subtitles 806A. For example, even though the audio is still being concurrently received by the real-time speech-to-text service and the computing device displaying transcription pane 802A is still connected to the current transcription instance of the audio, after selection of pause / resume element 804A, subtitles transcribed from the audio may not be displayed in subtitles 806A. Instead, the subtitles may be stored in temporary storage (e.g., on a computing device associated with transcription pane 802A, in a buffer storage on a server computing device hosting the real-time speech-to-text service) until a subsequent "resume" selection is made for pause / resume element 804A.

[0058] In this example, when the pause / resume element 804A is selected, the subtitles 806A are paused at the current speaker speech position 808A. Thus, as shown in the transcription pane 802B, even when additional audio from the speaker is received by the real-time speech-to-text service (via the computing device receiving the audio) and transcribed, as indicated by the current speaker speech position 808B, the content is not presented in the position 810 where it would have been presented if the pause / resume element 804A selection had not been received. However, when a subsequent selection of the pause / resume element 804B is made in the transcription pane 802C, the subtitles held in the temporary storage state (e.g., the buffered state) can be automatically presented, as indicated by the subtitles being moved forward / presented to the current speaker speech position 808C in the subtitles 806C.

[0059] In addition, although no transcription pane is shown with a scroll bar, it should be understood that the subtitles can be scrolled while they are being presented or while they are in a paused state. For example, a user can pause the presentation of subtitles, scroll up to what the user missed during the ongoing lecture, resume the presentation of subtitles, and scroll to the current active state in the subtitles. Other mechanisms for moving forward or backward in subtitles are conceivable. For example, a user can use voice commands to locate subtitles (e.g., "go back five minutes," "jump back to [concept A] in the lecture"). In the case of voice commands, natural language processing can be performed on the received commands / audio, and one or more tasks identified via the processing can be performed, the results of which can be presented in the transcription pane.

[0060] Figure 9A An exemplary method 900A for integrating speech-to-text transcription in a productivity application is provided. The method 900A begins at a start operation and flow moves to operation 902A.

[0061] In operation 902A, a first device sends a request to access a real-time speech-to-text transcription of an audio signal currently being received by a second device. That is, the second device is associated with a speaking user. In some examples, the second device can be utilized to request generation of a join code for a transcription instance associated with the audio. For example, the request to generate a join code can be received from a speech-to-text application on the second device, a translation application on the second device, or a productivity application on the second device.

[0062] The request to generate a join code can be received by the real-time speech to text service, and a join code can be generated. The join code can include a QR code, a barcode, one or more characters, an encrypted signal, etc. In some examples, the request to access the real-time speech to text transcription can include receiving the join code from the first device. In other examples, when the request to access the real-time speech to text transcription is received, the first device can then present a field for entering the join code. In any case, once the join code is entered on the first device (e.g., in a productivity application, in a pop-up window), the first device can join the transcription instance associated with the speaking user.

[0063] From operation 902A, flow continues to operation 904A, where the real-time speech-to-text transcription is presented in a transcription pane of a productivity application user interface on the first device. The productivity application may include, for example, a note-taking application, a word processing application, a presentation application, a spreadsheet application, and / or a task completion application.

[0064] From operation 904A, flow continues to operation 906A, where a selection of a word in the presented transcription is received. For example, the selection may include highlighting the word, underlining it, copying the input, and / or electronically grabbing it. The selection may be input via mouse input, touch input, stylus input, and / or verbal input.

[0065] From operation 906A, the process continues to operation 908A, in which a request is received to drag a word from the transcription pane and drop the word into a window outside the transcription pane in the productivity application. In some examples, when the drag is initiated, the word can be copied to a temporary storage device and the word can be pasted from the temporary storage device into the productivity application at the location where the drop was initiated. In other examples, the word can be copied directly from the transcription pane and pasted into the productivity application at the location where the drop was initiated (e.g., without first copying to a temporary storage device). In some examples, the location in the productivity application where the word is dropped can include a note portion related to the subject of the transcription. In additional examples, one or more language processing models can be applied to the transcription, and the type of subject matter to which the transcription relates can be determined. In such an example, the productivity application can present one or more saved notes related to the subject matter of the transcription.

[0066] From operation 908A, flow proceeds to operation 910A, where the word is caused to be presented in a window outside of the transcription pane in the productivity application. In some examples, the word can be automatically associated with a link. The link, if accessed, can cause the portion of the transcription including the word to be presented. In other examples, the link, if accessed, can cause one or more notes associated with the word to be presented.

[0067] Flow moves from operation 910A to an end operation, and method 900A ends.

[0068] Figure 9B An exemplary method 900B is provided for presenting definitions of words and / or phrases from a custom dictionary to be included in a transcription pane of a productivity application. The method 900B begins at a start operation and flow moves to operation 902B.

[0069] At operation 902B, a selection of a word in the transcription presented in the transcription pane of the productivity application is received. Figure 9A A portion of the described real-time speech-to-text transcription example is presented.

[0070] From operation 902B, flow continues to operation 904B, where a request is received to present a definition of the second word in the productivity application user interface. The request may include selecting a dictionary icon in the transcription pane. In other examples, a definition may be requested using a right-click of a mouse and a dictionary lookup process. Other mechanisms are contemplated.

[0071] From operation 904B, flow continues to operation 906B, where a custom dictionary associated with a user account associated with the speaking user is identified. The custom dictionary may be generated based at least in part on analyzing one or more documents associated with the speaking user (e.g., the speaking user's account). Those one or more documents may include lecture notes and / or presentation documents presented in association with the current lecture and transcription instance. In other examples, the custom dictionary may be associated with a department at a university and / or a group within an organization.

[0072] Flow continues from operation 906B to operation 908B, where the definition of the word from the custom dictionary is caused to be presented in the productivity application user interface.

[0073] Flow moves from operation 908B to an end operation, and method 900B ends.

[0074] Figure 9C An exemplary method 900C for pausing and resuming real-time speech-to-text captioning in a transcription pane of a productivity application is provided. The method 900C begins at a start operation and flow moves to operation 902C.

[0075] At operation 902C, a request to pause real-time speech-to-text transcription may be received. That is, as the user speaks, subtitles may be continuously added to the transcription in the transcription pane of the productivity application, and the user may select an option in the productivity application to pause the presentation of subtitles.

[0076] From operation 902C, flow continues to operation 904C, where the presentation of the real-time speech-to-text transcript in the transcription pane is paused. That is, while the speech may still be in the process of being received and processed by the real-time speech-to-text service, the presentation of the additional captions in the transcription pane may be stopped during the pause.

[0077] From operation 904C, flow proceeds to operation 906C, where the incoming real-time speech-to-text transcription is maintained in a buffered state on the receiving device while the real-time speech-to-text transcription is paused. That is, in this example, the speech processed by the real-time speech-to-text service for the current transcription instance and subsequent transcriptions / captions are maintained in temporary storage during the pause. The transcription can be maintained in temporary storage on a server device (e.g., a server device associated with the real-time speech-to-text service) and / or on the device on which the initial pause command was received.

[0078] Flow continues from operation 906C to operation 908C, where a request to resume real-time speech-to-text transcription is received.

[0079] From operation 908C, the process continues to operation 910C, where the real-time speech-to-text transcription held in the buffered state is presented in the transcription pane. That is, all subtitles held in the temporary storage device while the pause is in effect may be automatically presented in the transcription pane along with the previously presented subtitles.

[0080] Flow continues from operation 910C to operation 912C, where the presentation of the real-time speech-to-text transcription is resumed in the transcription pane. Thus, the subtitles generated by the real-time speech-to-text service from the time the transcription was resumed from its paused state can again be continuously presented in the transcription pane.

[0081] Flow moves from operation 912C to an end operation, and method 900C ends.

[0082] Figure 10 and Figure 11 A mobile computing device 1000 is shown, such as a mobile phone, smartphone, wearable computer (such as smart glasses), tablet computer, e-reader, laptop computer, or other AR-compatible computing device, with which embodiments of the present disclosure may be implemented. Figure 10, illustrates one aspect of a mobile computing device 1000 for implementing these aspects. In a basic configuration, the mobile computing device 1000 is a handheld computer having both input and output elements. The mobile computing device 1000 typically includes a display 1005 and one or more input buttons 1010 that allow a user to enter information into the mobile computing device 1000. The display 1005 of the mobile computing device 1000 may also be used as an input device (e.g., a touch screen display). If included, an optional side input element 1015 allows further user input. The side input element 1015 may be a rotary switch, a button, or any other type of manual input element. In alternative aspects, the mobile computing device 1000 may incorporate more or fewer input elements. For example, in some embodiments, the display 1005 may not be a touch screen. In yet another alternative embodiment, the mobile computing device 1000 is a portable telephone system, such as a cellular phone. The mobile computing device 1000 may also include an optional keypad 1035. The optional keypad 1035 may be a physical keypad or a "soft" keypad generated on the touch screen display. In various embodiments, output elements include a display 1005 for displaying a graphical user interface (GUI), a visual indicator 1020 (e.g., a light emitting diode), and / or an audio transducer 1025 (e.g., a speaker). In some aspects, the mobile computing device 1000 incorporates a vibration transducer for providing tactile feedback to the user. In yet another aspect, the mobile computing device 1000 incorporates input and / or output ports, such as an audio input (e.g., a microphone jack), an audio output (e.g., a headphone jack), and a video output (e.g., an HDMI port) for sending signals to or receiving signals from an external device.

[0083] Figure 11 1 is a block diagram illustrating the architecture of one aspect of a mobile computing device. That is, a mobile computing device 1100 can implement some aspects in conjunction with a system (e.g., architecture) 1102. In one embodiment, the system 1102 is implemented as a "smartphone" capable of running one or more applications (e.g., a browser, email, calendar, contact manager, messaging client, game, and media client / player). In some aspects, the system 1102 is integrated into a computing device, such as an integrated personal digital assistant (PDA) and a wireless phone.

[0084] One or more application programs 1166 can be loaded into memory 1162 and run on or in association with operating system 1164. Examples of applications include a phone dialer program, an email program, a personal information management (PIM) program, a word processing program, a spreadsheet program, an Internet browser program, a messaging program, and the like. System 1102 also includes a non-volatile storage area 1168 within memory 1162. Non-volatile storage area 1168 can be used to store persistent information that should not be lost when system 1102 loses power. Applications 1166 can use and store information in non-volatile storage area 1168, such as email or other messages used by the email application. A synchronization application (not shown) also resides on system 1102 and is programmed to interact with a corresponding synchronization application residing on the host computer to synchronize information stored in non-volatile storage area 1168 with corresponding information stored on the host computer. It should be understood that other applications may be loaded into the memory 1162 and executed on the mobile computing device 1100, including instructions for providing and operating a real-time speech-to-text platform.

[0085] The system 1102 has a power source 1170 that can be implemented as one or more batteries. The power source 1170 can also include an external power source, such as an AC adapter or a powered docking station that replenishes or charges the batteries.

[0086] System 1102 may also include a radio interface layer 1172 that performs the function of sending and receiving radio frequency communications. Radio interface layer 1172 facilitates wireless connectivity between system 702 and the "outside world" via a communications carrier or service provider. Transmissions to and from radio interface layer 1172 are controlled by operating system 1164. In other words, communications received by radio interface layer 1172 can be disseminated to application programs 1166 via operating system 1164, and vice versa.

[0087] The visual indicator 1020 can be used to provide a visual notification, and / or the audio interface 1174 can be used to generate an audible notification via the audio transducer 1025. In the illustrated embodiment, the visual indicator 1020 is a light emitting diode (LED), and the audio transducer 1025 is a speaker. These devices can be directly coupled to the power supply 1170 so that when activated, they remain on for the duration specified by the notification mechanism, even if the processor 1160 and other components may be turned off to save battery power. The LED can be programmed to remain on indefinitely until the user takes action to indicate the power-on status of the device. The audio interface 1174 is used to provide audible signals to the user and receive audible signals from the user. For example, in addition to being coupled to the audio transducer 1025, the audio interface 1174 can also be coupled to a microphone to receive audible input, such as to facilitate a telephone conversation. According to embodiments of the present disclosure, the microphone can also be used as an audio sensor to facilitate notification control, as described below. The system 1102 may also include a video interface 1176 that enables operation of the onboard camera 1030 to record still images, video streams, and the like.

[0088] The mobile computing device 1100 implementing the system 1102 may have additional features or functionality. For example, the mobile computing device 1100 may also include additional data storage devices (removable and / or non-removable), such as magnetic disks, optical disks, or tapes. Such additional storage Figure 11 denoted by non-volatile storage area 1168.

[0089] Data / information generated or captured by the mobile computing device 1100 and stored via the system 1102 can be stored locally on the mobile computing device 1100, as described above, or the data can be stored on any number of storage media that can be accessed by the device via the radio interface layer 1172 or via a wired connection between the mobile computing device 1100 and a separate computing device associated with the mobile computing device 1100, such as a server computer in a distributed computing network, such as the Internet. It should be understood that such data / information can be accessed via the mobile computing device 1100 via the radio interface layer 1172 or via a distributed computing network. Similarly, such data / information can be readily transferred between computing devices for storage and use according to well-known data / information transmission and storage means, including email and collaborative data / information sharing systems.

[0090] Figure 12is a block diagram illustrating the physical components (e.g., hardware) of a computing device 1200 that may implement aspects of the present disclosure. The computing device components described below may have computer-executable instructions for generating, presenting, and providing operations associated with real-time speech to text transcription and translation. In a basic configuration, the computing device 1200 may include at least one processing unit 1202 and system memory 1204. Depending on the configuration and type of the computing device, the system memory 1204 may include, but is not limited to, volatile memory (e.g., random access memory), non-volatile memory (e.g., read-only memory), flash memory, or any combination of such memories. The system memory 1204 may include an operating system 1205 suitable for running one or more productivity applications. For example, the operating system 1205 may be suitable for controlling the operation of the computing device 1200. Furthermore, embodiments of the present disclosure may be implemented in conjunction with a graphics library, other operating systems, or any other application, and is not limited to any particular application or system. This basic configuration is Figure 12 1208. The computing device 1200 may have additional features or functionality. For example, the computing device 1200 may also include additional data storage devices (removable and / or non-removable), such as magnetic disks, optical disks, or tapes. Such additional storage Figure 12 12 is shown by a removable storage device 1209 and a non-removable storage device 1210.

[0091] As described above, a plurality of program modules and data files may be stored in system memory 1204. When executed on processing unit 1202, program modules 1206 (e.g., speech transcription engine 1220) may perform processing including, but not limited to, aspects described herein. According to an example, speech transcription engine 1211 may perform one or more operations associated with receiving audio signals and converting those signals into a transcription that may be displayed in a productivity application. Translation engine 1213 may perform one or more operations associated with translating a transcription in a first language into one or more additional languages. Word definition engine 1215 may perform one or more operations associated with associating definitions or notes from a notebook application with words in the transcription included in a transcription pane. Note presentation engine 1217 may perform one or more operations associated with analyzing the transcription (e.g., utilizing natural language processing and / or machine learning), identifying relevant portions of the notebook application related to the transcription, and automatically presenting the portion of the notebook application.

[0092] Furthermore, embodiments of the present disclosure may be implemented in circuits comprising discrete electronic components, packaged or integrated electronic chips comprising logic gates, circuits utilizing microprocessors, or on a single chip comprising electronic components or a microprocessor. For example, embodiments of the present disclosure may be implemented via a system on a chip (SOC) wherein Figure 12Each or many of the components shown in can be integrated onto a single integrated circuit. Such a SOC device may include one or more processing units, a graphics unit, a communication unit, a system virtualization unit, and various application functions, all of which are integrated (or "burned") onto the chip substrate as a single integrated circuit. When operated via the SOC, the functions described herein with respect to the ability of the client to switch protocols can be operated via dedicated logic integrated on a single integrated circuit (chip) with the other components of the computing device 1200. Embodiments of the present disclosure may also be implemented using other technologies capable of performing logical operations, such as AND, OR, and NOT, including but not limited to mechanical, optical, fluid, and quantum technologies. In addition, embodiments of the present disclosure may be implemented within a general-purpose computer or in any other circuit or system.

[0093] The computing device 1200 may also have one or more input devices 1212, such as a keyboard, a mouse, a pen, an audio or voice input device, a touch or slide input device, and the like. It may also include (multiple) output devices 1214, such as a display, speakers, a printer, and the like. The above devices are examples, and other devices may also be used. The computing device 1200 may include one or more communication connections 1216 that allow communication with other computing devices 1250. Examples of suitable communication connections 1216 include, but are not limited to, radio frequency (RF) transmitters, receivers, and / or transceiver circuits; universal serial bus (USB), parallel, and / or serial ports.

[0094] The term computer-readable media as used herein may include computer storage media. Computer storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, or program modules. System memory 1204, removable storage device 1209, and non-removable storage device 1210 are all examples of computer storage media (e.g., memory storage devices). Computer storage media may include RAM, ROM, electrically erasable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage, cassette tape, magnetic tape, magnetic disk storage or other magnetic storage device, or any other manufactured product that can be used to store information and can be accessed by computing device 1200. Any such computer storage media may be part of computing device 1200. Computer storage media does not include carrier waves or other propagated or modulated data signals.

[0095] Communication media may be implemented by computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media. The term "modulated data signal" may describe a signal that has one or more characteristics set or changed in such a manner as to encode the signal in the signal. By way of example, and not limitation, communication media may include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.

[0096] Figure 13 One aspect of the architecture of a system for processing data received at a computing system from a remote source such as a personal / general-purpose computer 1304, a tablet computing device 1306, or a mobile computing device 1308 is shown, as described above. The content displayed at the server device 1302 can be stored in different communication channels or other storage types. For example, a directory service 1322, a web portal 1324, a mailbox service 1326, an instant messaging repository 1328, or a social networking site 1330 can be used to store various documents. Program modules 1206 can be used by clients communicating with the server device 1302, and / or program modules 1206 can be used by the server device 1302. The server device 1302 can provide data to and from client computing devices such as a personal / general-purpose computer 1304, a tablet computing device 1306, and / or a mobile computing device 1308 (e.g., a smartphone) via a network 1315. As examples, the computer systems described herein may be implemented in a personal / general purpose computer 1304, a tablet computing device 1306, and / or a mobile computing device 1308 (e.g., a smartphone). In addition to receiving graphics data that may be used for pre-processing at a graphics originating system or post-processing at a receiving computing system, any of these embodiments of the computing device may also retrieve content from a repository 1316.

[0097] For example, various aspects of the present disclosure are described above with reference to block diagrams and / or operational diagrams of methods, systems, and computer program products according to various aspects of the present disclosure. The functions / actions noted in the blocks may not occur in the order shown in any flowchart. For example, depending on the functions / actions involved, two blocks shown in succession may actually be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order.

[0098] The description and illustration of one or more aspects provided in this application are not intended to limit or restrict in any way the scope of the present disclosure as claimed. The aspects, examples and details provided in this application are considered sufficient to convey all and enable others to make and use the best mode of the claimed disclosure. The claimed disclosure should not be interpreted as being limited to any aspect, example or detail provided in this application. Whether shown and described in combination or separately, various features (both structures and methods) are intended to be selectively included or omitted to produce an embodiment with a specific feature set. After providing a description and illustration of the present disclosure, those skilled in the art can envision changes, modifications and alternative aspects that do not depart from the broader scope of the claimed disclosure within the spirit of the broader aspects of the general inventive concepts implemented in this application.

[0099] The various embodiments described above are provided as illustrations only and should not be construed as limiting the claims appended hereto. Those skilled in the art will readily appreciate that various modifications and changes may be made without following the exemplary embodiments and applications illustrated and described herein, and without departing from the true spirit and scope of the appended claims.

Claims

1. A computer-implemented method for integrating speech-to-text transcription in a productivity application, the computer-implemented method comprising: sending, by the first device, a request to access a real-time speech-to-text transcription of an audio signal currently being received by the second device; causing the real-time speech-to-text transcription to be presented in a transcription pane of a productivity application user interface on the first device; receiving a selection of a word in the presented transcription; receiving a request to cause a definition for the word to be presented in the productivity application user interface; identifying a custom dictionary associated with a user account of the second device; as well as A definition for the word from the custom dictionary is caused to be presented in the productivity application user interface.

2. The computer-implemented method of claim 1 , wherein the request to access the real-time speech-to-text transcription includes an accession code.

3. The computer-implemented method of claim 2 , wherein the opt-in code provides access to the real-time speech-to-text transcription for any computing device that provides the opt-in code to a real-time speech-to-text transcription service that receives the audio signal.

4. The computer-implemented method of claim 1 , further comprising: receiving a selection of a second word in the presented transcription; receiving a request to drag the second word from the transcription pane and drop the second word into a window outside of the transcription pane in the productivity application; as well as The second word is caused to be presented in the window outside the transcription pane in the productivity application. 5 . The computer-implemented method of claim 1 , wherein the custom dictionary is generated based at least in part on analyzing a document presented by the second device while the audio signal is being received by the second device.

6. The computer-implemented method of claim 5, wherein analyzing the document comprises: A neural network that has been trained to identify topical topics is applied to the documents.

7. The computer-implemented method of claim 1 , further comprising: receiving a selection of a second word in the presented transcription; receiving a request to cause a definition for the second word to be presented in the productivity application user interface; determining that the second word cannot be located in the custom dictionary; as well as A selectable option is generated to perform a web search for the second word to be presented.

8. The computer-implemented method of claim 1 , further comprising: receiving a selection of a second word in the presented transcription; receiving a request to associate an annotation with the second word in the productivity application user interface; as well as The annotation is associated with the second word in the productivity application user interface.

9. The computer-implemented method of claim 1 , further comprising: receiving a request to translate the transcription from a first language into which the audio signal was originally transcribed into a second language; as well as The real-time speech-to-text transcription is caused to be presented in the transcription pane in the second language.

10. The computer-implemented method of claim 1 , further comprising: receiving a request to pause the real-time speech-to-text transcription; pausing the presentation of the real-time speech-to-text transcription in the transcription pane; maintaining incoming real-time speech-to-text transcription in a buffered state on the first device while the real-time speech-to-text transcription is paused; receiving a request to resume the real-time speech-to-text transcription; causing the real-time speech-to-text transcription maintained in the buffered state to be presented in the transcription pane; as well as The presentation of the real-time speech-to-text transcription in the transcription pane is resumed.

11. A system for integrating speech-to-text transcription in a productivity application, comprising: a memory for storing executable program code; as well as one or more processors functionally coupled to the memory, the one or more processors responsive to computer-executable instructions contained in the program code and operative to: sending, by the first device, a request to access a real-time speech-to-text transcription of an audio signal currently being received by the second device; causing the real-time speech-to-text transcription to be presented in a transcription pane of a productivity application user interface on the first device; receiving a selection of a word in the presented transcription; receiving a request to cause a definition for the word to be presented in the productivity application user interface; identifying a custom dictionary associated with a user account of the second device; as well as A definition for the word from the custom dictionary is caused to be presented in the productivity application user interface.

12. The system of claim 11, wherein the request to access the real-time speech-to-text transcription includes an accession code.

13. The system of claim 11 , wherein the one or more processors are further responsive to the computer-executable instructions contained in the program code and operate to: receiving a selection of a second word in the presented transcription; receiving a request to drag the second word from the transcription pane and drop the second word into a window outside of the transcription pane in the productivity application; as well as The second word is caused to be presented in the window outside the transcription pane in the productivity application.

14. The system of claim 11 , wherein the one or more processors are further responsive to the computer-executable instructions contained in the program code and operate to: The custom dictionary is generated based on analyzing a document presented by the second device while the audio signal is being received by the second device.

15. The system of claim 11, wherein the one or more processors are further responsive to the computer-executable instructions contained in the program code and operate to: receiving a selection of a second word in the presented transcription; receiving a request to associate an annotation with the second word in the productivity application user interface; as well as The annotation is associated with the second word in the productivity application user interface.

16. The system of claim 11, wherein the one or more processors are further responsive to the computer-executable instructions contained in the program code and operate to: applying a machine learning model to the presented transcription, wherein the machine learning model has been trained to classify text into topic types; categorizing said transcripts presented into thematic types; receiving an instruction to save the rendered transcript; as well as The presented transcription is automatically saved in a location corresponding to the subject type.

17. The system of claim 11, wherein the one or more processors are further responsive to the computer-executable instructions contained in the program code and operate to: receiving a request to pause the real-time speech-to-text transcription; pausing the presentation of the real-time speech-to-text transcription in the transcription pane; maintaining incoming real-time speech-to-text transcription in a buffered state on the first device while the real-time speech-to-text transcription is paused; receiving a request to resume the real-time speech-to-text transcription; causing the real-time speech-to-text transcription maintained in the buffered state to be presented in the transcription pane; as well as The presentation of the real-time speech-to-text transcription in the transcription pane is resumed.

18. A computer-readable storage device comprising executable instructions that, when executed by one or more processors, facilitate integration of speech-to-text transcription in a productivity application, the computer-readable storage device comprising instructions executable by the one or more processors to: sending, by the first device, a request to access a real-time speech-to-text transcription of an audio signal currently being received by the second device; causing the real-time speech-to-text transcription to be presented in a transcription pane of a productivity application user interface on the first device; receiving a selection of a word in the presented transcription; receiving a request to cause a definition for the word to be presented in the productivity application user interface; identifying a custom dictionary associated with a user account of the second device; as well as A definition for the word from the custom dictionary is caused to be presented in the productivity application user interface.

19. The computer-readable storage device of claim 18, wherein the instructions are further executable by the one or more processors to: receiving a selection of a second word in the presented transcription; receiving a request to transfer the second word from the transcription pane to a window in a productivity application external to the transcription pane; and The second word is automatically caused to be presented in the window in the productivity application outside of the transcription pane.

20. The computer-readable storage device of claim 18, the instructions being further executable by the one or more processors to: receiving a request to pause the real-time speech-to-text transcription; pausing said presentation of said real-time speech-to-text transcription; maintaining incoming real-time speech-to-text transcription in a buffered state on the first device while the real-time speech-to-text transcription is paused; receiving a request to resume the real-time speech-to-text transcription; causing the real-time speech-to-text transcription maintained in the buffered state to be presented in the transcription pane; as well as The presentation of the real-time speech-to-text transcription in the transcription pane is resumed.

Citation Information

Patent Citations

  • Voice recognition device and system

    CN109841209A

  • Word flow annotation

    CN109844854A