Systems and methods for GPT guided neural punctuation for session speech

By identifying and correcting the defluency in the decoded audio data in the automatic speech recognition system, the problem of truncation in the middle of sentences is solved, and the readability and transcription quality of the audio data are significantly improved.

CN120129939APending Publication Date: 2025-06-10MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380075124.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-01
Filing Date
2023-10-14
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing automatic speech recognition systems are prone to truncation in the middle of sentences when segmenting audio data, resulting in a decrease in readability and quality of transcription, especially when dealing with natural language and spoken communication.

Method used

The readability of the audio data is improved by identifying and correcting the defluency in the decoded audio data, such as interjections, repetitions and words with low recognition scores. The system uses neural network models and teacher models to identify and correct inefficiency and generate more accurate and readable transcripts.

Benefits of technology

It significantly improves the readability and transcription quality of audio data, avoids truncation in the middle of sentences, and enhances the overall performance of the speech recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120129939A_ABST
    Figure CN120129939A_ABST
Patent Text Reader

Abstract

Some disclosed embodiments relate to obtaining decoded audio data comprising spoken utterances identified in the audio data, and identifying non-fluency in the decoded audio data. Upon determining that modifying the non-fluency will improve the readability score of the decoded audio data, the system generates a particular modification to modify the non-fluency and applies the particular modification to the decoded audio data. Updated decoded audio is then generated that reflects the particular correction. The updated decoded audio data has an improved readability over the decoded audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Automatic speech recognition systems and other speech processing systems are used to process and decode audio data to detect speech utterances (e.g., words, phrases, and / or sentences). The processed audio data is then used in various downstream tasks such as search-based queries, speech-to-text transcription, language translation, closed captioning, etc. Often, the processed audio data needs to be segmented into multiple audio segments before being transmitted to downstream applications or streamed to other processing.

[0002] Conventional systems are configured to perform audio segmentation for today's continuous speech based on timeout-driven logic. In such speech recognition systems, at the end of a detected word (i.e., when the audio has "timed out"), the audio is segmented after a certain amount of silence has elapsed. This timeout-based segmentation does not take into account the fact that someone can naturally pause between sentences while thinking about what they are going to say next. As a result, the words are often truncated in the middle of a sentence before someone has finished stating the sentence. This reduces the quality of the output of the data consumed by downstream post-processing components such as a punctuator or a machine translation component. Previous systems and methods have been developed that include neural network-based models that combine current auditory information and corresponding language signals to improve segmentation. However, even such methods, while superior to timeout-based logic, can suffer from over-segmenting the audio, resulting in some of the same problems as timeout-based logic segmentation.

[0003] For example, Figure 1A A conventional automatic speech recognition system including a decoder 104, a punctuator 108, and a user display 112 is depicted. Figure 1B An example of a conventional flow of the speech recognition system shown in FIG. 1 is illustrated. As shown, audio 102 includes: a spoken utterance (e.g., a spoken utterance such as audio 114 "I'm going to walk the dog at 10 o'clock tonight. After I walk him, I'll feed him") is used as an input to decoder 104, which decodes audio 102 and outputs decoded segment 106 (e.g., decoded segment 118 "I'm going to walk the dog at 10 o'clock tonight. I'm going to"). This decoded segment 106 is an input to punctuator 108, which punctuates decoded segment 106 to output punctuated output 110 (e.g., punctuated output 122 "I'm going to walk the dog at 10 o'clock tonight. I'm going to."). This punctuated output 110 is then transmitted to user display 112 for display to the user.

[0004] It is noted that, as Figure 1BAs shown, since it includes the partial sentence "I want", the system does not correctly punctuate the punctuated output. This reduces the viewing quality of the transcription on the user's display because the incorrect punctuated output is presented to the user. The system may be able to return and re-output a corrected version of the output, but a conventional system replaces the already-displayed incorrect output with the newly corrected output, which may be confusing to the user viewing the user's display, which dynamically changes with different outputs of the same portion of the audio data.

[0005] In addition, in some instances, due to certain disfluencies present in the decoded segments, the system is unable to correctly punctuate. These disfluencies stem from the different natures of spoken and written communication. For example, when people are speaking, they can pause, stutter, repeat words, or use interjections such as "um" (e.g., filler words). Due to these disfluencies, it may be difficult to produce a readable transcription of the spoken words.

[0006] In view of the foregoing, there is still a need for improved systems and methods for segmenting audio in order to generate more accurate, more readable transcriptions corresponding to the complete spoken utterances included in the audio and high-quality displays of those transcriptions.

[0007] The subject matter claimed herein is not limited to embodiments that solve any disadvantages or operate only in environments such as those described above. Rather, this background is provided only to illustrate an exemplary technical field in which some embodiments described herein may be practiced. Summary of the Invention

[0008] Embodiments of the present disclosure include systems and methods for generating improved transcriptions for spoken utterances identified in input audio data. Specifically, embodiments of the present disclosure relate to systems and methods for improving the readability of decoded audio data.

[0009] For example, systems are provided for obtaining decoded audio data including spoken utterances identified in audio data, identifying disfluencies in the decoded audio data, and determining that correcting the disfluencies will improve the readability of the decoded audio data. Once the system has identified the disfluencies and determined that they should be corrected to improve the readability of the decoded audio data, the system generates a specific correction to correct the disfluencies and applies the specific correction to the decoded audio data. Finally, updated decoded audio data is generated, which reflects the specific correction applied to the decoded audio data. In such an instance, the updated decoded audio data is characterized by improved readability over the decoded audio data.

[0010] The present invention content is provided to introduce in a simplified form a series of concepts that will be further described in the following detailed description. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to assist in determining the scope of the claimed subject matter.

[0011] Additional features and advantages will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the teachings herein. The features and advantages of the present invention may be realized and obtained by the means and combinations particularly pointed out in the appended claims. The features of the present invention will become more fully apparent from the following description and the appended claims, or may be learned by practice of the invention as described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] To describe the manner in which the above and other advantages and features can be obtained, a more particular description of the subject matter briefly described above will be presented by reference to specific embodiments shown in the drawings. It should be understood that these drawings only depict typical embodiments and should not be considered as limiting the scope. The embodiments will be described and explained with additional specificity and detail by using the drawings, in which:

[0013] Figures 1A to 1B Various example embodiments of existing speech recognition and transcription generation systems are shown.

[0014] Figure 2 A computing environment is shown in which a computing system incorporates and / or is used to execute the disclosed aspects of embodiments of the present disclosure.

[0015] Figures 3A to 3C Various examples and / or stages of a flowchart of a system configured to coordinate the transmission of a transcription output after a speech segment has been punctuated are shown.

[0016] Figures 4A to 4B An example flowchart for marking and correcting disfluencies that occur in decoded audio data during training and used during operation is shown.

[0017] Figure 5 An embodiment of a flowchart with multiple actions for generating an improved speech transcription by identifying and correcting disfluencies that occur in decoded audio data is shown.

[0018] Figure 6 An example embodiment of identifying and correcting disfluencies associated with punctuation marks included in decoded audio data is shown.

[0019] Figure 7 An example embodiment of identifying and correcting disfluencies associated with repeated words included in decoded audio data is shown.

[0020] Figure 8 Illustrates an example embodiment of identifying and correcting disfluencies associated with interjections included in decoded audio data.

[0021] Figure 9 Illustrates an example embodiment of identifying and correcting disfluencies associated with words having low recognition scores included in decoded audio data.

[0022] Figure 10 Illustrates an example embodiment of identifying and correcting disfluencies associated with a mismatched reading level between decoded audio data and a target reader.

[0023] Figure 11 Illustrates an example embodiment of identifying and correcting disfluencies associated with confidential information included in decoded audio data.

[0024] Figure 12 Illustrates an example embodiment of determining when and how to display updated decoded audio data.

[0025] Figure 13 Illustrates an example embodiment of displaying updated decoded audio data, including displaying an embedded link that directs a user to an original version of the decoded audio data.

[0026] Figure 14 Illustrates an example embodiment of different ways of displaying updated decoded audio data based on one or more attributes of the updated decoded audio data.

[0027] Figure 15 Illustrates an example embodiment of a user display having at least two different modes, one example embodiment for accessibility and one example embodiment for readability. Detailed Description

[0028] Embodiments of the present disclosure relate to systems and methods for generating transcripts of audio data. In this regard, it should be understood that some embodiments of the present disclosure specifically relate to improved systems and methods for improving the segmentation and punctuation of transcripts of audio data by avoiding the output of incomplete language fragments. Embodiments of the present disclosure provide many technical advantages over existing systems.

[0029] Cognitive services such as ASR systems cater to a wide variety of customers. Each customer wishes to optimize their experience with respect to wait time, accuracy, and cost of goods sold (COGS). Split improvements are key to affecting punctuation as the two are closely related. Many existing systems, including powerful neural network-based methods, incur high wait times and / or COGS. As such, these models are not viable for wait-time sensitive customers (e.g., as in streaming audio applications). Even for wait-time tolerant customers, existing speech recognition services introduce mid-sentence breaks after long stretches of uninterrupted speech (over-splitting). Readability is reduced when such breaks occur.

[0030] However, semantic segmenters, such as those included in the embodiments disclosed herein, are capable of achieving significant readability improvements without sacrificing accuracy, while improving the rendering of individual sentences more quickly compared to current generation. Accordingly, the embodiments of the present disclosure achieve significant improvements for all word-based languages, even without a neural model for segmentation. Additionally, this further improves machine translation performance.

[0031] One advantage of the embodiments of the present disclosure is that they provide significant improvements in the readability of closed captioning services. Such embodiments improve the accuracy of punctuation, which can in turn also help improve the overall functionality of semantic segmenters. Depending on customer constraints, a user can select from different parameters in order to customize the trade-off between wait time, precision, and COG. Such an approach allows for the best system / service level combination for both worlds (both segmentation and punctuation) given customer constraints.

[0032] Attention is now turned to Figure 2 , Figure 2 FIG. 200 is shown, which also includes (via network 230) (multiple) third-party systems 220 in communication with a computing system 210, the computing system 210 incorporating and / or operative to execute the disclosed aspects of the embodiments of the present disclosure. The (multiple) third-party systems 220 include one or more processors 222 and one or more hardware storage devices 224.

[0033] The computing system 210, for example, includes one or more processors (such as one or more hardware processors 212) and a memory (i.e., (multiple) hardware storage devices 240) storing computer-readable instructions 118, where one or more of the (multiple) hardware storage devices 240 can accommodate any number of data types and any number of computer-executable instructions 218. When the computer-executable instructions 218 are executed by one or more processors 212, the computing system 210 is configured to implement one or more aspects of the embodiments of the present disclosure through the computer-executable instructions 218. The computing system 210 is also shown to include (multiple) user interfaces 214 and (multiple) input / output (I / O) devices 216.

[0034] As Figure 2 shown, the (multiple) hardware storage devices 240 are shown as a single storage unit. However, it should be understood that the (multiple) hardware storage devices 240 are distributed storage that is distributed to several separate and sometimes remote systems and / or (multiple) third-party systems 220. The computing system 210 may also include a distributed system where one or more components of the computing system 210 are maintained / run by different discrete systems that are far from each other and each perform different tasks. In some instances, multiple distributed systems, such as in a distributed cloud environment, perform similar and / or shared tasks for implementing the functions of the present disclosure.

[0035] As described herein, the (multiple) hardware storage devices 240 are configured to store and / or cache different data types in the memory, including: audio data 241, decoded audio data 242, punctuated data 243, and updated decoded audio data 244. The (multiple) hardware storage devices 240 also store an ASR system 245, and the ASR system 245 includes at least a punctuator 246 and a disfluency tagger 247.

[0036] The audio data 241 includes both natural language audio and analog audio. The audio is obtained from multiple locations and applications. In some instances, natural language audio is extracted from previously recorded or downloaded files, such as video recordings with audio or audio-only recordings). Some examples of recordings include videos, podcasts, voicemails, voice memos, songs, etc. Natural language audio is also extracted from live streaming content, which is continuous live speech such as news broadcasts, phone calls, virtual or in-person meetings, etc. In some instances, previously recorded audio files are streamed. The audio data includes spoken utterances with or without a corresponding clean speech reference signal. Natural audio data is recorded from multiple sources, including applications, meetings including one or more speakers, surrounding environments including background noise and human speakers, etc. It should be understood that natural language audio includes one or more spoken languages of the world.

[0037] The decoded audio data 242 includes speech tags corresponding to the spoken utterances identified in the audio data 241 as the output of the ASR system. The decoded audio data 242 is then punctuated by the punctuator 246 using soft and / or hard punctuation. The punctuated data 243 is then analyzed by the disfluency tagger 247 configured to identify and tag disfluencies in the punctuated data 243. Among other disfluencies discussed below Figures 6 to 11 which may be related to interjections or filler words, repeated words, poor initial punctuation, low recognition score words, confidential words, mismatched reading comprehension score words, the system then determines whether the disfluency should be corrected and what corrections should be made. If the disfluency is corrected, the system generates updated decoded audio data 244 (e.g., corrected tags and / or punctuation marks corresponding to the identified spoken utterances).

[0038] Attention is now turned to Figures 3A to 3C , Figures 3A to 3C various examples and / or stages of a flowchart of a system configured to coordinate the transmission of a transcription output after a speech segment has been punctuated are shown. Figures 3A to 3C A decoder 304, a punctuator 308, a coordinator 312, and a user display 316 are shown. Attention is first turned to Figure 3A where the decoder 304 is configured to decode the spoken utterances identified in the input audio (e.g., the streaming audio data 302A) associated with the speaker 301 to generate decoded audio segments (e.g., the decoded segment 306A). In some instances, the decoded segment includes a speech data representation and / or a speech data transcription (i.e., speech token tags). The decoded segment is then punctuated by the punctuator 308 at one or more identified language boundaries within the decoded segment. In some instances, the decoder 304 is configured to identify the language boundaries. Additionally or alternatively, the punctuator 308 is configured to identify the language boundaries within the decoded segment and / or confirm previously identified language boundaries by the decoder. In some instances, the language boundaries are detected in the streaming audio data prior to transcription by the decoder.

[0039] A language boundary is a representative marker that is identified and / or generated to indicate the end of a complete sentence. In other words, a language boundary exists at the end of a complete sentence, typically just after the end of the last word of the sentence. Based on the language boundary, the correct punctuation to be placed at the language boundary (i.e., just after the last word of the sentence) can be determined. Additionally, a text or audio segment can be further segmented into at least a portion that includes a complete sentence and a second portion that may or may not include another complete sentence. It should be understood that a language boundary can also be detected at the end of an audio or text phrase where a speaker or writer has intentionally spoken or written a sentence fragment. In some instances, a language boundary is predicted when a speaker has paused for a predetermined amount of time. In some instances, a language boundary is determined based on the context of a first segment (or a first part of a segment) relative to a subsequent segment (or a subsequent part of the same segment).

[0040] Once the decoded segment 306A has been punctuated by the punctuator 308, the punctuated segment 310A is analyzed by the coordinator 312, which is configured to detect one or more parts of the correctly segmented and punctuated portions (e.g., complete sentences) of the punctuated segment 310A and output only the complete sentences (e.g., output 314A) to the user display 316.

[0041] Some example user displays or user interfaces include audiovisual displays such as a television or computer monitor, display outputs such as a mobile device or tablet, and interactive displays that receive user input. In some instances, the outputs are displayed continuously, with only a limited number of outputs displayed on the display, such as in the case of live captions for streaming audio / visual data. In some instances, the outputs are appended to each other to form a final transcription, and each output is displayed as part of the ongoing transcription as it is generated. Such a transcription can be displayed via a scrollable user interface. In some instances, the outputs are displayed only when all final outputs (i.e., correctly segmented and punctuated outputs) have been generated.

[0042] In some instances, the output 314A includes grammatically complete sentences, while in some instances, the output 314A includes grammatically incomplete or incorrect transcriptions but is still correctly segmented and punctuated because the output 314A includes portions of the initially decoded segment corresponding to intentional sentence fragments and / or intentional consecutive sentences. In some instances, the decoded segment includes a single complete sentence, multiple complete sentences, partial sentences, or a combination of complete and partial sentences in any order.

[0043] Now turn attention to Figure 3B , Figure 3B shown by Figure 3AAn example of input audio processed by the automatic speech recognition system shown. For example, streaming audio data 302B is obtained, which includes the spoken utterance "I will walk the dog at 10 o'clock tonight and I will feed him after walking him." The decoder 304 decodes the first audio segment and outputs a decoded segment 306B that includes "I will walk the dog at 10 o'clock tonight and I will". In this instance, the streaming audio data is initially segmented in this way due to a pause (i.e., speaker silence) indicated by "..." in the input audio data. Then, the punctuator 308 punctuates the decoded segment and outputs a punctuated segment 310B (e.g., "I will walk the dog at 10 o'clock tonight. I will.").

[0044] Then, the coordinator 312 is configured to detect which portions of the punctuated segments are complete segments and output only those one or more portions that are complete sentences. As Figure 3B shown, the coordinator 312 recognizes that "I will walk the dog at 10 o'clock tonight." is a complete sentence and generates an output 314B corresponding to the first portion of the punctuated segment. The second portion of the punctuated segment ("I will.", see portion 311) is determined to be an incomplete sentence and is therefore retained in the coordinator (and / or storage buffer) without being generated as an output that can be transmitted to the user display 316. The first portion of the punctuated segment (e.g., output 314B) is transmitted to the user display and presented on the user display.

[0045] Now turn attention to Figure 3C , Figure 3C which shows Figure 3B a continuation of the speech processing described in

[0046] In the case where no language boundary is detected in the initial segment, the computing system avoids outputting the initial segment of the decoded streamed audio data and continues to decode the streamed audio data until a subsequent segment of the decoded streamed audio data is generated and appended to the initial segment of the decoded streamed audio data. In this way, the system analyzes the appended segment to determine whether a language boundary exists.

[0047] In some embodiments, the computing system utilizes a cache that promotes improved timing for the output of different speech segments. For example, the system stores the initial segment of the decoded streamed audio data in the cache. Then, after outputting the first part of the initial segment, the system clears the cache for the first part of the initial segment of the decoded streamed audio data. In a further embodiment, while clearing the cache for the first part of the initial segment of the decoded streamed audio data, the system retains the second part of the segment of the decoded streamed audio data in the cache. Embodiments that utilize the cache in this way effectively manage the storage space of the cache by deleting the data that has already been output and retaining the data that is needed to continue to accurately generate the punctuated output, thereby improving the functionality of the computing system.

[0048] For example, when the second part of the initial segment of the decoded streamed audio is retained in the cache, the system is able to store subsequent segments of the decoded streamed audio data in the cache, where the subsequent segments of the decoded streamed audio data are appended to the second part of the initial segment of the decoded streamed audio data to form a new segment of the decoded streamed audio data.

[0049] The system then determines whether there is a subsequent language boundary within the new segment of the decoded streamed audio data. When it is determined that a subsequent language boundary exists, the system applies new punctuation at the subsequent language boundary and outputs the first part of the new segment of the streamed audio data that ends at the subsequent language boundary, while avoiding outputting the second part of the new segment that is temporally after the second part of the initial segment.

[0050] Embodiments of the present disclosure relate to providing systems, methods, and devices for automatic segmentation and punctuation as well as user-initiated segmentation and / or punctuation. For example, a decoder can be configured to generate decoded segments based on user commands and / or detected keywords identified within the streamed audio data. Similarly, a punctuator can be configured to punctuate the decoded segments based on user commands and / or detected keywords within the decoded segments.

[0051] Attention is now turned to Figure 4A , Figure 4A which shows an example flowchart of using a disfluency tagger to improve the readability of decoded audio data. As Figure 4AAs shown, once the system has generated a valid output, the system is able to analyze the output for any disfluencies. Although Figure 4A illustrates the disfluency tagging process after respunctuation (i.e., verification), it should be understood that disfluency tagging can also occur at different locations in the speech processing pipeline, including after the speech recognition backend or after the neural punctuator.

[0052] As Figure 4A shown, audio data 402 is processed by a speech recognition backend 404 (e.g., an ASR model), which outputs decoded audio data 406 (e.g., speech labels corresponding to the spoken utterances recognized in the audio data 402). The neural punctuator 408 then generates an initial set of punctuation symbols for the decoded audio data, applies the initial set of punctuation symbols, and generates punctuated data 410. In some instances, the system identifies tags that can be normalized (e.g., numbers, dates, addresses, etc.) and performs normalization 412 on the identified tags. Additionally, it should be understood that Figures 3A to 3C the coordinator shown in can be used to suppress and verify different parts of the decoded audio data to ensure optimal segmentation before the neural punctuator, normalization / respunctuation, and / or disfluency tagger. Thus, in some instances, Figure 4A the data 414 shown represents Figure 3A the output 314A shown.

[0053] Figure 4A Also shown is a disfluency tagger 416, which is configured to identify and tag disfluencies in the decoded audio data. Among other disfluencies to be discussed below with reference to Figures 6 to 11 these disfluencies can be related to interjections or filler words, repeated words, poor initial punctuation.

[0054] As Figure 4A shown, after punctuating the decoded audio data, the system applies a series of post-processing to further improve the punctuation. Some of these post-processing steps include using a teacher model 420 (e.g., a large-scale pre-trained model (LS-PTM), e.g., a generative model, or a new tokenizer based on GPT) as a weak tagger, using the identified disfluencies as additional guidance for improvement in the punctuation, and using the teacher model 420 as a predictor of punctuation points. In some instances, the teacher model 420 is used as a readability scorer, which reports how the readability scores of different sets of words that form similar sentences compare. When different sentences include the same words and only differ in punctuation, the scorer can report which sentence is more readable than the other(s).

[0055] Embodiments of the present disclosure relate to further improvements to the readability scoring process, where systems and methods are provided for creating weakly labeled punctuation marks in a fully automated manner, without the need for manual labeling efforts to generate training data. The weakly labeled data or a subset of the best weak labels is used as training data to fine-tune one or more production models. Subsequently, the system selects the model that produces the most readable text as determined by a scorer and / or a human evaluator.

[0056] The purpose of the teacher model 420 is to determine what the best punctuation label is, particularly for decoded audio data that includes speech disfluencies. In some instances, to save computational overhead and time, the teacher model 420 is only applied to the portion of the decoded audio data where disfluencies have been labeled / identified.

[0057] There are many different ways in which the teacher model 420 can correct disfluencies and improve the readability of the decoded audio data. For example, if there is a disfluency in the current sentence (e.g., "And um brought some new clothes."), it can be merged with its previous sentence (e.g., "I went to the mall."). In some instances, the decision of whether to merge sentences is based on determining whether the character count of the previous sentence plus the character count of the current sentence is below a predetermined maximum character length.

[0058] For sentence merging to occur, the teacher model 420 scores the previous and current sentences separately. Additionally, the teacher model 420 is configured to divide by the number of words in both the previous and current sentences to find the average score per word (i.e., the raw score). Then, the model calculates the average score per word for the two sentences joined together as one sentence using various different punctuation calculations. The sentences will be merged according to whatever punctuation is associated with the highest score. For example, in some instances, the model uses a comma to merge the sentences and calculates the average score per word for the comma score (e.g., "I went to the mall , And um brought some new clothes"). In some instances, the model only uses the character space between the two sentences to merge sentences without any punctuation and calculates a punctuation-less score (e.g., "I went to the mall and um brought some new clothes."). If the comma score is higher than the punctuation-less score, or other scores based on one or more different connecting punctuation marks, the previous sentence and the current sentence will be merged into one sentence, as the final output with a comma placed between the different sentences (i.e., where the original sentence boundaries were identified).

[0059] In another example, sometimes the disfluency does not occur at the beginning or the end of the sentence, so different techniques are needed to correct the disfluency. In such instances, the system creates a version of the original sentence without the disfluency, called the modified sentence. The system also creates multiple new expected sentences (e.g., the current modified sentence and the next modified sentence) by taking the modified sentence and splitting the modified sentence where the disfluency is marked. For example, if the original sentence is "I went kayaking today, um, but I felt cold." The disfluency is "um". Then the system generates the modified sentence by removing "um" (e.g., "I went kayaking today, but I felt cold."). Subsequently, the system uses the sentence boundary defined at the time position of the disfluency to split the modified sentence into two different sentences. Thus, the current modified sentence is: "I went kayaking today." And the next modified sentence is: "But I felt cold today."

[0060] Then, the system can score whether the next modified sentence should be merged with the current modified sentence and which punctuation mark should be used to merge the sentences without disfluency. For example, for two sentences now, the above merging logic can be applied to determine whether the merge results in a higher score, and which punctuation mark will result in the highest readability score if the sentences are merged.

[0061] Thus, in some instances, the system compares at least four different scores. A first score is calculated based on merging the sentences with a comma inserted just before the disfluency (e.g., "I went kayaking today, but I felt cold."). A second score is calculated based on merging the sentences with a character space (e.g., "I went kayaking today but I felt cold."). A third score is calculated based on merging the two sentences with a period (e.g., "I went kayaking today. But I felt cold."). A fourth score is calculated based on merging the sentences with a question mark (e.g., "I went kayaking today? But I felt cold."). For the third and fourth scores, the system checks whether the average per-word readability score of the two new sentences is better than the average per-word readability score of the original sentence. If this is the case, the system only considers the third option as a potential competitor and sets the readability score to the average of the two new sentences, otherwise the system considers it to have an infinite score.

[0062] If the selected score from the previous step is worse than the readability score of the original sentence, the system repeats the process, selects the best option among the generated options, but retains the disfluencies in the different sentence options. Additionally, if the disfluency appears at the end of the sentence, or if the sentence has a disfluency, the system considers merging the sentence with the subsequent sentence in order (contrary to the previous sentence as described above). The merging logic is similar to the above options, including avoiding merging sentences if the character length of the merged sentence exceeds a predetermined maximum character length. Alternatively, if there is no subsequent sentence to merge with the current sentence, the system is configured to adjust the punctuation, e.g., if a disfluency appears at the end of the sentence, adjust the punctuation at the end of the sentence and select the punctuation from among a period, a question mark, a comma, an exclamation point, or other punctuation marks, etc.

[0063] The teacher model 420 will perform all of the above analyses on each disfluency in each sentence such that the final tagged output is affected by multiple analysis phases as described above. Since most punctuation models are trained on written text generated in a standard-based and / or polished written communication style, conventional punctuators have difficulty handling and accurately punctuating decoded audio data that includes speech disfluencies that do not typically occur in written text. Spoken language is very spontaneous and includes disfluencies such as "um", "uh", "you know", "right", and repeated words. Even human taggers often find it difficult to tag disfluencies in text. This is why leveraging the teacher model 420 (i.e., a large pre-trained model) allows the ASR system to generate punctuated text while considering and addressing the deficiencies in the original speech data.

[0064] However, due to the size of the LS-PTM (e.g., the teacher model 420), a large pre-trained model is used as the teacher model to generate weak tags (selected from the sentence with the highest readability score as described above), and the weak tags are used as training data 422 to train a computationally more efficient punctuation model (e.g., the student model 426). Further improvement is achieved when the system is able to repunctuate the speech recognition output in sentence blocks that include sentences already defined by the semantic segmentation model. With this improvement, the system can improve the performance of the model, e.g., perform repunctuation with the LS-PTM after every seven sentences instead of every two sentences.

[0065] After training the student model 426, the system replaces the neural punctuator 408 with the student model that will be used during runtime. Now turn the attention to Figure 4B , Figure 4B FIG. shows a process flow diagram for performing speech-to-text transcription using the trained student model.

[0066] After fine-tuning the student model and selecting the student model with the highest readability score for the decoded audio data from various student models, the normalized and respunctuated 412 and the disfluency tagger 416 are then used, and the student model is used to process the decoded audio data 406 to generate updated decoded audio data 428. Then, the system is capable of sending the updated decoded audio data 428 and displaying the updated decoded audio data 428 on the user display 430. Because the updated decoded audio data 428 has a higher readability score than the original decoded audio data, the user experience is greatly improved. In some instances, during operation, the disfluency tagger 416 marks and corrects the disfluencies identified in the audio data received as input to the disfluency tagger. In some instances, during runtime, the audio data output from the normalized and respunctuated 412 is sent to and displayed at the user display 430 without undergoing tagging.

[0067] Attention is now turned to Figure 5 , Figure 5 which shows a flowchart 500 including various actions (action 510, action 520, action 530, action 540, action 550, and action 560) associated with an exemplary method that can be implemented by a computing system 210 for generating an improved speech transcription. For example, Figure 9 shows a method for generating updated decoded audio data by identifying and correcting disfluencies.

[0068] The first action shown includes the action of obtaining decoded audio data including the spoken utterances identified in the audio data (action 510). Then the system identifies the disfluencies in the decoded audio data (action 520). Subsequently, the system determines whether correcting the disfluencies will improve the readability score of the decoded audio data (action 530). The action of determining whether the disfluencies detected in the audio should be corrected or left unchanged, in order to improve the readability of the transcription of the utterance, the system applies a machine learning algorithm that has been trained to calculate scores for the readability of different utterances. In some cases, when making this determination, the system applies the same model to different types of spoken utterances. In other cases, different trained models are applied to different contexts (e.g., different data topics, different languages, different speakers, different educational levels of the speaker and / or intended audience, etc.).

[0069] In many instances, it is decided to fix the disfluencies to improve readability. However, in some embodiments, it is determined that the disfluencies should be retained to maintain readability and any changes will negatively impact the readability of the utterance. In fact, some disfluencies may be part of an intentional and deliberate speaking style, such as emphasizing a particular term or concept.

[0070] Once it is determined that correcting the disfluency will improve the readability score of the decoded audio data, the system generates a specific correction configured to correct the disfluency (action 540) and applies the specific correction to the decoded audio data (action 550). Finally, the system generates updated decoded audio data reflecting the specific correction (action 560). The updated decoded audio data has an improved readability score compared to the decoded audio data of the spoken utterance.

[0071] In some instances, decoded audio data is obtained by first acquiring audio data including a spoken utterance from a speaker and continuously decoding the streaming audio data to generate the decoded audio data. Then, the system determines whether a language boundary exists within an initial segment of the decoded audio data. When it is determined that a language boundary exists, a first portion of the initial segment that is temporally before the language boundary (e.g., prior to the language boundary) and a second portion of the initial segment that is temporally after the language boundary are identified. Subsequently, the first portion of the initial segment of the audio data is output while avoiding outputting the second portion of the initial segment, where the decoded audio data includes the first portion of the initial segment of the audio data.

[0072] After generating the updated decoded audio data, the system is configured to display the updated decoded audio data. In some instances, the acquired audio data is streaming audio data. In some instances, the acquired audio data is streaming audio data, where the updated decoded audio data is displayed as live captions of the streaming audio data. Additionally or alternatively, the acquired audio data is previously recorded data, where the updated decoded audio data is displayed as a complete transcription of the audio data.

[0073] Disfluency

[0074] It should be understood that there are many different types of disfluencies that can be corrected or selectively retained based on different attributes associated with the decoded audio data. For example, turning attention to Figures 6 to 11 , each of which shows different disfluency corrections that can be selectively made to improve the readability of the decoded audio data.

[0075] Now turning attention to Figure 6 , Figure 6 shows an example embodiment of identifying and correcting disfluencies associated with punctuation marks included in the decoded audio data. As Figure 6As shown, the decoded audio data 602 includes the sentence "Well, I think it will work. Without it." The system identifies a disfluency 604 associated with a punctuation mark at the split boundary between "work" and "without". The system determines that merging the two sentences will result in a higher readability score than the original sentences and applies a correction (e.g., applies correction 606). The updated decoded audio data 608 includes "Well, I think it will work without it." This reflects a correction to the punctuation mark (e.g., secondary punctuation mark) that caused the readability score of the original sentence to decrease. In this case, the interjection is retained, and a character space is used to merge the sentences, replacing the period in the initial punctuation set included in the decoded audio data 602.

[0076] Figure 7 An example embodiment is shown that identifies and corrects disfluencies associated with repeated words included in the decoded audio data. As Figure 7 shown, the decoded audio data 702 includes the sentences "Well, I think it will work. Without it." The system identifies a disfluency associated with the repeated word 704 (e.g., "I"). The system determines that removing one of the repeated words and merging the two sentences will result in a higher readability score than the original sentences and applies correction 706. The updated decoded audio data 708 includes "Well, I think it will work without it." This reflects a correction to the repeated word that caused the readability score of the original sentence to decrease. In summary, in some instances, when a disfluency is associated with a repeated word, applying a specific correction to the decoded audio data includes removing one or more of the repeated words.

[0077] Now turn attention to Figure 8 , Figure 8 An example embodiment is shown that identifies and corrects disfluencies associated with interjections included in the decoded audio data. As Figure 8 shown, the decoded audio data 602 includes the sentences "Well, I think it will work. Without it." The system identifies a disfluency associated with the interjections 804 (e.g., "Well" and "uh") at the start of the sentence. The system determines that removing the interjections and merging the two sentences will result in a higher readability score than the original sentence and applies a correction (e.g., applies correction 806). The updated decoded audio data 808 includes "I think it will work without it." This reflects a correction to the punctuation mark that caused the readability score of the original sentence to decrease. In this case, the interjections are removed, and a character space is used to merge the sentences, replacing the period in the initial punctuation set included in the decoded audio data 802.

[0078] Thus, as Figure 8 shown in the example of

[0079] Now turning attention to Figure 9 , Figure 9 which shows an example embodiment of identifying and correcting disfluencies associated with words having a low recognition score included in decoded audio data. As Figure 9 shown, the decoded audio data 902 includes the sentence "I am hanging out at the new stores in town."

[0080] The system determines the recognition score 904 for each tag included in the decoded audio data and determines that the tag for "hanging out" has a recognition score that fails to meet or is below a predetermined recognition score threshold. This low recognition score is used to identify the disfluency 906 associated with the word "hanging out."

[0081] The system determines that replacing "hanging out" with a word having a higher recognition score or a higher context score relative to other tags in the sentence will result in a higher readability score than the original sentence and applies the correction (e.g., applies correction 908). The updated decoded audio data 910 includes "I am thinking about the new stores in town," and the audio data 910 reflects a correction to the punctuation that caused the readability score of the original sentence to decrease. In this case, the word "thinking" is used to replace the word "hanging out" to improve the readability score.

[0082] Now turning attention to Figure 10 , Figure 10 which shows an example embodiment of identifying and correcting disfluencies associated with a mismatch in reading level between decoded audio data and a target reader. For example, as Figure 10 shown, the system identifies a first reading comprehension level 1004 of a target user 1002 to whom the updated decoded audio data will be presented. The system also identifies a second reading comprehension level 1008 associated with a particular word (e.g., diligent) included in the decoded audio data 1006.

[0083] The system compares reading levels and determines that a second reading comprehension level is higher than a first reading comprehension level, where disfluency 1010 is flagged based on determining that the second reading comprehension level is higher than the first reading comprehension level (i.e., a reading level mismatch). Thus, applying a specific correction 1012 to the decoded audio data includes replacing a specific word with a new word (e.g., effort) having a reading comprehension level equal to or less than the first reading comprehension level, which improves the readability of the updated decoded audio data 1014 for the target user while maintaining a similar semantic meaning of the sentence (i.e., conveying a similar message to the target reader).

[0084] Attention is now turned to Figure 11 , Figure 11 which shows an example embodiment of identifying and correcting disfluencies associated with confidential information included in decoded audio data. As Figure 11 shown, the decoded audio data 1102 includes the sentence "Their revenue exceeded $3 million." The system flags a disfluency associated with confidential information 1104 related to the disclosure of how much revenue the company made.

[0085] In some instances, a user may wish to share the decoded audio data outside the company but not disclose the confidential information / data. The system applies a correction 1106 such that the updated decoded audio data 1108 includes "Their revenue exceeded X." This reflects a correction of the confidential data (e.g., rewriting a previously disclosed amount) or replacing a specific word with a new word not associated with the confidential information.

[0086] User display settings

[0087] Attention is now turned to Figure 12 , Figure 12 which shows an example embodiment of determining when and how to display the updated decoded audio data. For example, the updated decoded audio data 1204A includes the statement: "I will feed him after walking the dog." The system analyzes the sentence and determines that there is no disfluency 1208, and the updated decoded audio data is displayed in a regular font 1216 on the user display 1214. The updated decoded audio data 1204B includes the sentence: "I will or may for I will feed him after walking the dog." The system analyzes the sentence and determines that a disfluency 1210 is identified but retained in the updated decoded audio data. The updated decoded audio data is displayed in an italic font 1218 on the user display 1214 to indicate the presence and retention of the identified disfluency. This facilitates bringing to the attention of the reader or sender a possible disfluency that may or may not require correction.

[0088] The updated decoded audio data 1204C includes the sentence "I will feed him after walking the dog." This reflects the correction of the interjection disfluency (e.g., "or maybe to") identified in the updated decoded audio data 1204B. Thus, the system identifies the corrected disfluency 1212 and displays the updated decoded audio data on the user display 1214 by bolding a portion of the sentence 1220 associated with the removed interjection. In this way, a user reading the sentence on the user display can know whether a disfluency has been identified and the location where the disfluency has been corrected based on the different formatting applied. This can be helpful, especially for helping to provide visibility to the user of the changes that have been made.

[0089] In view of the foregoing, it should be understood that various different techniques can be applied to identify corrected disfluencies in the updated decoded audio data and to apply a first visual format corresponding to the corrected disfluency that is different from the format applied to the rest of the decoded audio data. Specifically, as described and shown, when the updated decoded audio data is displayed on the user display, the user display reflects a first visual formatting and a second visual formatting to indicate to the user which part of the decoded audio data has been corrected. The system then displays the updated decoded audio data set on the user display, including the first visual formatting and the second visual formatting.

[0090] Attention is now turned to Figure 13 , Figure 13 An example embodiment of a related method for displaying updated decoded audio data is shown, including displaying an embedded link that directs the user to the original version of the decoded audio data (e.g., the data that existed before the disfluency was corrected). For example, the decoded audio data 1302 includes the sentence: "It is being used for for um the menu in the restaurant." The updated decoded audio data 1304 reflects the correction of disfluencies (e.g., the repeated word "for for" and the interjection "um"). The system identifies the corrected disfluencies and embeds a link 1306 associated with the corrected disfluencies within the updated decoded audio data, where the link directs the user to the unedited version of the decoded audio data. The system can then display the updated decoded audio data at the user display, including displaying the link as a selectable object.

[0091] In this example, user display 1308 shows updated decoded audio data 1310 with an embedded link corresponding to the remaining word "for". This link, when selected, is operable to cause user display 1308 to display the original decoded audio data 1312 (e.g., unedited). Additionally or alternatively, when the user clicks on the selectable link, the user display is configured to display only the modified portion 1314 (e.g., "for for um") of the original sentence, as shown.

[0092] Figure 14 Example embodiments illustrate different ways of displaying updated decoded audio data based on one or more attributes of the updated decoded audio data. In some embodiments, when displaying updated decoded audio data based on one or more attributes, the system may determine to retain or remove disfluencies. For example, decoded audio data 1402 includes the sentence: "I will or might feed him after I walk the dog." Updated decoded audio data 1404 includes the following statement: "I will feed him after I walk the dog." In some instances, the system selectively displays the original decoded audio data (i.e., retains the disfluency and thus overrides the identification of the disfluency) or removes the disfluency based on an identified attribute 1406 associated with the decoded audio data.

[0093] By way of example, if the auditory attribute of the speaker identification emphasizes the words "or might", e.g., with increased volume, the system may determine that the disfluency is desired and relevant to conveying the meaning of the sentence, and thus the system will display the original decoded audio data 1412 at user display 1408. In another example, if a linguistic attribute associated with the words "or might" is identified because those interjections do not add to or are unnecessary for conveying the overall meaning of the sentence, the system is configured to display the updated decoded audio data 1410 on the user display.

[0094] It should be understood that the system can deterministically select which data to display based on attributes of the context associated with the disfluency and / or attributes of the target reader (i.e., the end user of the user display).

[0095] It should also be understood that many different visual formatting modifications can be made to the decoded streaming audio segment before or during its display on the user display. Visual formatting can include font formatting (e.g., bold italic, underline, and / or strike-through), one or more fonts, different capitalization schemes, text coloring and / or highlighting, animation, or any combination thereof. While the following figures describe visual formatting modifications corresponding to different fonts, it should be understood that attribute analysis or external content links can be displayed according to any of the above visual formatting types.

[0096] Now turn attention to Figure 15 , Figure 15 which shows an example embodiment of a user display having at least two different modes, one for accessibility and one for readability. As Figure 15 shown, the computing system obtains decoded audio data 1502 that includes a transcription of spoken utterances identified in the audio data (e.g., "uh, okay I think it'll work. Without it.").

[0097] The system then generates different updated decoded audio data based on an optimization of the readability or accessibility of the user display. For example, optimizing for accessibility involves preserving disfluencies in the final output. Thus, the system identifies disfluency 1504 in the decoded audio data of the spoken utterance associated with the initial punctuation included in the decoded audio data 1502.

[0098] In this example, the system determines that correcting the disfluency will improve the readability of the transcription while still maintaining accessibility to the exact words spoken. After generating a specific correction to correct the disfluency, the system applies the specific correction 1506 to the decoded audio data in order to generate updated decoded audio data 1508 that includes the spoken utterance reflecting the specific correction (e.g., "uh, okay, I think it'll work without it."). The updated decoded audio data has improved readability over the decoded audio data of the spoken utterance but maintains high accessibility.

[0099] In addition, the system generates updated decoded data optimized for readability. Thus, the system identifies disfluency 1510 in the decoded audio data of the spoken utterance associated with an interjection included in the decoded audio data 1502. The system determines that correcting the disfluency will improve the readability of the transcription. After generating a specific correction to correct the disfluency, the system applies the specific correction 1512 to the decoded audio data in order to generate updated decoded audio data 1514 (e.g., "I think it'll work without it."), which includes the spoken utterance reflecting the specific correction that removes the interjection as they do not contribute to the overall meaning of the sentence. The updated decoded audio data has improved readability over the decoded audio data of the spoken utterance.

[0100] As Figure 15As shown, the user display 1516 is configured with multiple modes (e.g., an accessibility mode 1518, which is configured to display updated decoded audio data 1520 optimized for accessibility; and a readability mode 1522, which is configured to display updated decoded audio data 1524 optimized for readability). The system or the user can switch between any mode based on user input or an automatically identified attribute associated with the decoded audio data 1502.

[0101] For example, in some instances, when it is determined that disfluencies in the updated decoded audio data should be removed, the system selects the readability mode to display the updated decoded audio data. Alternatively, once it is determined that the disfluencies should be retained, the accessibility mode is selected to display the decoded audio data.

[0102] In addition, in some instances, the system receives user input indicating whether to use the readability mode or the accessibility mode, and selects which mode to use based on the user input. Additionally or alternatively, the computing system automatically identifies an attribute associated with the target user or an attribute associated with the target output (e.g., live captions or discrete transcriptions), and selects which mode to use based on the identified attribute.

[0103] In view of the foregoing, it should be understood that embodiments of the present disclosure provide improved technical benefits over conventional automatic speech recognition systems that are not well trained to punctuate and display decoded audio data with speech disfluencies. Technical advantages include the ability to improve segmentation and punctuation based on the represented disfluencies. Additionally, the readability of the final output of the ASR system is improved by providing better punctuation (if the disfluencies are retained) or removing the disfluencies to better match the standard written communication style. This further improves additional downstream tasks, such as summarization, question and answer queries, or other natural language processing tasks.

[0104] Additionally, training is improved by training a student model on training data using LS-PTM, where the training data illustrates different ways of handling disfluencies in the decoded audio data. Since the system considers several different options for correcting disfluencies, the user display can be optimized for different criteria, including readability and / or accessibility.

[0105] Example Computing System

[0106] Embodiments of the present invention may include or utilize a special purpose or general purpose computer including computer hardware (e.g., computing system 210), as discussed in more detail below. Embodiments within the scope of the present invention also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media accessible by a general purpose or special purpose computer system.

[0107] A computer-readable medium (e.g., Figure 2 hardware storage device 240) storing computer-executable instructions (e.g., Figure 2 computer-executable instructions 218) is a physical hardware storage medium / device that excludes a transmission medium. A computer-readable medium that carries computer-executable instructions or computer-readable instructions (e.g., computer-executable instructions 218) in one or more carrier waves or signals is a transmission medium. Thus, by way of example and not limitation, embodiments of the present invention can include at least two distinctly different types of computer-readable media: physical computer-readable storage media / devices and transmission computer-readable media.

[0108] Physical computer-readable storage media / devices are hardware and include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage (such as CDs, DVDs, etc.), magnetic disk storage, or other magnetic storage devices, or any other hardware that can be used to store the desired program code means in the form of computer-executable instructions or data structures and that is accessible by a general purpose or special purpose computer.

[0109] A "network" (e.g., Figure 2 network 230) is defined as one or more data links capable of transmitting electronic data between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computer via a network or another communication connection (wired, wireless, or a combination of wired or wireless), the computer properly views the connection as a transmission medium. Transmission media can include networks and / or data links that can be used to carry the desired program code means in the form of computer-executable instructions or data structures and that are accessible by a general purpose or special purpose computer. Combinations of the above are also included within the scope of computer-readable media.

[0110] In addition, once reaching various computer system components, program code devices in the form of computer-executable instructions or data structures can be automatically transferred from a transmission computer-readable medium to a physical computer-readable storage medium (and vice versa). For example, computer-executable instructions or data structures received via a network or data link can be cached in RAM within a network interface module (e.g., "NIC") and then ultimately transferred to the computer system RAM and / or a less volatile computer-readable physical storage medium at the computer system. Thus, a computer-readable physical storage medium can be included in computer system components that also (or even primarily) utilize a transmission medium.

[0111] Computer-executable instructions include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a particular function or group of functions. Computer-executable instructions can be, for example, binary code, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims need not be limited to the above-described features or acts. Instead, the described features and acts are disclosed as example forms for implementing the claims.

[0112] Those skilled in the art will understand that the present invention can be practiced in a network computing environment having many types of computer system configurations, including personal computers, desktop computers, laptop computers, messaging processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, pagers, routers, switches, and the like. The present invention can also be practiced in a distributed system environment where local and remote computer systems that are network-linked (either via hardwired data links, wireless data links, or a combination of hardwired and wireless data links) both perform tasks. In a distributed system environment, program modules can be located in local and remote memory storage devices.

[0113] Alternatively or additionally, the functions described herein can be performed, at least in part, by one or more hardware logic components. By way of example, and not limitation, illustrative types of hardware logic components that can be used include field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), and the like.

[0114] The present invention can be implemented in other specific forms without departing from its basic characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. Thus, the scope of the present invention is indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalents of the claims are embraced within their scope.

Claims

1. A method implemented by a computing system for modifying decoded audio data based on a spoken utterance that appears in audio data for generating the decoded audio data, for improving the readability of a transcription of the decoded audio data, the method comprises: obtaining decoded audio data including a spoken utterance recognized in the audio data; identifying disfluencies in the decoded audio data; determining that correcting the disfluencies will improve the readability score of the decoded audio data; generating a specific correction for correcting the disfluencies; applying the specific correction to the decoded audio data; and generating updated decoded audio data reflecting the specific correction, the updated decoded audio data having improved readability for the spoken utterance over the decoded audio data.

2. The method according to claim 1, wherein obtaining the decoded audio data further comprises: obtaining audio data including a spoken utterance from a speaker; continuously decoding the audio data to generate decoded audio data; determining whether a language boundary exists within an initial segment of the decoded audio data; when a language boundary is determined to exist, identifying a first portion of the initial segment that is temporally before the language boundary, and identifying a second portion of the initial segment that is temporally after the language boundary; and outputting the first portion of the initial segment of the audio data while avoiding outputting the second portion of the initial segment, wherein the decoded audio data includes the first portion of the initial segment of the audio data.

3. The method according to claim 1, after applying the specific correction and generating the updated decoded audio data, displaying the updated decoded audio data at a user display.

4. The method according to claim 3, wherein the audio data is streaming audio data, and wherein the updated decoded audio data is displayed as live captions for the streaming audio data.

5. The method according to claim 1, wherein the decoded audio data includes initial punctuation, the method further comprises: generating secondary punctuation configured to facilitate correction of disfluencies associated with the initial punctuation, wherein applying the specific correction to the decoded audio data comprises: replacing the primary punctuation with the secondary punctuation.

6. The method according to claim 1, wherein the disfluency is associated with a repeated word, and wherein applying the specific correction to the decoded audio data comprises: removing the repeated word.

7. The method according to claim 1, wherein the disfluency is associated with an interjection, wherein applying the specific correction to the decoded audio data comprises: removing the interjection.

8. The method according to claim 1, further comprises: identifying a first reading comprehension level of a target user; Identify a second reading comprehension level associated with a specific word included in the decoded audio data; Determine that the second reading comprehension level is higher than the first reading comprehension level, wherein the disfluency is identified based on determining that the second reading comprehension level is higher than the first comprehension level; And Wherein applying the specific correction to the decoded audio data includes: replacing the specific word with a new word having a reading comprehension level equal to or less than the first reading comprehension level.

9. The method according to claim 1, further comprising: Determine an identification score associated with a specific word included in the decoded audio data, the identification score being associated with the accuracy of the computing system decoding a portion of the audio data corresponding to predicting the specific word; And Determine that the identification score is below an identification score threshold, wherein the disfluency is identified based on determining that the identification score is below the identification score threshold; Wherein applying the specific correction to the decoded audio data includes: replacing the specific word with a new word having a higher identification score than the specific word.

10. The method according to claim 1, further comprising: Identify a specific word included in the decoded audio data, the specific word being associated with confidential information; Wherein applying the specific correction to the decoded audio data includes: redacting the specific word or replacing the specific word with a new word not associated with confidential information.

11. The method according to claim 1, further comprising: Identify a corrected disfluency in the updated decoded audio data; Apply a first visual formatting corresponding to the corrected disfluency, the corrected disfluency being different from a second visual formatting corresponding to the decoded audio data, such that when the updated decoded audio data is displayed at a user display, the user display reflects the first visual formatting and the second visual formatting to indicate to the user that a portion of the decoded audio data has been corrected; And Display the updated decoded audio data at the user display, the updated decoded audio data including the first visual formatting and the second visual formatting.

12. The method according to claim 1, further comprising: Identify a corrected disfluency within the updated decoded audio data; Embed a link associated with the corrected disfluency within the updated decoded audio data, wherein the link directs a user to an unedited version of the decoded audio data; And Display the updated decoded audio data at a user display, the updated decoded audio data including displaying the link as a selectable object.

13. The method according to claim 1, further comprising: Identify an attribute associated with the disfluency; And Based on the attribute, determine to retain the disfluency in the updated decoded audio data.

14. The method according to claim 1, wherein the improved readability is based on an improved readability score.

15. The method according to claim 1, wherein the specific correction is generated by a student model, and the student model is trained using training data generated by the LS-PTM.