Method and system for real-time transcription of an audio data stream into text and for time-synchronous playback of the audio data stream and the text.

The method and system address the challenge of synchronizing text with audio in internet radio by real-time transcription and display, using a server to process complex AI models and ensure millisecond-precise synchronization and efficient resource use.

DE102024107191B4Active Publication Date: 2025-12-24RADIOZEIT GMBH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
DE102024107191
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-02-20
Filing Date
2024-03-13
Publication Date
2025-12-24
Estimated Expiration
2044-03-13

AI Technical Summary

Technical Problem

Existing internet radio stations with high spoken content face challenges in displaying synchronized text due to low playback volume, background noise, and hearing impairments, and existing solutions like downloadable video files are impractical due to high costs and lack of real-time capability.

Method used

A method and system that transcribe spoken words from internet radio streams into text in real-time with a delay of up to one minute, using a server to process complex AI models and synchronize the text display with audio playback on end devices, employing multiple AI models for transcription, correction, and translation.

Benefits of technology

Enables accurate, millisecond-precise synchronization of audio and text playback, reducing hallucinations and resource usage, and supporting translation and music visualization, while maintaining real-time processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000012_0000
    Figure 00000012_0000
  • Figure 00000013_0000
    Figure 00000013_0000
  • Figure 00000014_0000
    Figure 00000014_0000
Patent Text Reader

Abstract

Method for real-time transcription of a continuous audio data stream (1) from an internet radio station into text, by a server (2), and for playing back the audio data stream and time-synchronously displaying the transcribed text on an end device (3, 3-1, 3-2, 3-3), wherein the server (2) performs the following steps: Retrieving a continuous audio data stream (1) from the internet radio station; Extracting or generating (15) time-sequential audio bit data sets (ABD), for example as WAV files, from the continuous audio data stream (1); Determining (21) split time points (Z) within the audio bit data sets (ABD), wherein each split time point (Z) specifies a time of a pause between sentences or words, using an initial AI model to detect sentence pauses and speech pauses between words; Assembling (23) successive audio bit data sets (ABD) into a respective audio data chunk (AC), the start time and end time of which are determined by a respective split time (AZ) such that the start time of a currently assembled audio data chunk (AC) is prior to an end time of an earlier assembled and temporally adjacent audio data chunk (AC), and the temporally adjacent audio data chunks (AC) overlap with a predetermined minimum buffer length (PL), in particular a buffer length of at least one audio bit data set (ABD); Transcribing (24) a respective audio data chunk (AC) into an associated transcript chunk (TC) using a second, in particular generative, AI model for transcription, wherein the second AI model is more complex than the first AI model, wherein a respective transcript chunk (TC) specifies a language contained in the audio data chunk (AC) as text annotated with time information, wherein the text is represented as segments and The token is specified, and the time information specifies the start and end time of each segment or token in the audio data chunk; Filtering out overlapping duplicate text from a given transcript chunk (TC) to obtain a non-overlapping transcript chunk (TC); Transmitting each non-overlapping transcript chunk (TC) as the current part of a time-annotated text data stream to the terminal (3, 3-1, 3-2, 3-3); and Transmitting the continuous audio data stream as temporally sequential audio data segments (ADS1, ADS2), wherein the audio data segments contain information about their occurrence times from the start of the audio data stream, where the terminal (3, 3-1, 3-2, 3-3) performs the following steps: Receiving the temporally sequential audio data segments (ADS1, ADS2) as an audio data stream from the server (2); Receiving the transcript chunks (TC) as a text data stream from the server (2); Playing back the audio data stream and time-synchronized display of a text from the text data stream, wherein the relative start and end times of segments and tokens stored in the transcript chunks (TC) are converted into absolute start and end times with reference to the start of the audio data stream, and the terminal device (3, 3-1, 3-2, 3-3) displays a respective transcript chunk segment at the occurrence time of an audio data segment and preferably highlights a respective token in the displayed segment if its absolute start time is before and its absolute end time is after the occurrence time.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to a method and system for real-time transcription of a continuous audio data stream from an internet radio station into text by a server, and for playing back the audio data stream and time-synchronously displaying the transcribed text on an end device. State of the art

[0002] Nowadays, there is a huge variety of internet radio stations whose audio streaming content can be accessed and received via the internet using suitable devices such as smartphones, tablets, smart TVs, TV sticks, smartwatches, smart speakers, and internet-enabled digital radios. However, internet radio stations with a high proportion of spoken content and little music often present the problem that the spoken words are difficult to understand. This can be due to a necessary low playback volume, background noise near the device, or hearing impairment on the part of the user. Ideally, in such cases, the audio streaming content of stations with a high proportion of spoken content would also be displayed as text on the device's screen. For such a solution to be widely accepted and truly user-friendly, it is crucial that the text is displayed synchronously with the spoken words.

[0003] One approach for such a system with synchronized speech output and text display could involve providing the streaming content as downloadable or streamable video files. However, this approach is impractical due to the high costs for the operator and the large amounts of data that would need to be transferred. A significant drawback of this approach would also be the system's lack of real-time capability, as the text could only be integrated into the video after the currently played speech or speaker's voice has finished.

[0004] Therefore, there is a need for a method and / or a system that automatically converts the speech contained in streaming content into text during the continuous retrieval of the streaming content, i.e., in real time, but with the allowance of a processing-related continuous delay of approximately one minute, and displays the text synchronously together with the speech playback.

[0005] From US 2020 / 0126538 A1, a speech recognition system is known in which audio data for long expressions is divided into several overlapping segments.

[0006] From US 2021 / 0295846 A1, a speech / text processing system is known that uses a dynamic speech model selection to use speech models that correspond to a likely pronunciation of a spoken word.

[0007] From US 2023 / 0096543 A1, a system and a method for providing automated language translation in real time are known. Disclosure of the invention

[0008] The object of the invention is to provide a method and a system which make it possible to automatically and in real time create a corresponding text from a language or spoken words contained in an audio streaming content, and to display the created text synchronously with the playback of the audio streaming content on an end device.

[0009] In this application, the term "language" is generally understood to mean the linguistic content or spoken words and sentences of an audio data stream. Only in connection with a translation of the transcribed text into a foreign language or another national language does the term "language" in this application have a country-specific reference, as in "mother tongue," "foreign language," "original language," and "national language."

[0010] To implement the method or system, it is necessary to delay the original signal from the internet radio station by a sufficiently low minimum amount. For example, a total delay of 50 seconds is planned, with, for instance, 30 seconds allocated for recording the audio signal and another 20 seconds for overlapping, transcribing, and, if necessary, translating the audio signal. This constant delay, as explained above, does not contradict the feature of real-time transcription. The term "real-time" here means that the presented method does not cause the intended maximum total delay to be exceeded.

[0011] The object of the invention is achieved by a method having the features of claim 1. Furthermore, the object of the invention is achieved by a system having the features of claim 10. Advantageous embodiments are described in the additional features of the dependent claims.

[0012] A method is proposed for the real-time transcription of a continuous audio data stream from an internet radio station into text by a server, and for playing back the audio data stream and displaying the transcribed text in a time-synchronous manner on an end device.

[0013] The audio stream is played back on the end device, and the transcribed text is displayed synchronously. The actual processing and transcription take place on the server, which is preferably an internet server. This allows multiple end devices with limited processing power to use a central server, which, for example, can process a very complex or even the most complex AI model for transcription using GPUs.

[0014] The server retrieves a continuous audio stream from the internet radio station and extracts or generates time-sequential audio bit data from this stream. These audio bit data can be, for example, WAV files (WAV = Waveform Audio File Format). Each audio bit data record might represent, for example, a 5-second audio signal interval.

[0015] To avoid cut start and end times of the audio data stream in the middle of a word, it is proposed to first identify ideal cut points, for example after a sentence or in a pause between two words, and then provide these cut points to a process for joining or assembling the short audio bit data sets into an audio data chunk.

[0016] The server therefore determines split points or intersections within the audio bit data sets, with each split point specifying a pause between sentences or words, using an initial AI model to detect sentence pauses and pauses between words. This first AI model is smaller and therefore less complex and computationally intensive than a second AI model mentioned later. The first AI model is selected, adapted, and trained to focus primarily on timing rather than the accuracy of the transcript.

[0017] The server assembles successive audio bit records into a single audio data chunk. The start and end times of each chunk are determined by a specific split time point, such that the start time of a currently assembled audio data chunk precedes the end time of a previously assembled and temporally adjacent audio data chunk. The temporally adjacent audio data chunks overlap by a predetermined minimum buffer length. This predetermined minimum buffer length is preferably at least one audio bit record, for example, 5 seconds.

[0018] The server transcribes each audio data chunk into an associated transcript chunk using a second AI model for transcription. This second AI model is preferably a generative AI model and is more complex than the first. For example, the second AI model is Whisper™ from OpenAI™. Each transcript chunk specifies the language (in the sense of sentences and words) contained within the audio data chunk as text annotated with time information. The text is specified as segments and tokens. A segment represents, for example, a sentence, and a token represents a word or part of a word. The time information specifies the start and end times of each segment or token within the audio data chunk. The timestamps, i.e., the annotated start and end times, refer to the relative position within the chunk. Time 0 indicates the starting point without any overlap.Due to overlap, the first timestamps can therefore become negative (for example, a maximum of -10s).

[0019] The server then filters out overlapping duplicate text from each transcript chunk to obtain a transcript chunk free of overlaps. This filtering process is performed incrementally, first searching the previous chunk for the first five contiguous tokens of a new chunk that end at a time greater than 0 seconds. If multiple contiguous tokens are not found, the number of contiguous tokens is reduced by one until one is found. If even a single token is not found in the previous chunk, no filtering takes place, and the entire chunk is retained.

[0020] The server then transfers each non-overlapping transcript chunk as the current part of a time-annotated text data stream to the end device.

[0021] In addition, the server performs a transmission of the continuous audio data stream as temporally sequential audio data segments, for example according to the DASH or HLS streaming protocol, whereby the audio data segments contain information about their occurrence times from the start of the audio data stream.

[0022] The end device receives the sequential audio data segments as an audio data stream from the server. The end device also receives the transcript chunks as a text data stream from the server.

[0023] The terminal plays back the audio stream and displays a time-synchronized text from the stream, converting the relative start and end times of segments and tokens stored in the transcript chunks into absolute start and end times relative to the start of the audio stream. At the occurrence time of an audio segment, the terminal displays a corresponding transcript chunk segment and preferentially highlights a token within the displayed segment if its absolute start time (relative to the segment and token) is before and its absolute end time is after the occurrence time of the audio segment.

[0024] The playback of the audio stream and the time-synchronized display of text from the text stream, as described above, require synchronization of the audio and text stream data. The implementation of this synchronization is illustrated below. The end device receives the audio stream using a streaming protocol, such as a DASH stream containing an OPUS audio codec or an HLS stream containing an AAC audio codec. These streaming protocols provide millisecond-accurate information about when the stream was started on the server side, as they themselves deliver the audio stream in individual audio segments via the HTTP protocol.

[0025] Based on the millisecond-precise position of the audio data stream or audio data segments, the end device can retrieve the relevant part of the transcript, for example, a section within a 30-second transcript chunk. A sequential number for the chunk can be calculated according to: Chunk_Number=[Playback time / 30s]

[0026] The transcript chunks are transmitted, for example, in JSON format or as a subtitle data track in WebVTT (Web Video Text Tracks) or TTML (Timed Text Markup Language) format as part of a DASH or HLS stream. Each chunk contains a list of segments (e.g., sentences), each of which in turn contains a list of tokens (e.g., (partial) words). Every segment and every token contains a start and end time in milliseconds. The end device reads these segments with their respective tokens and stores them in a mapping table. The relative start and stop times within the transcript chunks are then converted to absolute times from the start of the audio stream on the server side, according to: t_a=Chunk_Number*30+t_r where t_r is the relative start time or end time within the transcript chunk, and t_a is the absolute start time or end time since the stream started on the server.

[0027] This mapping table allows the terminal device to display the correct segment by retrieving the segment at the current time in milliseconds, and in the same way also to highlight the current (partial) word or token.

[0028] The presented method enables the playback of the audio data stream with a time-synchronized display of the transcribed text on the terminal device, accurate to the millisecond range.

[0029] When transcribing with current generative AI models, such as Whisper™ from OpenAI™, so-called hallucinations still occur. For example, after completed sentences, copyright notices or general information about subtitling are displayed during a longer audio silence. The reason for this is that such notices were apparently present in the training data used by the AI ​​models, and the model has learned to generate them after the transcription is complete. Other examples include closing phrases like "See you next time," which were not present in the audio stream but appear in the transcript.

[0030] In a further development of the presented method, the server uses a hallucination database to detect an erroneously generated segment within a transcript chunk. This segment, created during transcription by a hallucination of the second AI model without corresponding audio data, is then removed along with its timing information. Filtering is performed using an expandable filter table. Each entry in the table results in the removal of every occurrence of that entry found in the transcribed text.

[0031] In a further development of the presented method, the server uses a third AI model to correct the content of transcript chunks. A correction of a content error in one or more transcript chunks is performed by sending a query to the third AI model. This third AI model is preferably a large-language model, such as GPT-4® from OpenAI™. The percentage of corrected words is limited to a predefined maximum value, preferably 5%. If the predefined maximum value is exceeded and / or a predefined correction time is exceeded, the correction process is aborted.

[0032] In a further development of the presented method, the server uses an additional AI model to determine the probability of each occurrence of speech, music, and / or noise in the audio data stream and / or individual audio data chunks at a predefined sampling rate, for example, once per second. The server only transcribes the audio data chunks into transcript chunks if the determined speech-related probability value is above a predefined threshold and the music-related probability value is below a predefined threshold. This saves resources, thereby reducing power consumption and costs. It also reduces the occurrence of potential hallucinations. To achieve this, threshold exceedances and / or falls below the thresholds trigger further splitting points.Intersection points are sent to the transcription process via interprocess communication. Sections that do not need to be transcribed are then cut from the audio data stream or the audio bit data sets. These now missing sections in the transcript are reinserted by incrementing all subsequent timestamps by the duration of the missing section.

[0033] The server preferentially informs the end device of the determined probability values ​​or threshold exceedances / falling below the limit for the purpose of adjusting the display mode. When music is playing, the end device can then display the spectrum analyzer or frequency spectra described below, along with the album cover of the respective song, instead of the transcript.

[0034] In a further development of the presented method, the server determines a Fast Fourier Transform (FFT) dataset for each audio data chunk and makes it available to the end device for retrieval. Based on the retrieved FFT datasets, the end device visualizes the frequency spectra of the audio data chunks, particularly for pure music audio chunks. This avoids the need for complex processing of the raw audio signal data on the end device, provided such processing is even possible there. For example, an FFT dataset is provided on the server for each audio data chunk in a separate JSON file for retrieval by the end device. According to this further development, an FFT is performed on the server, and the resulting data, such as 16 frequency bands, is sampled at a rate of, for example, 50 Hz.The data volume, for example, corresponds to a maximum of 80 kBytes for 30 seconds, which is one fifth of the audio data to be transmitted.

[0035] A further development of the presented method focuses on server-side, time-synchronized translation of the transcript into a foreign language. To display the translated transcript synchronously with the audio stream, it is translated on the server side, and the timestamps of the original language transcript are also transferred. Since the translated texts may be longer, shorter, or in a different order than the original, the timestamps must be adjusted. For this purpose, the transcript chunk is translated in the first step.

[0036] In this further development of the presented method, it is accordingly provided that the server continuously translates transcript chunks into a specified foreign language using an AI model and creates foreign language transcript chunks with adapted time information and transmits them to the end device, and the end device synchronizes the foreign language transcript chunks, based on the adapted time information, with the playback of the audio data stream.

[0037] In a further training module of the aforementioned program, the server is designed to translate transcript chunks using a large-language model (LLM), such as a GPT® version like GPT-3.5 or GPT-4® from OpenAI™, and a translation AI model, such as one from DeepL® SE. The server prioritizes the first translation result generated by the LLM over the second translation result generated by the translation AI model if the first translation result is available to the server within a predefined timeframe, for example, within 7 seconds. If the predefined timeframe is exceeded, only the second translation result, which is available within 1 to 2 seconds, is used.

[0038] Once the translated transcript is available, the text can be divided into segments and tokens. First, the individual words are evenly distributed across segments, ensuring the same number of segments in the translation. The start and end times of the translated segments are then adjusted so that the segments, now containing roughly the same number of words, are evenly spaced within the chunk's time frame. Finally, within each segment, the individual letters of the words are evenly distributed as tokens. This ensures that each letter appears approximately the same length in the frontend. Although the translation doesn't always display the exact word being spoken, it still allows for easy reading, and the display points in the original language and translation are as synchronized as possible.

[0039] Further training and development of the presented procedure is aimed at the synchronous playback of the translation with a synthetically generated voice.

[0040] The system envisions that the server, using an additional AI model, will generate an extra audio stream with a synthetic voice from the transcript chunks translated into a foreign language and transmit it to the end device. Alongside the first playback process for the original audio stream retrieved from the internet radio station, the end device will initiate a second playback process for the additional audio stream. The first playback process controls the playback position of the second. In the event of a pause in playback, or if the two playback processes are no longer synchronized to within a millisecond plus / minus of an allowed variance, the second playback process will adjust to the playback position of the first.

[0041] The implementation of a so-called "ducking" process is based on information from the AI ​​model for music recognition. Here, information about the presence of speech is transmitted to the end device at regular intervals, analogous to the transmission of FFT information. If it is detected that no relevant speech is present, the translated audio track is set to a volume of 0%, and the original audio track is set to 100% of the desired volume. As soon as translated speech becomes relevant, the volume of the original audio track is reduced to a lower value, for example, 25% of the desired volume. The volume of the translated audio track is then increased to 100%. This crossfade occurs dynamically within a few seconds, for example, 2 seconds.Similarly, the crossfade back to sections without speech occurs by reducing the volume level of the translated audio track to 0%, while simultaneously raising the level of the original audio track back to 100% volume.

[0042] A further development of the presented method focuses on coordinating parallel transcription processes to improve server hardware utilization. The transcription is intended to take place on dedicated expansion cards within the server. For example, two Nvidia GPUs per server are mentioned. These GPUs accelerate the processing of the generative AI by approximately a factor of 10. Since multiple internet radio stations are to be available on the end device, it is important to process as many data streams as possible in parallel. Starting the transcription process requires additional time. Therefore, it is not optimal to start and stop a process for each data chunk. Consequently, a method is proposed in which already initialized processes can process newly arriving data.For this purpose, ongoing processes are informed via interprocess communication when new audio bit data sets, such as WAV files, are available. These are then read in and processed, taking into account the split points or intersections. In a specific example case, the AI ​​model used, for example Whisper™ from OpenAI™, requires at least 2 GB of RAM on the aforementioned expansion card, which limits the number of parallel AI processes to, for example, 6 per GPU. Therefore, audio bit data sets from multiple senders are processed per transaction process.

[0043] According to this further development of the presented method, the server is designed to receive and transcribe audio bit data sets from several continuous audio data streams or transmitters in parallel, and to operate parallel transcription processes for transcribing audio data chunks using a respective second AI model, wherein the server further assigns audio data chunks from several audio data streams to a respective running transcription process for transcription, and wherein the server operates separate transcription processes for different national languages ​​or foreign languages.

[0044] Furthermore, a system for the real-time transcription of a continuous audio stream from an internet radio station into text, by a system server, and for playing back the audio stream and displaying the transcribed text synchronously on a system terminal is presented. The server and the terminal constitute the system and are configured to execute the above-described procedure and its subsequent development and training.

[0045] Furthermore, a computer program or computer program product with executable instructions is presented which, when executed by a server, cause the server to execute the above-presented procedure and its further training and development, and, when executed by an end device, cause the end device to execute the above-presented procedure and its further training and development.

[0046] Finally, a data carrier containing the computer program presented above is introduced, wherein the data carrier is a computer-readable storage medium, a server accessible via the Internet, an electronic signal, an optical signal and / or a radio signal.

[0047] The features of the method and system presented above can be combined arbitrarily within the scope of the invention, provided that such a combination is not obviously contradictory. Features of the claims to which the words "in particular" and "preferred" are assigned are to be understood as optional and non-limiting and serve to specify particularly advantageous embodiments of the present invention. Brief description of the drawings

[0048] To illustrate the presented method and system, exemplary implementations are now presented with reference to the following figures. Fig. Figure 1 schematically illustrates a system for real-time transcription of an audio data stream from an internet radio station by a server into text and for playing back the audio data stream and time-synchronously displaying the transcribed text on an end device according to an embodiment of the invention. Fig. Figure 2 schematically illustrates the processing processes of a transcription unit of the server according to a further embodiment of the invention; Fig. Figure 3 shows an exemplary compilation of audio bit data sets into audio data chunks according to an embodiment of the present invention. Fig. Figure 4 illustrates a coordination of parallel transcription processes according to an embodiment of the present invention.

[0049] In the figures shown, identical or similar elements, components, values ​​and intervals are designated across figures using the same reference symbols. Exemplary embodiments of the invention

[0050] Fig. Figure 1 schematically illustrates a system for real-time transcription of an audio data stream from an internet radio station into text by a server and for playing back the audio data stream and displaying the transcribed text synchronously on an end device according to an embodiment of the invention. The system comprises a server 2 and an end device 3. The server 2 is configured as an internet server for retrieving, receiving, and processing a continuous audio data stream 1 provided by an internet radio station. The server 2 is configured to extract or generate time-sequential audio bit data records ABD and time-sequential audio data segments ADS1, ADS2 from the retrieved audio data stream 1, for example, by using a Fig. 1 illustrated audio / video processing program 15, which is in Fig. 1 is specified as the open-source program ffmpeg.

[0051] The time-sequential audio bit data sets ADB are stored in WAV format or as WAV files (WAV = Waveform Audio File Format) with a predefined audio signal interval length of, for example, 5 seconds, which will be described in detail later with reference to Fig. The transcription unit 20 described in section 2 is provided. The transcription unit 20 of server 2 is configured to continuously determine time-annotated transcript chunks TC as the current part of a time-annotated text data stream based on the temporally sequential audio bit data records ABD. The transcript chunks TC represent the speech contained in the received audio data stream 1, i.e., spoken sentences and words or segments and tokens, and are transmitted from server 2 to terminal device 3 as the current part of the time-annotated text data stream.

[0052] The temporally sequential audio data segments ADS1 and ADS2 are provided for transmission to the terminal device 3 with a predefined time delay of, for example, 50 seconds, according to the DASH streaming protocol (based on an OPUS audio codec) or the HLS streaming protocol (based on an AAC audio codec). This time delay is predefined or selected such that, in addition to a recording time interval of, for example, 30 seconds for recording the audio signal contained in the audio data stream, a processing time interval of, for example, another 20 seconds is provided for processing the audio signal (with the length of the recording time interval) by the transcribing unit 20. The audio data segments ADS1 and ADS2 contain information about their occurrence times from the start of the audio data stream, since the streaming protocols used provide millisecond-accurate information about when the audio data stream or the audio data segments began.The stream was started on the server side.

[0053] As in Fig. As shown schematically in Figure 1, the server 2 can transmit the generated audio data segments ADS1, ADS2, and the transcript chunks TC determined by the transcribing unit to a variety of end devices 3. The end device 3 can be a universally applicable device, for example, an iOS® device 3-1, such as an iPad or iPhone, or an Android® device or desktop device 3-2, or a device 3-3 specifically designed for visualizing internet radio speech. The end device 3 is connected to the server 2 for wireless or wired data communication via the internet. The end device 3 is configured to receive the temporally sequential audio data segments ADS1, ADS2 as an audio data stream and the transcript chunks TC as a text data stream from the server 2. The end device 3 is further configured to play back the audio data stream and to display a time-synchronized text from the text data stream.The terminal 3 converts the relative start and end times of segments and tokens stored in the transcript chunks TC into absolute start and end times relative to the start of the audio stream. Furthermore, at the occurrence time of an audio data segment ADS1, ADS2, the terminal 3 displays a corresponding segment, such as a sentence, from a transcript chunk TC if the segment's absolute start time is before and its absolute end time is after the occurrence time. Additionally, the terminal 3 can visually highlight a specific token, such as a word or part of a word, within the currently displayed segment if the token's start time is before and its end time is after the current absolute playback time of the audio data segment.

[0054] Fig. 2 schematically illustrates processing processes with reference to Fig. 1. The transcription unit 20 of a server mentioned in a further embodiment of the invention processes the time-sequential audio bit data records ABD described above in order to determine, as a result of the processing processes explained below, non-overlapping transcript chunks TC as the current part of a time-annotated text data stream.

[0055] In a preprocessing step 21, the transcription unit 20 determines split points AZ within the audio bit data sets ABD using an initial AI model to detect sentence pauses and pauses between words. Each split point AZ (also referred to as a cut point or "split point") specifies the point in time of a pause between sentences or words. Compared to a second AI model mentioned later, this first AI model is smaller and therefore less complex and computationally intensive. The first AI model is selected and trained to focus primarily on timing rather than the accuracy of the transcript.

[0056] In a further preprocessing step 22, the transcription unit 20 uses another AI model to determine the probability value of each occurrence of speech (W1), music (W2), and / or noise (W3) in one or more consecutive audio bit data sets ABD at a predefined sampling rate, for example, once per second. The determined probability values ​​W1, W2, and W3 are then used in a subsequent compilation step 23 to decide whether or not the current audio bit data sets ABD should be transcribed into transcript chunks, i.e., into text. Only if the determined speech-related probability value W1 is above a predefined threshold SW1 and the music-related probability value W2 is below a predefined threshold SW2 does the transcription unit 20 perform a transcription of the current audio bit data sets ABD.

[0057] In the compilation process 23 (also referred to as the joining process), the transcription unit 20 assembles temporally successive audio bit data records ABD into a respective audio data chunk AC. The start and end times of an audio data chunk are determined by a respective split time AZ, calculated using the preprocessing process 21, such that the start time of a currently compiled audio data chunk AC precedes the end time of a previously compiled and temporally adjacent audio data chunk AC, and the temporally adjacent audio data chunks overlap by a predetermined minimum buffer length PL. The predetermined minimum buffer length is preferably the duration of an audio bit data record ABD and is therefore 5 seconds in one described embodiment.

[0058] In a transcription process 24, the transcription unit 20 transcribes an audio data chunk AC, currently determined by the compilation process 23, into an associated transcript chunk TC using a second AI model for transcription. The second AI model is a generative AI model, such as Whisper™ from OpenAI™, and is therefore more complex than the first AI model. As a result of the transcription process 24, each transcript chunk TC specifies the speech (in the sense of spoken sentences and words) contained in the audio data chunk AC as text annotated with time information. The text is specified as segments and tokens. A segment represents, for example, a sentence, and a token represents a word or part of a word. The time information specifies the start and end times of each segment or token in the audio data chunk AC.The timestamps, i.e., the annotated start and end times, refer to the relative position within the chunk. The transcription process then filters out any overlapping duplicate text from each transcript chunk to obtain a non-overlapping transcript chunk (TC). This process of filtering out overlapping duplicate text occurs incrementally. First, the process searches the previous chunk for the first five contiguous tokens of a new chunk that end at a time greater than 0 seconds. If several contiguous tokens are not found, the number of contiguous tokens is reduced by one until one is found. If even a single token is not found in the previous chunk, no filtering takes place, and the entire chunk is retained.

[0059] To improve the result of the transcription process 24, the transcription unit 20 provides error correction processes 25 and 26 that eliminate or at least reduce transcription errors contained in the transcript chunks.

[0060] A hallucination filtering process 25 uses a hallucination database to identify erroneously generated segments in a given transcript chunk TC. These segments were created during transcription by a hallucination of the second AI model without corresponding audio data. The process then removes the identified erroneously generated segment and its timing information from the transcript chunk. Filtering is performed using an expandable filter table. Each entry in the table results in the removal of every occurrence of that entry found in the transcribed text.

[0061] A content correction process 26 uses a third AI model to correct a content error in one or more transcript chunks (TC). This third AI model is a large-language model, such as GPT-4® from OpenAI™. The percentage of corrected words is limited to a predefined maximum value, preferably 5%. If this maximum value is exceeded and / or a predefined correction time is exceeded, the correction process is aborted.

[0062] After the transcription process 24 and the optional error correction processes 25, 26, the time information of the transcript chunks is adjusted in a timestamp adjustment process 27 to take into account time ranges in which the audio data stream contains no speech or spoken words but only music, and therefore no transcript chunks were created, as determined, for example, by the preprocessing process 22.

[0063] As a result of the execution of transcription process 24 and timestamp adjustment process 27, as well as the optional error correction processes 25 and 26, the transcription unit 20 delivers non-overlapping transcript chunks TC, each of which specifies a language or spoken text contained in the associated audio data chunk AC as annotated text with time information. These non-overlapping transcript chunks TC are then transmitted from server 2 to terminal 3.

[0064] Optionally, the transcription unit 20 can translate the transcription result described above—that is, from spoken language in one language to text in the same language—into text in another language or a foreign language. In a translation process 30, the transcription unit 20 translates transcript chunks TC of an original language into transcript chunks TCt of a specified foreign language. In translation process 30, each transcript chunk TC is translated into the specified foreign language by two parallel processes, 30-1 and 30-2. The first process, 30-1, uses a Large Language Model (LLM), for example, a GPT® version such as GPT-3.5 or GPT-4® from OpenAI™. The second process, 30-2, uses a translation AI model, for example, from DeepL® SE.If the translation result of LLM process 30-1 is available within a predefined translation time, for example, within 7 seconds, it is prioritized over the other translation result of process 30-2. If the predefined translation time is exceeded, only the translation result of process 30-2 is used, which is available, for example, within 1 to 2 seconds. Subsequently, in process 30-3, the timestamps of the translated text, i.e., the time information of the transcript chunks TCt, are adjusted appropriately so that the translated text can be displayed on the end device at least sentence-synchronously with the original language being played.

[0065] Fig. Figure 3 shows an exemplary assembly of audio bit data sets ABD into audio data chunks AC according to an embodiment of the present invention. The temporally sequential audio bit data sets ABD are in Fig. 3. Illustrated as wav data records or wav files wav_x ​​(where x ranges from 1 to 19) with a respective audio interval length of 5 seconds. As shown in Fig. As illustrated in Figure 3, the audio data stream is stored with a delay of 30 seconds, or in six audio bit data sets of 5 seconds each. Split times AZ, as results of the preprocessing process described above, are also shown in Figure 21. Fig. 3. Audio data chunks are represented in chronological order AZ1 to AZ5. Each audio data chunk AC is specified by its start time and end time, each of which is a split time AZ. For a currently assembled audio data chunk AC (for example, AC2), its start time (for example, AZ1) must precede the end time (AZ2) of a previously assembled and temporally adjacent audio data chunk AC (for example, AC1). Accordingly, the end time of audio data chunk AC2 is split time AZ4, which follows split time AZ3, which in turn is the start time of audio data chunk AC3. Temporally adjacent audio data chunks (for example, AC1 and AC2) must overlap by a predefined minimum buffer length PL, which is specified in Fig. 3 is exemplified as >5 seconds and thus as the duration of more than one audio bit data set ABD.

[0066] Fig. Figure 4 illustrates the coordination of parallel transcription processes for improved server utilization according to an embodiment of the present invention. As shown in Fig. As illustrated in Figure 4, the server receives three continuous audio data streams or transmitters in parallel. Two transcription processes that have already been initialized or started are continuously assigned audio data chunks from the three different audio data streams or transmitters to be transcribed. The ongoing transcription processes are informed via interprocess communication when new audio bit data records or audio data chunks are available. These are then read in and processed, taking into account the split points or intersection points. Reference symbol list 1 Audio data stream 2 servers 3, 3-1, 3-2, 3-3 Terminal 15 Audio / Video Processing Program 20 transcription units 21, 22 Preprocessing process 23. Compilation process 24 Transcription process 25 Hallucination filtering process 26 Content correction process 27 Timestamp Adjustment Process 30, 30-1, 30-2, 30-3 Translation process ABD Audio Bit Data Set AC audio data chunk AZ Division date ADS1, ADS2 audio data segment(s) PL buffer length or overlap length TC Transcript Chunk TCt transcript chunk of a translated text

Claims

[1] Method for real-time transcription of a continuous audio data stream (1) from an internet radio station into text, by a server (2), and for playing back the audio data stream and time-synchronously displaying the transcribed text on an end device (3, 3-1, 3-2, 3-3), wherein the server (2) performs the following steps: Retrieving a continuous audio data stream (1) from the internet radio station; Extracting or generating (15) time-sequential audio bit data sets (ABD), for example as WAV files, from the continuous audio data stream (1); Determining (21) split time points (Z) within the audio bit data sets (ABD), wherein each split time point (Z) specifies a time of a pause between sentences or words, using an initial AI model to detect sentence pauses and speech pauses between words; Assembling (23) successive audio bit data sets (ABD) into a respective audio data chunk (AC), the start time and end time of which are determined by a respective split time (AZ) such that the start time of a currently assembled audio data chunk (AC) is prior to an end time of an earlier assembled and temporally adjacent audio data chunk (AC), and the temporally adjacent audio data chunks (AC) overlap with a predetermined minimum buffer length (PL), in particular a buffer length of at least one audio bit data set (ABD); Transcribing (24) a respective audio data chunk (AC) into an associated transcript chunk (TC) using a second, in particular generative, AI model for transcription, wherein the second AI model is more complex than the first AI model, wherein a respective transcript chunk (TC) specifies a language contained in the audio data chunk (AC) as text annotated with time information, wherein the text is represented as segments and The token is specified, and the time information specifies the start and end time of each segment or token in the audio data chunk; Filtering out overlapping duplicate text from a given transcript chunk (TC) to obtain a non-overlapping transcript chunk (TC); Transmitting each non-overlapping transcript chunk (TC) as the current part of a time-annotated text data stream to the terminal (3, 3-1, 3-2, 3-3); and Transmitting the continuous audio data stream as temporally sequential audio data segments (ADS1, ADS2), wherein the audio data segments contain information about their occurrence times from the start of the audio data stream, where the terminal (3, 3-1, 3-2, 3-3) performs the following steps: Receiving the temporally sequential audio data segments (ADS1, ADS2) as an audio data stream from the server (2); Receiving the transcript chunks (TC) as a text data stream from the server (2); Playing back the audio data stream and displaying a text of the text data stream in a time-synchronized manner, wherein the relative start and end times of segments and tokens stored in the transcript chunks (TC) are converted into absolute start and end times with reference to the start of the audio data stream, and the terminal device (3, 3-1, 3-2, 3-3) displays a respective transcript chunk segment at the occurrence time of an audio data segment and preferably highlights a respective token in the displayed segment if its absolute start time is before and its absolute end time is after the occurrence time. [2] Method according to claim 1, wherein the server (2) uses a hallucination database to detect an erroneously generated segment in a respective transcript chunk (TC), which was generated during transcription by a hallucination of the second AI model without corresponding audio data, in particular by comparison with a filter table, and removes the detected erroneously generated segment and its time information from the transcript chunk (25). [3] Method according to claim 1 or 2, wherein the server (2) uses a third AI model, in particular a large language model, to correct the content of transcript chunks (TC) (26), wherein the proportion of corrected words is limited to a predetermined maximum value, preferably 5%, and the correction is aborted when a predetermined correction time is exceeded. [4] Method according to one of claims 1 to 3, wherein the server (2) determines a respective probability value of a current occurrence of speech, music and / or noise in the audio data stream and / or respective audio data chunks at a predetermined determination rate, for example once per second, using a further AI model (22), and the server transcribes the audio data chunks (AC) into transcript chunks (TC) only if the determined speech-related probability value is above a predetermined threshold and the music-related probability value is below a predetermined threshold, wherein preferably the server (2) communicates the determined probability values ​​or threshold exceedances / falling below to the terminal device (3, 3-1, 3-2, 3-3) for the purpose of adjusting the display mode. [5] Method according to one of claims 1 to 4, wherein the server (2) determines a respective FFT data set of a Fast Fourier Transform of the audio data for a respective audio data chunk (AC) and makes it available for retrieval by the terminal device (3, 3-1, 3-2, 3-3), for example in a respective JSON file per audio data chunk, and the terminal device visualizes frequency spectra of the audio data chunks on the basis of retrieved FFT data sets, in particular for pure music audio data chunks. [6] Method according to any one of claims 1 to 5, wherein the server (2) continuously translates transcript chunks (TC) into a specified foreign language using an AI model (30-1, 30-2) and creates foreign language transcript chunks (TCt) with adapted time information (30-3) and transmits them to the terminal device (3, 3-1, 3-2, 3-3), and the terminal device (3, 3-1, 3-2, 3-3) synchronizes the foreign language transcript chunks (TCt) with the playback of the audio data stream, based on the adapted time information. [7] Method according to claim 6, wherein the server (2) translates the transcript chunks (TC) by parallel use of a large language model (30-1) and a translation AI model (30-2), wherein the server prioritizes a first translation result produced with the large language model (30-1) over a second translation result produced with the translation AI model (30-2) if the first translation result is available to the server within a predetermined translation time, and if the predetermined translation time is exceeded, only the second translation result is used. [8] Method according to claim 6 or 7, wherein the server (2) generates an additional audio data stream with a synthetic voice from the transcript chunks (TCt) translated into a foreign language by using a further AI model and transmits it to the terminal device (3, 3-1, 3-2, 3-3), and wherein the terminal device (3, 3-1, 3-2, 3-3) starts a second playback process to play the additional audio data stream in addition to the first playback process to play the original audio data stream retrieved from the internet radio station, and the first playback process controls a playback position of the second playback process. [9] Method according to any one of claims 1 to 8, wherein the server (2) receives and transcribes audio bit data sets (ABD) of several continuous audio data streams or transmitters in parallel and operates parallel transcription processes for transcribing audio data chunks (AC) using a respective second AI model, wherein the server (2) further assigns audio data chunks (AC) of several audio data streams to a respective running transcription process for transcription, and wherein the server (2) operates separate transcription processes for different languages. [10] System (2, 3, 3-1, 3-2, 3-3) for real-time transcription of a continuous audio data stream (1) from an internet radio station into text, by a server (2) of the system, and for playing back the audio data stream and time-synchronously displaying the transcribed text on an end device (3, 3-1, 3-2, 3-3) of the system, wherein the server (2) and the end device (3, 3-1, 3-2, 3-3) are configured to perform the method according to any one of claims 1 to 9. [11] Server (2) for the system according to claim 10, wherein the server (2) is preferably equipped with a plurality of GPUs for accelerated processing of the generative AI. [12] Terminal device (3, 3-1, 3-2, 3-3) for the system according to claim 10. [13] Computer program with executable instructions which, when executed by a server (2), cause the server (2) to execute a method according to one of claims 1 to 9 and, when executed by an end device (3, 3-1, 3-2, 3-3), cause the end device to execute a method according to one of claims 1, 4 to 6, or 8. [14] Data carrier containing a computer program according to claim 13, wherein the data carrier is a computer-readable storage medium, a server accessible via the Internet, an electronic signal, an optical signal and / or a radio signal.

Citation Information

Patent Citations

  • Speech recognition with sequence-to-sequence models

    US20200126538A1

  • Dynamic model selection in speech-to-text processing

    US20210295846A1

  • Systems and methods for providing real-time automated language translations

    US20230096543A1