Method and system for the real-time transcription of an audio data stream into text and for the time-synchronous reproduction of the audio data stream and the text

A server-based AI system transcribes internet radio audio streams into synchronized text on terminal devices, addressing the challenge of real-time transcription and display for improved user experience.

WO2025176252A1PCT designated stage Publication Date: 2025-08-28RADIOZEIT GMBH

Patent Information

Application Number
PCT/DE2025/100146
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-13
Filing Date
2025-02-07
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing systems fail to provide real-time transcription of audio streaming content with high spoken content into synchronized text due to high costs and lack of real-time capability, especially for internet radio stations with low playback volume or background noise, and hearing impairments.

Method used

A method and system that utilize a server to process audio streams using AI models to transcribe speech into text in real-time with a delay of up to one minute, displaying the text synchronously with audio playback on terminal devices, employing multiple AI models for accuracy and efficiency.

Benefits of technology

Enables real-time transcription and synchronized text display on terminal devices, reducing costs and improving user experience by ensuring text is displayed in sync with audio, even in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure DE2025100146_28082025_PF_FP_ABST
    Figure DE2025100146_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method and a system for the real-time transcription of a continuous audio data stream (1) from an Internet radio transmitter into text by means of a server (2), and for playing back the audio data stream and time-synchronously displaying the transcribed text on a terminal (3, 3-1, 3-2, 3-3). The continuous audio data stream is divided into temporally continuous audio data segments (ADS1, ADS2) and transmitted to a terminal. The terminal (3, 3-1, 3-2, 3-3) is designed to receive the temporally continuous audio data segments (ADS1, ADS2) as an audio data stream, to receive the transcript chunks (TC) as a text data stream, and to play back the audio data stream and time-synchronously display text of the text data stream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method and system for real-time transcription of an Android data stream into a text and for time-synchronous playback of the Android data stream and the text

[0002] The present invention relates to a method and system for real-time transcribing a continuous audio data stream of an Internet radio station into a text, by a server, and for playing the audio data stream and time-synchronously displaying the transcribed text on a terminal device.

[0003] State of the art

[0004] Nowadays, there is a huge variety of internet radio stations whose audio streaming content can be accessed or received via the internet using suitable devices such as smartphones, tablets, smart TVs, TV sticks, SMS watches, smart speakers, and internet-enabled digital radios. Particularly with internet radio stations with a high proportion of spoken content and little music content, the spoken words are often difficult to understand, for example due to the necessary low playback volume, background noise near the device, or the hearing impairment of the device user. It would be desirable if the audio streaming content from stations with a high proportion of spoken content were also displayed as text on the device screen in such cases. For high acceptance and genuine user comfort of such a solution, it is important that the text is displayed synchronously with the spoken words.

[0005] One approach for such a system with synchronized speech output and text display could be to provide the streaming content as retrievable or streamable video files. However, this approach is impractical due to the high cost for the operator and the large amounts of data that would need to be transmitted. A very significant disadvantage of this approach would also be the system's lack of real-time capability, since the text could only be integrated into the video after the currently playing speech or speaker's voice has ended.

[0006] Therefore, there is a need for a method and / or a system that automatically converts the speech contained in a streaming content into text during the continuous retrieval of the streaming content, i.e. in real time, but with the admissibility of a continuous delay in the range of approximately one minute due to processing effort, and displays the text synchronously together with the speech reproduction. Disclosure of the invention

[0007] The object of the invention is to provide a method and a system which make it possible to automatically and in real time create a corresponding text from a speech or spoken words contained in an audio streaming content, and to display the created text synchronously with the playback of the audio streaming content on a terminal device.

[0008] In this application, the term "language" generally refers to the language content composed of words and sentences, or the spoken part of an audio data stream. Only in the context of a translation of the transcribed text into a foreign language or another national language does the term "language" in this application have a country reference, such as "native language," "foreign language," "original language," or "national language."

[0009] To implement the method or system, it is necessary to delay the source signal from the internet radio transmitter to a sufficient minimum. For example, a total delay of 50 seconds is provided, with 30 seconds for recording the audio signal and a further 20 seconds for overlapping, transcribing, and translating the audio signal if necessary. However, this constant delay explained above does not contradict the characteristic of real-time transcribing. The term "real-time" refers, in this case, to the fact that the proposed method does not exceed the specified maximum total delay.

[0010] The object of the invention is achieved by a method having the features of claim 1. Furthermore, the object of the invention is achieved by a system having the features of claim 10. Advantageous embodiments emerge from the additional features of the dependent claims.

[0011] A method is proposed for real-time transcribing a continuous audio data stream of an Internet radio station into a text by a server, and for playing the audio data stream and time-synchronously displaying the transcribed text on a terminal device.

[0012] The audio stream content is played on the device, and the transcribed text is displayed synchronously. The actual processing and transcription takes place on the server, which is preferably a Taternet server. This allows multiple devices with little computing power to use a central server, which, for example, uses GPUs to process a very complex or even the most complex AI model for transcription.

[0013] The server retrieves a continuous audio data stream from the internet radio station and extracts or generates temporally continuous audio bit records from the continuous audio data stream. The audio bit records can, for example, be WAV files (WAV ::: Waveform Audio File Format). Each audio bit record represents, for example, a 5-second audio signal interval.

[0014] In order to avoid mid-word cut start and end points of the audio data stream, it is proposed to first identify ideal cut points, for example after a sentence or in a pause between two words, and then to make these cut points available to a process for joining or assembling the short audio bit data sets into an audio data chunk.

[0015] The server therefore determines splitting points or intersection points within the audio bit data sets, with each splitting point specifying a pause between sentences or words, using a first AI model to detect sentence pauses and speech pauses between words. Compared to a second AI model mentioned later, the first AI model is a smaller and thus less complex and computationally intensive AI system. The first AI model is selected, adapted, or trained in such a way that attention is paid primarily to timing and less to the correctness of the transcript.

[0016] The server compiles consecutive audio bit data records into a respective audio data chunk, whose start time and end time are determined by a respective determined splitting time such that the start time of a currently compiled audio data chunk lies before the end time of a previously compiled and temporally adjacent audio data chunk, and the temporally adjacent audio data chunks overlap by a predetermined minimum buffer length. The predetermined minimum buffer length is preferably a buffer length of at least one audio bit data record, for example, 5 seconds.

[0017] The server transcribes each audio data chunk into an associated transcript chunk using a second AI model for transcription. The second AI model is preferably a generative AI model and is more complex than the first AI model. The second AI model is, for example, Whisper from QpenAI. Each transcript chunk specifies a language contained in the audio data chunk (in the form of sentences and words) as text annotated with punctuation information. The text is specified as segments and tokens. A segment, for example, represents a sentence, and a token represents a word or part of a word. The time information specifies the start and end times of each segment or token in the audio data chunk. The timestamps, i.e., the annotated start and end times, refer to the relative position within the chunk. Time 0 marks the starting point without overlap.With overlap, the first timestamps can therefore become negative (for example, a maximum of -10s).

[0018] The server then filters out overlapping duplicate text from each transcript chunk to obtain an overlap-free transcript chunk. The process for filtering out overlapping duplicate text occurs step by step, first searching the previous chunk for the first 5 contiguous tokens of a new chunk that end at a time >0 seconds. If multiple contiguous tokens are not found, the number of contiguous tokens is reduced by 1 until it is found. If even a single token is not found in the previous chunk, no filtering occurs, and the entire chunk is retained.

[0019] The server then transmits a respective overlap-free transcript chunk as the current part of a time-annotated text data stream to the end device.

[0020] In addition, the server transmits the continuous audio data stream as temporally continuous audio data segments, for example according to the DASH or HLS Freanting protocol, wherein the audio data segments contain information about their occurrence times from the start of the audio data stream.

[0021] The end device receives the temporally continuous audio data segments as an audio data stream from the server. The end device also receives the transcript chunks as a text data stream from the server.

[0022] The terminal device plays the audio data stream and displays a text of the text data stream in a time-synchronized manner, whereby the relative start and end times of segments and tokens stored in the transcript chunks are converted into absolute start and end times with reference to the start of the audio data stream. The terminal device displays a respective transcript chunk segment at the time of occurrence of an audio data segment and preferably visually highlights a respective token in the displayed segment if its absolute start time (relative to segment and token) lies before and its absolute end time after the time of occurrence of the audio data segment. The playback of the audio data stream described above and the time-synchronized display of a text of the text data stream require synchronization of the data of the audio data stream and the text data stream, the implementation of which is illustrated by way of example below.The end device receives the audio data stream using a streaming protocol, for example, as a DASH stream containing an OPUS audio codec, or as an HLS stream containing an AAC audio codec. These streaming protocols provide millisecond-accurate information about when the stream was started on the server side, as they themselves deliver the audio data stream in individual audio data segments via the HTTP protocol.

[0023] Based on the respective millisecond-accurate position of the audio data stream or audio data segments, the device can retrieve the relevant part of the transcript, for example, a 30-second section of the transcript chunk. For this purpose, a consecutive number for the chunk can be calculated according to:

[0024] Chuuk number = [playback time / 30s]

[0025] The broadcast script chunks are transmitted, for example, in JSON format or as a subtitle data track in WebVTT (Web Video Text Tracks) or TTML (Timed Text Markup Language) format as part of a DASH or HLS stream. This contains a list of segments (e.g., sentences), each of which in turn contains a list of tokens (e.g., words or parts of words). Each segment and token contains a start time and an end time in milliseconds. The end device reads these segments with the respective tokens and stores them in a mapping table. The relative

[0026] Start and stop times within the transcript chunks are converted to absolute times from the start of the audio stream on the server side according to: t__a“Chunk_Num.mer*30-rt__r where t_r is the relative start and end time within the transcript chunk, and t__a is the absolute start and end time since the start of the stream on the server.

[0027] This mapping table allows the end device to display the correct segment by retrieving the segment at the current time in milliseconds and to highlight the current (partial) word or token in the same way. The presented method enables playback of the audio data stream with a time-synchronized display of the transcribed text on the end device in the millisecond range.

[0028] When transcribing with current generative AI models, such as OpenAI's Whisper, so-called hallucinations still occur. For example, copyright notices or general information about subtitling are displayed after completed sentences during a long period of audio silence. The reason for this is that such notices were apparently present in the AI ​​model's training data, and the model learned to generate them after transcription. Further examples include farewell phrases such as "See you next time," which were not present in the audio stream but appear in the transcript.

[0029] In a further development of the presented method, the server uses a hallucination database to detect an incorrectly generated segment in a respective transcript chunk. This segment was generated during transcription by a hallucination of the second AI model without corresponding audio data. The server then removes the detected, incorrectly generated segment and its time information from the transcript chunk. Filtering is performed using an expandable filter table. Each entry in the table results in every occurrence of the entry found in the transcribed text being removed.

[0030] In a further development of the presented method, the server corrects the content of transcript chunks using a third AI model. A correction of a content error in one or more transcript chunks is performed by sending a query to the third AI model. The third AI model is preferably a large language model, such as OpenAI's G.PT-4. The proportion of corrected words is limited to a specified maximum value, preferably 5%. If the specified maximum value is exceeded and / or if a specified correction time is exceeded, the correction is aborted.

[0031] In a further development of the presented method, the server uses an additional AI model to determine a respective probability value for a current occurrence of speech, music, and / or noise in the audio data stream and / or respective audio data chunks at a predetermined determination rate, e.g., once per second. The server only transcribes the audio data chunks into transcript chunks if the determined speech-related probability value is above a predetermined threshold and the music-related probability value is below a predetermined threshold. This saves resources and thus reduces power consumption and costs. Furthermore, the occurrence of possible hallucinations is reduced. For this purpose, further division times or thresholds are used to determine whether the threshold values ​​are exceeded or not met.Cut points are sent to the transcription process via audio processing communication. Sections that do not need to be transcribed are then cut from the audio data stream or audio bit data sets. These missing sections in the transcript are then reinserted by incrementing all subsequent timestamps by the duration of the section.

[0032] The Serv r It preferably communicates the determined probability values ​​or threshold violations to the terminal device for display mode adjustment. When music is playing, the terminal device can then display the spectrum analyzer or frequency spectra described below, along with an album cover of the respective song, instead of the transcript.

[0033] In a further development of the presented method, the server determines a respective FFT data set of a Fast Fourier Transformation of the audio data for each audio data chunk and makes it available to the end device for retrieval. On the basis of retrieved FFT data sets, the end device visualizes frequency spectra of the audio data chunks, in particular for pure music audio data chunks. This avoids complex processing of the raw data of the audio signal on the end device, provided this is even possible on the end device. An FFT data set is made available, for example, in a respective JSON file for each audio data chunk on the server for retrieval by the end device. According to this further development, for example, an FFT is performed on the server, and the resulting data, such as 16 frequency bands, is sampled at a rate of, for example, 50 Hz.For example, the data volume corresponds to a maximum of 80 kBytes for 30 seconds, which is one fifth of the audio data to be transmitted.

[0034] A further development of the presented method is aimed at the server-side, time-synchronous translation of the transcript into a foreign language. To display the translated transcript synchronously with the audio data stream, it is translated on the server side, and the timestamps of the original language transcript are also incorporated. Since the translated texts can be longer, shorter, or in a different order than the original, the timestamps must be adjusted. To achieve this, the transcript chunk is translated in the first step.

[0035] In this further development of the presented method, it is accordingly provided that the server continuously translates transcript chunks into a given foreign language by using a Kl model and creates foreign language transcript chunks provided with adapted time information and transmits them to the terminal device, and the terminal device synchronizes the foreign language transcript chunks with the playback of the audio data stream on the basis of the adapted time information.

[0036] In a further development of the training presented above, the server also translates the transcript chunks using a large language model, for example, a GPT version such as GPT3.5 or GPT4 from OpenAI, and a translation AI model, for example, from DeepL SE. The server prioritizes a first translation result generated with the large language model over a second translation result generated with the translation AI model if the first translation result is available to the server within a specified translation time, for example, within 7 seconds. If the specified translation time is exceeded, only the second translation result, which is available within 1 to 2 seconds, is used.

[0037] Once the translated transcript is available, the text can be divided into segments and tokens. To do this, the individual words are first divided evenly into segments so that the same number of segments is achieved in the translation. The start and end times of the translated segments are then distributed so that the segments, which now contain roughly the same number of words, are evenly distributed within the time span of the chunk. In the final step, the individual letters of the words are evenly distributed in each segment as its tokens. This means that each letter can be roughly the same length in the frontend. Although the currently spoken word is not always displayed in the translation, it is easy to follow along, and the display points in the original language and the translation are thus as synchronized as possible.

[0038] A further development of the recently presented training and further education of the presented procedure is aimed at the synchronous playback of the translation with a synthetically generated voice.

[0039] The server generates an additional audio data stream with a synthetic voice from the transcript chunks translated into a foreign language using an additional audio model and transmits it to the terminal device. In addition to the first playback process for playing the original audio data stream retrieved from the internet radio station, the terminal device starts a second playback process for playing the additional audio data stream. The first playback process controls a playback position of the second playback process. In the event of a playback pause, or if both playback processes are no longer synchronized to within a millisecond plus / minus a permitted variance, the second playback process is adjusted to the playback position of the first playback process.

[0040] So-called "ducking" is implemented using information from the Kl model for music recognition. Information about the presence of speech is transmitted to the end device at regular intervals, analogous to the transmission of FFT information. If it is detected that no relevant speech is present, the translated audio track is set to a volume of 0%, and the original audio track is set to 100% of the desired volume. As soon as translated audio becomes relevant, the original audio track's volume is set to a lower value, for example, 25% of the desired volume. The volume of the translated audio track is increased to 100%. This crossfading occurs dynamically within a few seconds, for example, 2 seconds.Similarly, the crossfade back to sections without speech takes place, where the volume level of the translated audio track is reduced to 0%, and at the same time the level of the original audio track is increased back to 100% of the volume.

[0041] A further development of the presented method is aimed at coordinating parallel transcription processes in order to improve the utilization of the server hardware. Transcription is planned to take place on special expansion cards in the server. As an example, two Nvidia GPUs per server are mentioned here. These GPUs accelerate the processing of the generative AI by a factor of approximately 10. Since several internet radio stations are to be available on the end device, it is important to process as many data streams as possible in parallel. Starting the transcription process requires additional time. It is therefore not optimal to start and then terminate a process for each data chunk. Therefore, a method is proposed in which already initialized processes can process newly incoming data.For this purpose, running processes are informed via interprocess communication when new audio bit data sets, such as WAV files, are available. These are then read in and processed taking into account the recording times or intersection points. In a specific example case, the AI ​​model used, for example, Whisper from OpenAI, requires at least 2 GB of RAM on the expansion card mentioned above, which limits the number of parallel AI processes to, for example, 6 per GPU. Therefore, audio bit data sets from multiple transmitters are processed per transaction process.

[0042] According to this development of the presented method, it is provided that the server receives and transcribes audio bit data records of several continuous audio data streams or transmitters in parallel and operates parallel transcribing processes for transcribing audio data chunks using a respective second Kl model, wherein the server further assigns audio data chunks of several audio data streams to a respective ongoing transcribing process for transcription, and wherein the server operates separate transcribing processes for different national languages ​​or foreign languages.

[0043] Furthermore, a system is presented for real-time transcribing a continuous audio data stream from an internet radio station into text, using a server of the system, and for playing the audio data stream and synchronously displaying the transcribed text on a terminal device of the system. The server and the terminal device form the system and are configured to execute the method presented above and its further developments and enhancements.

[0044] Furthermore, a computer program or computer program product with executable instructions is presented which, when executed by a server, cause the server to execute the method presented above and its further developments and further training, and which, when executed by a terminal device, cause the terminal device to execute the method presented above and its further developments and further training.

[0045] Finally, a data carrier with the computer program presented above is presented, wherein the data carrier is a computer-readable storage medium, a server accessible via the Internet, an electronic signal, an optical signal and / or a radio signal.

[0046] The features of the method and system presented above can be combined with one another in any way within the meaning of the invention, provided that such a combination is not obviously contradictory. Features of the claims to which the words "in particular" and "preferably" are assigned are to be understood as optional and non-limiting and serve to specify particularly advantageous embodiments of the present invention.

[0047] Short description of the drawings

[0048] To clarify the presented method and the presented system, implementation examples are now presented with reference to the following figures.

[0049] Fig. 1 schematically illustrates a system for real-time transcribing an audio data stream from an internet radio station into text by a server, and for playing the audio data stream and synchronously displaying the transcribed text on a terminal device according to an embodiment of the invention. Fig. 2 schematically illustrates processing steps of a transcribing unit of the server according to another embodiment of the invention;

[0050] Fig. 3 shows an example of assembling audio bit data sets into audio data chunks according to an embodiment of the present invention.

[0051] Fig. 4 illustrates a coordination of parallel transcription processes according to an embodiment of the present invention.

[0052] In the figures shown, identical or similar elements, components, values ​​and intervals are designated by the same reference symbols throughout the figures.

[0053] Embodiments of the invention

[0054] Fig. 1 schematically illustrates a system for real-time transcribing an audio data stream from an Internet radio station, by a server, into text and for playing the audio data stream and synchronously displaying the transcribed text on a terminal device according to an embodiment of the invention. The system comprises a server 2 and a terminal device 3. The server 2 is designed as an Internet server for retrieving, receiving, and processing a continuous audio data stream 1 provided by an Internet radio station. The server 2 is designed to extract or generate temporally continuous audio bit data records ABD and temporally continuous audio data segments ADS1, ADS2 from the retrieved audio data stream 1, for example by using an audio / video processing program 15 illustrated in Fig. 1, which is indicated in Fig. 1 as the open source program ßmpeg.

[0055] The temporally continuous audio bit data sets ADB are stored in WAV format or as WAV files (WAV = :: Waveform Audio File Format) with a predetermined audio signal interval of, for example, 5 seconds is provided to a transcribing unit 20, described in detail later with reference to Fig. 2. The transcribing unit 20 of the server 2 is configured to continuously determine time-annotated transcript chunks TC as the current part of a time-annotated text data stream based on the temporally continuous audio bit data records .ABD. The transcript chunks TC represent the speech contained in the received audio data stream 1, i.e., spoken sentences and words or segments and tokens, and are transmitted from the server 2 to the terminal 3 as the current part of the time-annotated text data stream.

[0056] The temporally continuous audio data segments ADS1, ADS2 are provided for transmission to the terminal 3 according to the DASH streaming protocol (based on an OPUS audio codec) or the HLS streaming protocol (based on an AAC audio codec) with a predetermined time delay of, for example, 50 seconds. This time delay is predetermined or selected such that in addition to a recording time interval for recording the audio signal contained in the audio data stream of, for example, 30 seconds, a processing time interval for processing the audio signal (with the length of the recording time interval) by the transcribing unit 20 of, for example, a further 20 seconds is provided.The audio data segments ADS1, ADS2 contain information about their occurrence times from the start of the audio data stream, since the streaming protocols used provide millisecond-accurate information about when the audio data stream was started on the server side.

[0057] As schematically illustrated in Fig. 1, the server 2 can transmit the generated audio data segments ADS1, ADS2 and the transcription chunks TC determined by the transcribing unit to a variety of terminal devices 3. The terminal device 3 can be designed as a universally applicable device, for example an iOS device 3-1, such as an iPad or iPhone, or an Android device or desktop device 3-2, or as a device 3-3 specifically developed for visualizing Internet radio speech. The terminal device 3 is connected to the server 2 for wireless or wired data communication via the Internet. The terminal device 3 is designed to receive the temporally continuous audio data segments ADS1, ADS2 as an audio data stream and the transcription chunks TC as a text data stream from the server 2. The terminal device 3 is further designed to play the audio data stream and to display a text of the text data stream in a time-synchronized manner.In doing so, the terminal device 3 converts the relative start times and end times of segments and tokens stored in the transcript chunks TC into absolute start times and end times with reference to the start of the audio data stream. Furthermore, the terminal device 3 displays a respective segment, for example a sentence, of a transcript chunk TC at an occurrence time of an audio data segment ADS1, ADS2 if the absolute start time of the segment is before and the absolute end time of the segment is after the occurrence time. In addition, the terminal device 3 can visually highlight a respective token, for example a word or part of a word, in the currently displayed segment if the start time of the token is before and the end time of the token is after the current absolute playback time of the audio data segment.

[0058] Fig. 2 schematically illustrates processing processes of the transcription unit 20 of a server mentioned with reference to Fig. 1 according to another embodiment of the invention. The transcription unit 20 processes the above-described, temporally continuous audio bit data sets ABD to determine, as a result of the processing processes explained below, non-overlapping transcript chunks TC as the current part of a time-annotated text data stream.

[0059] In a preprocessing step 21, the transcribing unit 20 determines splitting points AZ within the audio bit data sets ABD using a first AI model to detect sentence pauses and speech pauses between words. Each splitting point AZ (also referred to as a split point) specifies the point in time of a pause between sentences or words. Compared to a second AI model mentioned later, the first AI model is a smaller and thus less complex and less computationally intensive AI system. The first AI model is selected or trained in such a way that attention is paid primarily to timing and less to the correctness of the transcript.

[0060] In a further preprocessing process 22, the transcribing unit 20 uses a further KI model to determine a respective probability value of a current occurrence of speech (W1), music (W2), and / or noise (W3) in one or more consecutive audio bit data sets ABD at a predetermined determination rate of, for example, once per second. The determined probability values ​​W1, W2, W3 are used in a subsequent compilation process 23 to decide whether or not the current audio bit data sets ABD should be transcribed into transcript chunks, i.e., into text. Only if the determined speech-related probability value W1 is above a predetermined threshold value SW1 and the music-related probability value W2 is below a predetermined threshold value SW2 does the transcribing unit 20 perform a transcription for the current audio bit data sets ABD.

[0061] In the compilation process 23 (also referred to as the animation joining process), the transcribing unit 20 compiles temporally successive audio bit data records ABD into a respective audio data chunk AC. The start time and the end time of an audio data chunk are determined by a respective division time AZ determined by the preprocessing process 21 such that the start time of a currently compiled audio data chunk AC lies before the end time of a previously compiled and temporally adjacent audio data chunk AC, and the temporally adjacent audio data chunks overlap by a predetermined minimum buffer length PL. The predetermined minimum buffer length is preferably a duration of an audio bit data record ABD and, in one described embodiment, is thus 5 seconds.

[0062] In a transcription process 24, the transcription unit 20 transcribes an audio data chunk AC currently determined by the compilation process 23 into an associated transcript chunk TC using a second AI model for transcription. The second AI model is a generative AI model, for example, Whisper from OpenAI, and is thus more complex than the first AI model. As a result of the transcription process 24, a respective transcript chunk TC specifies a language contained in the audio data chunk AC (in the sense of spoken sentences and words) as text annotated with time information. The text is specified as segments and tokens. A segment represents, for example, a sentence, and a token represents a word or part of a word. The time information specifies the start time and end time of a respective segment or token in the audio data chunk AC.The timestamps, i.e., the annotated start and end times, refer to the relative position within the chunk. The transcription process 24 then filters out overlapping duplicate text from a respective transcript chunk to obtain an overlap-free transcript chunk TC. The process for filtering out overlapping duplicate text occurs step by step, first searching the previous chunk for the first 5 contiguous tokens of a new chunk that end at a time >0 seconds. If multiple contiguous tokens are not found, the number of contiguous tokens is reduced by 1 until it is found. If even a single token is not found in the previous chunk, no filtering takes place, and the entire chunk is retained.

[0063] To improve the result of the transcription process 24, the transcription unit 20 provides error correction processes 25 and 26 that eliminate or at least reduce transcription errors contained in the transcript chunks.

[0064] Eia Hallucination Filtering Process 25 determined by using a

[0065] The hallucination database identifies an incorrectly generated segment, which was generated during transcription by a hallucination of the second AI model without corresponding audio data, in a respective transcript chunk (TC) and removes the identified, incorrectly generated segment and its time information from the transcript chunk. Filtering is performed using an expandable filter table. Each entry in the table results in every occurrence of the entry found in the transcribed text being removed.

[0066] A content correction process 26 corrects a content error in one or more transcript chimes (TC) by using or querying a third AI model. The third AI model is a large language model, such as OpenAI's GPT-4.

[0067] The percentage of corrected words is limited to a specified maximum, preferably 5%. If this maximum is exceeded and / or a specified correction time is exceeded, the correction process is aborted.

[0068] After the transcribing process 24 and the optional error correction processes 25, 26, the time information of the transcript chunks is adjusted in a time stamp adjustment process 27 to take into account time periods in which the audio data stream contains no speech or spoken words but only music, and thus no transcript chunks were created, as determined, for example, by the preprocessing process 22.

[0069] As a result of performing the transcribing process 24 and the timestamp adjustment process 27, as well as the optional error correction processes 25 and 26, the transcribing unit 20 generates overlap-free transcript chunks TC, each of which specifies a speech or spoken word contained in the associated audio data chunk AC as text annotated with time information. These overlap-free transcript chunks TC are subsequently transmitted from the server 2 to the terminal 3.

[0070] Optionally, the transcription unit 20 can translate the transcription result presented above—i.e., from speech in a national language into text in the same national language—into a text in another language or foreign language. In a translation process 30, the transcription unit 20 translates transcript chunks TC of an original language into transcript chunks TCt of a given foreign language. In the translation process 30, each transcript chunk TC is translated into the given foreign language by two parallel processes 30-1 and 30-2. The first process 30-1 uses a large language model (LLM), for example, a GPT version such as GPT3.5 or GPT4 from OpenAI. The second process 30-2 uses a translation AI model, for example, from DeepL SE.If the translation result of LLM process 30-1 is available within a specified translation time, for example, within 7 seconds, it is prioritized over the other translation result of process 30-2. If the specified translation time is exceeded, only the translation result of process 30-2, which is available within 1 to 2 seconds, for example, is used. Subsequently, in a process 30-3, the timestamps of the translated text, i.e., the time information of the transcript chunks TCt, are appropriately adjusted so that the translated text can be displayed on the end device at least sentence-synchronously with the original language being played back.

[0071] Fig. 3 shows an example of a compilation of audio bit data records ABD to form audio data chunks AC according to an embodiment of the present invention. The temporally successive audio bit data records ABD are illustrated in Fig. 3 as wav data records or wav files way_x (with x from 1 to 19) with a respective audio interval length of 5 seconds. As illustrated in Fig. 3, the audio data stream is stored with a delay of 30 seconds or 6 audio bit data records of 5 seconds each. Splitting times AZ as results of the preprocessing process 21 described above are also illustrated in Fig. 3 in chronological order AZ1 to AZ5. A respective audio data chunk AC is specified by its start time and its end time, which are each a splitting time AZ.For a currently compiled audio data chunk AC (e.g., AC2), its start time (e.g., AZ1) must be temporally before an end time (AZ2) of a previously compiled and temporally adjacent audio data chunk AC (e.g., AC1). Accordingly, the end time of the audio data chunk AC2 is the splitting time point AZ4, which temporally follows the splitting point AZ3; which in turn is the start time point of the audio data chunk AC3. Temporarily adjacent audio data chunks (e.g., AC1 and AC2) must overlap by a predetermined minimum buffer length PL, which is specified in Fig. 3 as >5 seconds and thus as the duration of more than one audio bit data record ABD.

[0072] Fig. 4 illustrates the coordination of parallel transcription processes for improved server utilization according to an embodiment of the present invention. As illustrated in Fig. 4, the server receives three continuous audio data streams or transmitters in parallel. Two already initialized or started transcription processes are continuously assigned audio data chunks from the three different audio data streams or transmitters to be transcribed. The running transcription processes are informed via interprocess communication when new audio bit data sets or audio data chunks are available. These are then read in and processed taking into account the splitting times or intersection points.

[0073] Reference list

[0074] 1 audio data stream

[0075] 2 sex

[0076] 3, 3-1, 3-2, 3-3 terminal

[0077] 15 Audio video processing program

[0078] 20 Transcription Unit

[0079] 21, 22 Pre-processing process

[0080] 23 Compilation process

[0081] 24 Transcription process

[0082] 25 Hallucination filtering process

[0083] 26 Content correction process

[0084] 27 Timestamp adjustment process

[0085] 30, 30-1, 30-2, 30-3 Translation process

[0086] ABD audio bit record

[0087] AC, A.C1, AC2, AC3 audio data chunk

[0088] AZ, AZ1, AZ2, AZ3, AZ4, AZ5 allocation time

[0089] ADS1, ADS2 audio data segment(s)

[0090] PL buffer length or overlap length

[0091] TC transcript chunk

[0092] TCt Transcript Chuitk of a translated text

Claims

Patent claims 1. A method for real-time transcribing a continuous audio data stream (1) of a hitemetradio transmitter into a text, by a server (2), and for playing the audio data stream and synchronously displaying the transcribed text on a terminal (3, 3-1, 3-2, 3-3), wherein the server (2) performs the following steps: Retrieving a continuous audio data stream (1) from the internet radio transmitter; Extracting or generating (15) temporally continuous audio bit data sets (ABD), for example as WAV files, from the continuous audio data stream (1); Determining (21) division times (AZ) within the audio bit data sets (ABD), wherein a respective division time (AZ) specifies a time of a pause between sentences or words, using a first Kl model for detecting sentence pauses and speech pauses between words; Dividing (23) successive audio bit data sets (ABD) into a respective audio data chunk (AC), the start time and end time of which are determined by a respective determined division time (AZ) such that the start time of a currently compiled audio data chunk (AC) lies before an end time of a previously compiled and temporally adjacent audio data chunk (AC), and the temporally adjacent audio data chunks (AC) overlap with a predetermined minimum length (PL), in particular a buffer length of at least one audio bit data set (ABD); Transcribing (24) a respective audio data chunk (AC) into an associated transcript chunk (TC) using a second, in particular generative, Kl model for transcription, wherein the second Kl model is more complex than the first Kl model, wherein a respective transcript chunk (TC) specifies a language contained in the audio data chunk (AC) as text annotated with time information, wherein the text is specified as segments and tokens, and the time information specifies the start time and end time of a respective segment or token in the audio data chunk: Filtering out an overlapping duplicate text from a respective transcript chunk (TC) to obtain an overlap-free transcript chunk (TC); Transmitting a respective overlap-free transcript chunk (TC) as the current part of a time-annotated text data stream to the terminal device (3, 3-1, 3-2, 3-3); and Transmitting the continuous audio data stream as temporally continuous audio data segments (ADS1, ADS2), for example according to the DASH or HLS streaming protocol, wherein the audio data segments contain information about their occurrence times from the start of the audio data stream, wherein the terminal device (3, 3-1, 3-2, 3-3) performs the following steps: Receiving the temporally continuous audio data segments (ADS1, ADS2) as an audio data stream from the server (2); Receiving the transcript chunks (TC) as a text data stream from the server (2); Playing the audio data stream and time-synchronized display of a text of the text data stream, wherein the relative start times and end times of segments and tokens stored in the transcript chunks (TC) are converted into absolute start times and end times with reference to the start of the audio data stream, and the terminal (3, 3-1, 3-2, 3-3) displays a respective transcript chunk segment at an occurrence time of an audio data segment and preferably optically highlights a respective token in the displayed segment if its absolute start time is before and its absolute end time is after the occurrence time.

2. The method according to claim 1, wherein the server (2) uses a hallucination database to identify an erroneously generated segment in a respective transcript chunk (TC), which segment was generated during transcribing by a hallucination of the second Kl model without corresponding audio data, in particular by means of a comparison with a filter table, and removes the identified, erroneously generated segment and its time information from the transcript chunk (25).

3. Method according to claim 1 or 2, wherein the server (2) corrects (26) the content of TransGipt-Chuuks (TC) by using a third Kl model, which is in particular a large language model, wherein the proportion of corrected words is limited to a predetermined maximum value, preferably 5%, and the correction is aborted when a predetermined correction time period is exceeded.

4. The method according to one of claims 1 to 3, wherein the server (2) determines (22) a respective probability value of a current occurrence of speech, music and / or noise in the audio data stream and / or respective audio data chunks at a predetermined determination rate, for example once per second, by using a further KI model, and the server transcribes the audio data chunks (AC) into transcript chunks (TC) only if the determined speech-related probability value is above a predetermined threshold value and the music-related probability value is below a predetermined threshold value, wherein preferably the server (2) communicates the determined probability values ​​or threshold exceedances / undershoots to the terminal (3, 3-1, 3-2, 3-3) for the purpose of display adaptation.

5. The method according to one of claims 1 to 4, wherein the server (2) determines a respective FFT data set of a fast Fourier transform of the audio data for a respective audio data chunk (AC) and makes it available to the terminal (3, 3-1, 3-2, 3-3) for retrieval, for example in a respective JSON file per audio data chunk, and the terminal visualizes frequency spectra of the audio data chunks, in particular for pure music audio data chunks, on the basis of retrieved FFT data sets.

6. The method according to one of claims 1 to 5, wherein the server (2) continuously translates transcript chunks (TC) into a predetermined foreign language (30-1, 30-2) by means of a KI model and creates foreign language transcript chunks (TCt) provided with adapted time information (30-3) and transmits them to the terminal (3, 3-1, 3-2, 3-3), and the terminal (3, 3-1, 3-2, 3-3) displays the foreign language transcript chunks (TCt) synchronously with the playback of the audio data stream on the basis of the adapted time information.

7. The method according to claim 6, wherein the server (2) translates the transcript chunks (TC) by means of parallel use of a large language model (30-1), for example a GPT version of OpenAI, and a translation AI model (30-2), for example DeepL, wherein the server prioritizes using a first translation result created with the large language model (30-1) over a second translation result created with the translation AI model (30-2) if the first translation result is available to the server within a predetermined translation time period, and uses only the second translation result if the predetermined translation time period is exceeded.

8. The method according to claim 6 or 7, wherein the server (2) generates an additional audio data stream with a synthetic voice from the transcript chunks (TCt) translated into a foreign language by using a further Kl model and transmits it to the terminal (3, 3-1, 3-2, 3-3), and wherein the terminal (3, 3-1, 3-2, 3-3) starts a second playback process for playing the additional audio data stream in addition to the first playback process for playing the original audio data stream retrieved from the hitemetradio transmitter, and the first playback process controls a playback position of the second playback process.

9. The method according to one of claims 1 to 8, wherein the server (2) receives and transcribes audio bit data records (ABD) of a plurality of continuous audio data streams or transmitters in parallel and operates parallel transcribing processes for transcribing audio data chunks (AC) using a respective second Kl model, wherein the server (2) further assigns audio data chunks (AC) of a plurality of audio data streams for transcription to a respective ongoing transcribing process, and wherein the server (2) operates separate transcribing processes for different languages.

10. System (2, 3, 3-1, 3-2, 3-3) for real-time transcribing a continuous audio data stream (1) of an Internet radio station into text, by a server (2) of the system, and for playing the audio data stream and time-synchronously displaying the transcribed text on a terminal (3, 3-1, 3-2, 3-3) of the system, wherein the server (2) and the terminal (3, 3-1, 3-2, 3-3) are designed to carry out the method according to one of claims 1 to 9.

11. Server (2) for the system according to claim 10, wherein the server (2) is preferably equipped with a plurality of GPUs for accelerated processing of the generative AI.

12. Terminal (3, 3-1, 3-2, 3-3) for the system according to claim 10.

13. Computer program with executable instructions which, when executed by a server (2), cause the server (2) to carry out a method according to one of claims 1 to 9 and, when executed by a terminal (3, 3-1, 3-2, 3-3), cause the terminal to carry out a method according to one of claims 1, 4 to 6, or 8.

14. A data carrier with a computer program according to claim 13, wherein the data carrier is a computer-readable storage medium, a server accessible via the Internet, an electronic signal, an optical signal and / or a radio signal.

Citation Information

Patent Citations

  • Speech audio pre-processing segmentation

    US11049502B1

  • Computer-implemented method of transcribing an audio stream and transcription mechanism

    US11594227B2

  • Caption and / or Metadata Synchronization for Replay of Previously or Simultaneously Recorded Live Programs

    US20140259084A1

  • Speech recognition with sequence-to-sequence models

    US20200126538A1

  • Audio highlighter

    US20210064327A1

Cited By

  • Asynchronous alignment and accurate dynamic truncation method for audio frame and streaming recognition text

    CN122177119A

  • Asynchronous alignment of audio frames with streaming recognized text and accurate dynamic truncation method

    CN122177119B