Transcribing audio input

The method addresses the challenges of transcribing audio inputs in sections by combining overlapping words from adjacent transcripts, resulting in improved transcription quality and reduced delays in real-time applications.

WO2025120252A1PCT designated stage expired Publication Date: 2025-06-12ELISA OYJ
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/FI2024/050601
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-11-08
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Transcribing audio inputs in sections can introduce technical challenges, particularly in real-time applications where speech is split across sections, leading to incorrect transcription of words.

Method used

A computer-implemented method that obtains sections of an audio input, transcribes each section into a transcript, removes the last word from each transcript, and combines overlapping words from adjacent transcripts to create a cohesive result transcript, improving transcription quality.

Benefits of technology

The method enhances the quality of transcription by accurately combining overlapping words from adjacent sections, reducing delays in real-time transcription, and improving simultaneous interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FI2024050601_12062025_PF_FP_ABST
    Figure FI2024050601_12062025_PF_FP_ABST
Patent Text Reader

Abstract

According to an embodiment, a computer-implemented method (100) for transcribing an audio input, the method comprising: obtaining (101) a first section of an audio input; transcribing (102) the first section of the audio input into a first transcript; removing (103) at least a last word from the first transcript; obtaining (104) a second section of the audio input, wherein the first section overlaps with the second section; transcribing (105) the second section into a second transcript; finding (106) at least one overlapping word in the first transcript and the second transcript; and combining (107) the first transcript and the second transcript into a result transcript based on the at least one overlapping word in the first transcript and the second transcript.
Need to check novelty before this filing date? Find Prior Art

Description

TRANSCRIBING AUDIO INPUTTECHNICAL FIELD

[0001] The present disclosure relates to language processing, and more particularly to a computer-implemented method for transcribing an audio input , a computing device , and a computer program product .BACKGROUND

[0002] In many applications , an audio input needs to be transcribed into text in sections . This may be the case , for example , when the audio input comprises an audio stream and speech in the audio stream needs to be transcribed into text in real time . However, transcribing text in sections can introduce various technical challenges .SUMMARY

[0003] This summary is provided to introduce a selection of concepts in a s implif ied form that are further described below in the detailed description . This summary is not intended to identify key features or essential features of the claimed subj ect matter, nor is it intended to be used to limit the scope of the claimed subj ect matter .

[0004] It is an obj ective to provide a computer-implemented method for transcribing an audio input , a computing device , and a computer program product . The foregoing and other obj ectives are achieved by the features of the independent claims . Further implementation forms are apparent from the dependent claims , the description and the figures .

[0005] According to a first aspect , computer-implemented method for transcribing an audio input comprises : obtaining a first section of an audio input ; transcribing the first section of the audio input into a first transcript ; removing at least a last word from the first transcript ; obtaining a second section of the audio input , wherein the first section overlaps with the second section ; transcribing the second section into a second transcript ; finding at least one overlapping word in the first transcript and the second transcript ; and combining the first transcript and the second transcript into a result transcript based on the at least one overlapping word in the first transcript and the second transcript . The method can, for example , improve the quality of transcription .

[0006] In an implementation form of the first aspect , an overlap of the first section and the second section is in the range 0 . 5 - 5 seconds . The overlap can be sufficient to ensure that the overlap comprise at least one whole word .

[0007] In another implementation form of the first aspect , a length the first section and / or a length ofthe second section i s in the range 1 - 20 seconds . The first section can be of sufficient length to improve transcription .

[0008] In another implementation form of the first aspect , a ratio between a length of the audio input and a length the first section and / or a ratio between a length of the audio input and a length of the second section is greater than or equal to three . The first section can be of sufficient length to improve transcription .

[0009] In another implementation form of the first aspect , the method further comprises , after the transcribing the second section into the second transcript , removing at least a last word from the second transcript . The second transcript can be more easily combined with following transcripts without repeated words .

[0010] In another implementation form of the first aspect , the obtaining the second section of the audio input comprises : obtaining a subsection of the first section ; obtaining a third section of the audio input , wherein the third section is contiguous with the subsection of the first section; and combining the subsection of the first section and the third section, thus obtaining the second section . The second section can be obtained with an appropriate overlap with the first section .

[0011] In another implementation form of the first aspect , the method further comprises , after the obtaining the subsection of the first section, storing thesubsection of the first section in a buffer . The buffer can be used to efficiently store subsections .

[0012] In another implementation form of the first aspect , the method further comprises , after combining the subsection of the first section with the third section, emptying the buffer and storing a subsection of the second section in the buffer .

[0013] In another implementation form of the first aspect , the method further comprises : obtaining a fourth section of the audio input , wherein the fourth section is contiguous with the subsection of the second section ; combining the subsection of the second section and the fourth section, thus obtaining a fifth section ; transcribing the fifth section into a third transcript ; finding at least one overlapping word in the second transcript and the third transcript ; and combining the second transcript and the third transcript into the result transcript based on the at least one overlapping word in the second transcript and the third transcript

[0014] In another implementation form of the first aspect , the method further comprises performing realtime transcription and / or simultaneous interpretation based on the result transcript . The method can improve real-time transcription and / or simultaneous interpretation .

[0015] In another implementation form of the first aspect , the audio input comprises an audio stream and / or an audio stream of a voice call .

[0016] According to a second aspect , a computing device compri ses at least one processor and at least one memory including computer program code , the at least one memory and the computer program code being configured to , with the at least one proces sor, cause the computing device to perform the method according to the first aspect .

[0017] In an implementation form of the second aspect , the at least one memory and the computer program code are further configured to , with the at least one processor, cause the computing device to perform real-time transcription and / or simultaneous interpretation based on the result transcript .

[0018] According to a third aspect , a computer program product comprises program code configured to perform the method according to the first aspect when the computer program product is executed on a computer .

[0019] Many of the attendant features wil l be more readily appreciated as they become better understood by reference to the following detailed description considered in connection with the accompanying drawings .DESCRIPTION OF THE DRAWINGS

[0020] In the following, example embodiments are described in more detail with reference to the attached figures and drawings , in which :

[0021] Fig . 1 illustrates a flow chart representation of a method according to an embodiment ;

[0022] Fig . 2 illustrates a schematic representation of audio input transcription according to a comparative example ;

[0023] Fig . 3 illustrates a schematic representation of audio input transcription according to an embodiment ;

[0024] Fig . 4 illustrates a schematic representation of audio input transcription according to another embodiment ;

[0025] Fig . 5 illustrates a flow chart representing of a method according to an embodiment ; and

[0026] Fig . 6 illustrates a schematic representation of a computing device according to an embodiment .

[0027] In the following, like reference numerals are used to des ignate li ke parts in the accompanying drawings .DETAILED DESCRIPTION

[0028] In the following description, reference is made to the accompanying drawings , which form part of the disclosure , and in which are shown, by way of illustration, specific aspects in which the present disclosure may be placed . It is understood that other aspects may be utilised, and structural or logical changes may be made without departing from the scope of the present disclosure . The following detailed description, therefore , is not to be taken in a limiting sense , as the scope of the present disclosure is defined by the appended claims .

[0029] For instance , it is understood that a disclosure in connection with a described method may also hold true for a corresponding device or system configured to perform the method and vice versa . For example , if a specific method step is described, a corresponding device may include a unit to perform the described method step, even if such unit is not explicitly described or il lustrated in the f igures . On the other hand, for example , if a specific apparatus is described based on functional units , a corresponding method may include a step performing the described functionality, even if such step is not explicitly described or illustrated in the figures . Further, it is understood that the features of the various example aspects described herein may be combined with each other, unless specifically noted otherwise .

[0030] Fig . 1 illustrates a flow chart representation of a method according to an embodiment .

[0031] According to an embodiment , a computer-implemented method 100 for transcribing an audio input comprises obtaining 101 a first section of an audio input .

[0032] The audio input may also be referred to as audio data, audio information, an audio f ile , an audio stream, or similar .

[0033] The audio input may comprise , for example , a batch of audio as an audio file or a continuous audio stream . The audio input may comprise speech .

[0034] The method 100 may further comprise transcribing 102 the first section of the audio input into a first transcript .

[0035] The method 100 may further comprise removing103 at least a last word from the first transcript .

[0036] In some embodiments , the method 100 may comprise removing a last word from the first transcript .

[0037] The method 100 may further comprise obtaining104 a second section of the audio input , wherein the first section overlaps with the second section .

[0038] The second section can be obtained in various ways , such as those disclosed herein .

[0039] Herein, two sections may overlap if the sections have at least one common subsection .

[0040] The method 100 may further comprise transcribing 105 the second section into a second transcript .

[0041] The method 100 may further comprise finding 106 at least one overlapping word in the first transcript and the second transcript .

[0042] The method 100 may further comprise combining 107 the first transcript and the second transcript into a result transcript based on the at least one overlapping word in the first transcript and the second transcript .

[0043] The method 100 may be utilised in, for example , speech-to-text technologies . For example , the method 100 may be implemented in a speech-to-text bot configured to obtain information from users by, for example , phone . The result transcript may then be used for further dataprocessing of the obtained information . Alternatively or additionally, the method may be utili zed in, for example , real-time transcription in contact centre applications , real-time translation, and / or real-time speech interpretation and event generation in, for example , an emergency call .

[0044] At least some embodiments disclosed herein can improve the quality of speech transcription .

[0045] According to an embodiment , the method further comprises performing real-time transcription and / or simultaneous interpretation based on the result transcript .

[0046] At least some embodiments disclosed herein can reduce delays in real-time transcription .

[0047] Herein, real-time transcription may refer to transcribing the audio input while the audio input is still being generated . For example , the audio input can comprise an audio stream comprising speech from a speaker and the speaker can still be speaking while the speech is being transcribed .

[0048] Herein, simultaneous interpretation may refer to transcribing the audio input and translating the transcript while the audio input is sti ll being generated . For example , the audio input can comprise an audio stream comprising speech from a speaker and the speaker can still be speaking whi le the speech is being transcribed and the transcript is being translated . Further, the translation may also be converted to speech using atext-to-speech conversion while the speaker is still speaking .

[0049] According to an embodiment , the audio input comprises an audio stream and / or an audio stream of a voice call .

[0050] Fig . 2 illustrates a schematic representation of audio input transcription according to a comparative example .

[0051] In applications where speech is transcribed in real time , it may be advantageous to start transcribing a part of a speech input before the speaker has stopped speaking . For example , in simultaneous interpretation, a person may speak for 30 seconds . I f a system performing the simultaneous interpretation waits until the person stops speaking before transcribing the speech, there can be a silence lasting for over 30 seconds as the speech is transcribed, the transcription is translated, and speech is generated based on the translation . Instead, the speech can be transcribed in, for example , three- second sections and the sections can be translated, and speech generated . This way, the delay can be reduced to closer to three seconds rather than 30 second . However, this can introduce challenges to the transcription as illustrated in the comparative example of Fig . 2 .

[0052] In the comparative example of Fig . 2 , an audio input 201 comprising eight words is transcribed in three sections . As is illustrated, the words in the audio input 201 can be of varying lengths and the words may not be split cleanly between the three sections . Thevarying length of the words is illustrated as varying spacing of the words in Fig . 2 . In practical situations , the word length variance can be even greater than what is illustrated .

[0053] For example , the third word in the audio input 201 is split between a first section 202 and a second section 203 and the sixth word is split between the second section 203 and a third section 204 . Thus , a first part of the third word is transcribed into a first transcript 205 and a second section of the third word is transcribed into a second transcript 206 . Similarly, a first part of the sixth word i s transcribed into the second transcript 206 and a second section of the sixth word is transcribed into a third transcript 207 . This is indicated by the round brackets in Fig . 2 . For example, the transcript of the first and second part of the third word is referred to as (word3 )

[0001] and (word3 )

[0002] , respectively .

[0054] Since the first and second part of the third and sixth word are transcribed separately, the resulting transcript is incorrect in most cases . When the first 205 , second 206 , and third transcript 207 are combined into a result transcript 208 , the incorrectly transcribed words are carried over .

[0055] The issues illustrated in the comparative example of Fig . 2 can be especially typical in languages with long words , such as many compound words .

[0056] Fig . 3 illustrates a schematic representation of audio input transcription according to an embodiment .

[0057] In the embodiment of Fig . 3 , an audio input 201 comprising eight words is transcribed in three sections . As is illustrated in the audio input 201 of the embodiment of Fig . 3 , the words in the audio input 201 can be of varying lengths and the words may not be spl it cleanly between the three sections .

[0058] According to an embodiment , a length the first section and / or a length of the second section is in the range 1 - 20 seconds ( s ) .

[0059] According to an embodiment , a length the third section is in the range 1 - 20 seconds .

[0060] Alternatively, a length of the first section, a length of the second section, and / or a length of the third section may be in the range 1 - 15 s , 2 - 20 s , 2 - 15 s , 5 - 15 s , or 7 - 13 s .

[0061] A length may also be referred to as a temporal duration, a duration, a temporal length, or similar .

[0062] According to an embodiment , a ratio between a length of the audio input and a length the first section and / or a ratio between a length of the audio input and a length of the second section is greater than or equal to three .

[0063] According to an embodiment , a ratio between a length of the audio input and a length the third section is greater than or equal to three .

[0064] Alternatively, a ratio between a length of the audio input and a length the first section, a ratio between a length of the audio input and a length of the second section , and / or a length of the audio input anda length the third section is greater than or equal to two, four, five, six, seven, eight, nine, or ten.

[0065] A length of the first section, a length of the second section, and / or a length of the third section can be clearly and / or significantly shorter than the length of the audio input.

[0066] In the embodiment of Fig. 3, a first section 302 is transcribed into a first transcript 305 and the last word of the first transcript 305 is removed. A second section 303 is transcribed into a second transcript 306.

[0067] The first section 302 overlaps with the second section 303. Thus, the third word is transcribed into the first transcript 305 and the second transcript 306. The first section 302 comprises a part of the fourth word. Thus, the fourth word is probably transcribed incorrectly in the first transcript 305 as illustrated in Fig. 3. Similarly, the second section 303 comprises a part of the second word. Thus, the second word is probably transcribed incorrectly in the second transcript 306 as illustrated in Fig. 3.

[0068] In the embodiment of Fig. 3, the at least one overlapping word in the first transcript and the second transcript comprises the third word. In other embodiments, the at least one overlapping word may comprise any number of words.

[0069] According to an embodiment, an overlap of the first section and the second section is in the range 0.5 - 5 seconds (s) . Alternatively, the overlap of the firstsection and the second section may be 0.5 - 10 s, 0.5 - 7 s, 0.5 - 4 s, 0.5 - 3 s, 1 - 5 s, 1 - 4 s, or 1 - 3 s .

[0070] The overlap of the first section and the second section may be configured to be sufficient to ensure that at least one full spoken word in the audio input 201 is included in the overlap.

[0071] In some embodiments, if the first transcript 305 and the second transcript 306 comprise a plurality of overlapping words, the first transcript 305 and the second transcript 306 can be combined into the result transcript 310 based on a last overlapping word in the plurality of overlapping words.

[0072] The first transcript 305 and the second transcript 306 can be combined into a result transcript 310 based on the at least one overlapping word in the first transcript 305 and the second transcript 306. The first transcript 305 and the second transcript 306 can be combined, for example, in a such a way that the at least one overlapping word is not repeated in the result transcript 310. For example, the first transcript 305 and the second transcript 306 can be combined by taking words from the start of the first transcript 305 up to and including the at least one overlapping word and taking words from a first word after the at least one overlapping word of the second transcript 306 until the end of the second transcript 306. Alternatively, the first transcript 305 and the second transcript 306 can be combined by taking words from the start of the firsttranscript 305 up to but not including the at least one overlapping word and taking words from the at least one overlapping word of the second transcript 306 until the end of the second transcript 306 .

[0073] For example , in the embodiment of Fig . 3 , the first transcript 305 and the second transcript 306 can be combined into the result transcript 310 by taking wordl , word2 , and word3 from the first transcript 305 and taking word4 , word5 , and word6 from the second transcript 306 or by taking wordl and word2 from the first transcript 305 and taking word3 , word4 , word5 , and word6 from the second transcript 306 .

[0074] According to an embodiment , the method 100 further comprises , after the transcribing the second section into the second transcript , removing at least a last word from the second transcript .

[0075] The removing at least a last word from the second transcript 306 can be performed before combining the first transcript and 305 the second transcript 306 into the result transcript 310 based on the at least one overlapping word in the first transcript 305 and the second transcript 306 .

[0076] For example , in the embodiment of Fig . 3 , word7 is removed from the second transcript 306 .

[0077] A fifth section 304 can be transcribed into a third transcript 307 . The second section 303 overlaps with the fifth section 304 . Thus , the sixth word is transcribed into the second transcript 306 and the third transcript 307 . The second section 303 comprises a partof the seventh word . Thus , the seventh word is probably transcribed incorrectly in the second transcript 306 as illustrated in Fig . 3 . Similarly, the fifth section 304 comprises a part of the fifth word . Thus , the fifth word is probably transcribed incorrectly in the third transcript 307 as illustrated in Fig . 3 .

[0078] By removing at least the last word from the second transcript 306 , the second transcript 306 can be combined with the third transcript 307 in a similar fashion as combining the f irst transcript 305 with the second transcript 306 .

[0079] It should be appreciated that the second transcript 306 can be combined with the third transcript 307 in various ways . For example , in the embodiment of Fig .3 , the first transcript 305 and the second transcript 306 are first combined into the result transcript 310 and the result transcript 310 is then combined with the third transcript 307 . Alternatively, in some other embodiments , the second transcript 306 and the third transcript 307 can be combined separately .

[0080] As can be appreciated from the disclosure , the result transcript 310 can function as a result buffer . As more sections are transcribed, the resulting transcripts can be added to the result transcript . The result transcript can be also referred to as a result buffer, a result transcript buffer, or similar .

[0081] Fig . 4 illustrates a schematic representation of audio input transcription according to another embodiment .

[0082] According to an embodiment , the obtaining the second section of the audio input comprises : obtaining a subsection of the first section ; obtaining a third section of the audio input , wherein the third section is contiguous with the subsection of the first section; and combining the subsection of the first section and the third section, thus obtaining the second section .

[0083] For example , in the embodiment of Fig . 4 , a subsection 401 of the first section 302 is combined with a third section 402 , thus obtaining the second section . The third section 402 is contiguous with the subsection 401 of the first section 302 .

[0084] The subsection 401 of the first section 302 can correspond to the overlap of the first section 302 and the second section 303 . Thus , any disclosure herein in relation to the overlap of the first section 302 and the second section 303 may apply to the subsection 401 of the first section 302 .

[0085] The subsection 401 of the first section 302 should be long enough that the subsection 401 probably comprises at least one whole word of the audio input 201 but not so long that it significantly affects the transcription load .

[0086] According to an embodiment , the method 100 further comprises , after the obtaining the subsection of the f irst section , storing the subsection of the first section in a buffer .

[0087] The buffer 405 can also be referred to as an audio buffer, a subsection buffer, or similar .

[0088] For example , in the embodiment of Fig . 4 , the subsection 401 of the f irst section 302 is stored in a buffer 405 and the combined with the third section 402 , resulting in the second section 303 . The second section 303 can then be processed, for example , in a similar manner to the embodiment of Fig . 3 .

[0089] For example , in the case of the audio input 201 corresponding to an audio stream, the buffer can be used to efficiently store the subsection 401 of the first section 302 in order to combine the subsection 401 with the third section 402 when the third section 402 becomes available .

[0090] According to an embodiment , the method 100 further comprises , after combining the subsection of the first section with the third section, emptying the buffer and storing a subsection of the second section in the buffer .

[0091] For example , in the embodiment of Fig . 4 , the subsection 403 of the second section 303 is stored in a buffer 405 and combined with a fourth section 404 , resulting in the fifth section 304 . The fifth section 304 can then be processed, for example , in a similar manner to the embodiment of Fig . 3 . It should be appreciated that the subsection 403 of the second section 303 can also be referred to as a subsection 403 of the third section 402 .

[0092] According to an embodiment , the method 100 further comprises : obtaining a fourth section of the audio input , wherein the fourth section is contiguous with thesubsection of the second section; combining the subsection of the second section and the fourth section, thus obtaining a fifth section ; transcribing the fifth section into a third transcript ; finding at least one overlapping word in the second transcript and the third transcript ; and combining the second transcript and the third transcript into the result transcript based on the at least one overlapping word in the second transcript and the third transcript .

[0093] For example , in the embodiment of Fig . 4 , a fourth section 404 is contiguous with the subsection 403 of the second section 402 . The subsection 403 of the second section 402 i s stored in the buffer 405 and the subsection 403 and the fourth section 404 are combined, thus obtaining the fifth section 304 . The fifth section 304 is transcribed into the third transcript 307 and the second transcript 306 and the third transcript 307 are combined into the result transcript 310 based on the at least one overlapping word in the second transcript 306 and the third transcript 307 .

[0094] It should be appreciated that the second transcript 306 and the third transcript 307 can be combined into the result transcript 310 in various ways . For example , in the embodiment of Fig . 4 , since the result transcript 310 already comprises the second transcript 306 , the third transcript 307 can be combined with the result transcript 310 and the result can be again stored in the result transcript 310 . Thus , the result transcript 310 can function as a result buffer . In otherembodiments , the combining can be performed in other ways .

[0095] The embodiments of Figs . 3 and 4 can be generali zed for any number of sections . For example , the result transcript 310 can be implemented as a buffer and more sections can be transcribed in the manner disclosed herein and the result can be added to the buffer as long as there are more sections to transcribe . For example , in some embodiments , the audio input 201 can comprise an audio stream . Thus , more sections may be obtained as previous sections are transcribed and added to the result transcript 310 . This may be the case , for example, in simultaneous interpretation, where the result transcript 310 can be translated at the same time as more audio is obtained via the audio stream .

[0096] Fig . 5 illustrates a flow chart representing of a method according to an embodiment .

[0097] The embodiment of Fig . 5 illustrates one example of how any number of sections can be transcribed . The illustrated operations can be looped as long as there are more sections to transcribe .

[0098] In operation 501 , the content of the buffer can be obtained . The buf fer can comprise a subsection of a previously transcribed section as described herein . I f no sections have been transcribed, the buffer can be empty .

[0099] In operation 502 , a section can be obtained . The obtained section can be a section to be transcribednext . This section may be referred to as the current section .

[0100] In operation 503 , the content of the buffer and the section can be transcribed . The result may be referred to as the current transcript .

[0101] In operation 504 , the buffer can be emptied, and a subsection of the current section can be stored in the buffer . This content can then be obtained from the buffer in operation 501 when the next section is transcribed .

[0102] In operation 505 , it can be checked whether a previous transcript is empty . The previous transcript may refer to a result of a previous transcription . For example , if the procedure of Fig . 5 is performed multiple times , the previous transcript may have been obtained during the previous execution of operation 503 . I f the previous transcript is empty, the procedure can move to operation 507 . I f the previous transcript is not empty, the procedure can move to operation 506 .

[0103] In operation 507 , it can be checked whether the current transcript comprises more than one word . I f the current transcript comprises more than one word, the procedure can move to operation 512 . I f the current transcript does not comprise more than one word, the procedure can move to operation 508 .

[0104] In operation 508 , the current transcript can be added to the result transcript .

[0105] In operation 506, the previous transcript and the current transcript can be searched for the at least one overlapping word.

[0106] In operation 509, if there is at least one overlapping word, the procedure can move to operation 510. If there is no overlapping word, the procedure can move to operation 511.

[0107] In operation 510, if the at least one overlapping word is the last word of the previous transcript, the procedure can move to operation 511. If the at least one overlapping word is not the last word of the previous transcript, the procedure can move to operation 512.

[0108] In operation 511, the last word of the previous transcript can be added to the result transcript.

[0109] In operation 512, the current transcript without the last word of the current transcript can be added to the result transcript.

[0110] In operation 513, if the procedure is to be stopped, the procedure can move to operation 514. Otherwise, the procedure can move to operation 501. The procedure can be stopped, for example, when there are no more sections to be transcribed.

[0111] In operation 514, the last word of the last transcript can be added to the result transcript.

[0112] The procedure can be repeated as long as the procedure is not stopped in operation 513.

[0113] Fig. 6 illustrates a schematic representation of a computing device according to an embodiment.

[0114] According to an embodiment, a computing device 600 comprises at least one processor 601 and at least one memory 602 including computer program code. The at least one memory 602 and the computer program code may be configured to, with the at least one processor 601, cause the computing device 600 to perform the method 100.

[0115] The computing device 600 may comprise at least one processor 601. The at least one processor 601 may comprise, for example, one or more of various processing devices, such as a co-processor, a microprocessor, a digital signal processor (DSP) , a processing circuitry with or without an accompanying DSP, or various other processing devices including integrated circuits such as, for example, an application specific integrated circuit (ASIC) , a field programmable gate array (FPGA) , a microprocessor unit (MCU) , a hardware accelerator, a special-purpose computer chip, or the like.

[0116] The computing device 600 may further comprise a memory 602. The memory 602 may be configured to store, for example, computer programs and the like. The memory 602 may comprise one or more volatile memory devices, one or more non-volatile memory devices, and / or a combination of one or more volatile memory devices and nonvolatile memory devices. For example, the memory 602 may be embodied as magnetic storage devices (such as hard disk drives, magnetic tapes, etc.) , optical magnetic storage devices, and semiconductor memories (such asmask ROM, PROM (programmable ROM) , EPROM (erasable PROM) , flash ROM, RAM ( random access memory) , etc . ) .[01 1 7] The computing device 600 may further comprise other components not illustrated in the embodiment of Fig . 6 . The computing device 600 may comprise , for example , an input / output bus for connecting the computing device 600 to other devices . Further, a user may control the computing device 600 via the input / output bus and / or the computing device 600 may obtain the audio input via the input / output bus .

[0118] When the computing device 600 is configured to implement some functionality, some component and / or components of the computing device 600 , such as the at least one processor 601 and / or the memory 602 , may be configured to implement this functionality . Furthermore , when the at least one processor 601 is configured to implement some functionality, this functionality may be implemented using program code comprised, for example , in the memory .

[0119] The computing device 600 may be implemented at least partially using, for example , a computer, some other computing device , or similar .

[0120] According to an embodiment , the at least one memory 602 and the computer program code are further configured to , with the at least one processor 601 , cause the computing device 600 to perform real-time transcription and / or simultaneous interpretation based on the result transcript .

[0121] The method 100 and / or the computing device 600 may be utilised in, for example , automatic speech recognition (ASR) application such as in a so-called voice- bot . A voicebot may be configured to obtain information from users by, for example , phone and convert the voice information into text information using ASR . The method 100 can be used to reduce delays in speech transcription . The voicebot may further be configured to further process , such as classify, the text information . The voicebot can, for example , ask questions about, for example , basic information from a customer in a customer service situation over the phone , obtain the answers us ing ASR and the method 100 , and save the information in a system . Thus , the customer service situation can be made more efficient and user experience can be improved .

[0122] Any range or device value given herein may be extended or altered without losing the effect sought . Also any embodiment may be combined with another embodiment unless explicitly disallowed .

[0123] Although the subj ect matter has been described in language specific to structural features and / or acts , it is to be understood that the subj ect matter defined in the appended claims is not necessarily limited to the specific features or acts described above . Rather, the specific features and acts described above are disclosed as examples of implementing the claims and other equivalent features and acts are intended to be within the scope of the claims .

[0124] It will be understood that the benefits and advantages described above may relate to one embodiment or may relate to several embodiments . The embodiments are not limited to those that solve any or all of the stated problems or those that have any or all of the stated benefits and advantages . It wil l further be understood that reference to ' an ' item may refer to one or more of those items .

[0125] The steps of the methods described herein may be carried out in any suitable order, or simultaneously where appropriate . Additionally, individual blocks may be deleted from any of the methods without departing from the spirit and scope of the subj ect matter described herein . Aspects of any of the embodiments described above may be combined with aspects of any of the other embodiments described to form further embodiments without losing the effect sought .

[0126] The term ' comprising ' is used herein to mean including the method, blocks or elements identified, but that such blocks or elements do not comprise an exclusive list and a method or apparatus may contain additional blocks or elements .

[0127] It will be understood that the above description is given by way of example only and that various modif ications may be made by those ski lled in the art . The above specification, examples and data provide a complete description of the structure and use of exemplary embodiments . Although various embodiments havebeen described above with a certain degree of particularity, or with reference to one or more individual embodiments , those skilled in the art could make numerous alterations to the disclosed embodiments without departing from the spirit or scope of this specification .

Claims

CLAIMS :

1. A computer-implemented method (100) for transcribing an audio input, the method comprising: obtaining (101) a first section of an audio input ; transcribing (102) the first section of the audio input into a first transcript; removing (103) at least a last word from the first transcript; obtaining (104) a second section of the audio input, wherein the first section overlaps with the second section; transcribing (105) the second section into a second transcript; finding (106) at least one overlapping word in the first transcript and the second transcript; and combining (107) the first transcript and the second transcript into a result transcript based on the at least one overlapping word in the first transcript and the second transcript.

2. The computer-implemented method (100) according to claim 1, wherein an overlap of the first section and the second section is in the range 0.5 - 5 seconds .

3. The computer-implemented method (100) according to claim 1 or claim 2, wherein a length thefirst section and / or a length of the second section is in the range 1 - 20 seconds .4 . The computer-implemented method ( 100 ) according to any preceding claim, wherein a ratio between a length of the audio input and a length the first section and / or a ratio between a length of the audio input and a length of the second section is greater than or equal to three .5 . The computer-implemented method ( 100 ) according to any preceding claim, the method further comprising, after the transcribing the second section into the second transcript , removing at least a last word from the second transcript .6 . The computer-implemented method ( 100 ) according to any preceding claim, wherein the obtaining the second section of the audio input comprises : obtaining a subsection of the first section ; obtaining a third section of the audio input, wherein the third section is contiguous with the subsection of the first section ; and combining the subsection of the first section and the third section, thus obtaining the second section .7 . The computer-implemented method ( 100 ) according to claim 6 , the method further comprising, afterthe obtaining the subsection of the first section, storing the subsection of the first section in a buffer .8 . The computer-implemented method ( 100 ) according to claim 7 , wherein the method further comprises , after combining the subsection of the first section with the third section, emptying the buffer and storing a subsection of the second section in the buffer .9 . The computer-implemented method ( 100 ) according to claim 8 , the method further comprising : obtaining a fourth section of the audio input , wherein the fourth section is contiguous with the subsection of the second section ; combining the subsection of the second section and the fourth section, thus obtaining a fifth section ; transcribing the fifth section into a third transcript ; finding at least one overlapping word in the second transcript and the third transcript ; and combining the second transcript and the third transcript into the result transcript based on the at least one overlapping word in the second transcript and the third transcript .10 . The computer-implemented method ( 100 ) according to any preceding claim, the method further comprising performing real-time transcription and / or simultaneous interpretation based on the result transcript .11 . The computer-implemented method ( 100 ) according to any preceding claim, wherein the audio input comprises an audio stream and / or an audio stream of a voice call .12 . A computing device , comprising at least one processor and at least one memory including computer program code , the at least one memory and the computer program code conf igured to , with the at least one processor, cause the computing device to perform the method according to any preceding claim .13 . The computing device according to claim 12 , wherein the at least one memory and the computer program code are further configured to , with the at least one processor, cause the computing device to perform real-time transcription and / or simultaneous interpretation based on the result transcript .14 . A computer program product comprising program code configured to perform the method according to any of claims 1 - 11 when the computer program product is executed on a computer .

Citation Information

Patent Citations

  • Semiautomated relay method and apparatus

    US20180270350A1