A method of speech to text transcription
Patent Information
- Application Number
- GB2024000038
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-09
- Estimated Expiration
- 2044-01-02
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field of the invention The present disclosure relates to a computer-implemented method of speech to text transcription, in particular a method of self-improving speech to text transcription. Background Speech recognition programs are well known in the art, often relying on speech to text transcriptions. However, current speech to text transcription methods implemented in realtime are faced with a trade off between accuracy, and transcription responsiveness or speed. For example, speech to text transcriptions may rely on transcribing large portions of audio data or complete audio data to improve transcription accuracy by accounting for context of the audio data. However, this approach results in significant time lags between receiving a stream of audio data and generating a transcript. As such, these approaches are poorly suited to real-time speech to text transcription applications where quick transcription is required. By contrast, real-time speech to text transcriptions which prioritise fast transcription are often highly inaccurate as the context of the audio data is minimal, meaning the resulting transcriptions are more prone to error. There is therefore a need to enable fast, responsive, and accurate real-time speech to text transcriptions which may be utilised across a host of applications. Summary of the invention Aspects of the invention are as set out in the independent claims and optional features are set out in the dependent claims. Aspects of the invention may be provided in conjunction with each other and features of one aspect may be applied to other aspects. An aspect of the invention relates to a computer-implemented method of speech to text transcription. The method comprises obtaining a stream of audio data, segmenting the stream of audio data into a plurality of audio segments, and obtaining a window comprising at least a first audio segment of the plurality of audio segments. The method further comprises transcribing the window via speech to text transcription to generate a first transcription of at least the first audio segment, and outputting a transcript comprising the first transcription of the first audio segment. The method then comprises updating the window such that the window comprises at least the first audio segment and a second audio segment of the plurality of audio segments, and re-transcribing the window via speech to text transcription to generate at least a second transcription of the first audio segment and a first transcription of the second audio segment. The method then comprises updating the transcript based on the second transcription of the first audio segment and the first transcription of the second audio segment. This method may be advantageous to improve speech to text transcription for applications where quick transcription is required by outputting an initial transcript quickly based on the stream of audio data, and updating the output transcript as the stream of audio data progresses and more context is acquired, thereby improving the accuracy of the transcript through successive transcriptions whilst still providing fast and responsive transcription. The initial speech to text transcription may be performed by the same speech to text model(s) as the re-transcription, wherein the accuracy may be improved based on the changing audio context within the transcription window. This may be advantageous to reduce the required computational resources. However, the skilled person will also understand that different speech to text model(s) may also be used for the initial transcription and re-transcription. The method may further comprise outputting the updated transcript. The audio data of the first audio segment may precede the audio data of the second audio segment in the stream of audio data. For example, the first audio segment may correspond to the first segment of audio data in the stream of audio data. In some examples, the stream of audio data is a live or real-time stream of audio data. In such examples, obtaining the stream of audio data comprises obtaining a real-time stream of audio data, and segmenting the stream of audio data into a plurality of audio segments is performed in real-time. In some examples, the transcript may be output in real-time. This may be advantageous to facilitate responsive speech to text transcription in real-time. The skilled person will understand that, for the purposes of this application, “real-time” need not be limited to mean “instantaneously” or “at the exact actual time during which a process or event occurs.” Instead, “real-time” may be construed to encompass “at a time during the live stream of audio data.” This may be advantageous to facilitate outputting a transcript during the stream of audio data, rather than delaying generation of the transcript, and / or the general method, until after the live audio stream has ended, for example when the maximum context of the audio stream is available. The definition of “real-time” may encompass processing delays, for example but not limited to delays arising from segmenting live audio data for processing. Alternatively, obtaining the audio data may comprise obtaining static audio data. The method may also be advantageous for use with static (e.g., pre-recorded, or complete audio data) for providing fast and accurate transcriptions. The method may further comprise: generating a confidence score associated with the first transcription of the first audio segment; generating a confidence score associated with the second transcription of the first audio segment; identifying a preferred transcription of the first audio segment based on the confidence score associated with the first transcription of the first audio segment and the confidence score associated with the second transcription of the first audio segment; and wherein updating the transcript based on the second transcription of the first audio segment and a first transcription of the second audio segment comprises updating the transcript to comprise (i) the preferred transcription of the first audio segment, and (ii) first transcription of the second audio segment. Identifying the preferred transcription of an audio segment, such as but not limited to the first audio segment, may comprise comparing the confidence score associated with the first transcription of the first audio segment and the confidence score associated with the second transcription of the first audio segment, and identifying which transcription of the first audio segment has the highest confidence score. This may be advantageous to ensure that the transcript is updated only with transcriptions having a higher confidence, thereby ensuring only the most accurate transcriptions of segments are saved to the output transcript. In some examples, the confidence score may be based, at least in part, on the recency of the transcription. Alternatively, weightings may be applied to each confidence score based, at least in part, on the recency of the transcription. For example, wherein most recent transcriptions have a higher confidence score, or are weighted more favourably, to identify the preferred transcription. Additionally, or instead, the confidence score may be based on the confidence score of the speech to text model based on the accuracy of the transcription. The method may further comprise updating the window iteratively such that the window comprises sequential audio segments up to a maximum window size. The method steps of re-transcribing the window and updating the transcript may be repeated for each window iteration. Implementing a maximum window size may be advantageous to ensure that transcription of the window remains quick as the method progresses through the audio stream. For example, updating the window may comprise obtaining a current window using sliding window segmentation, wherein the current window comprises up to a maximum number of audio segments. Once the maximum window size is reached, the method may further comprise updating the window based on sliding window segmentation to exclude the oldest audio segment from the current window, and include the next sequential audio segment of the plurality of audio segments in the current window. This may be advantageous to ensure that transcription of the window remains quick and responsive, for example by avoiding repeatedly requiring repeated transcription of every audio segment. The transcription of older audio segments excluded from the current window may be fixed or “locked” within the output transcript. The method may further comprise, for each window iteration: re-transcribing the window via speech to text transcription to generate a transcription of each audio segment in the current window; outputting a confidence score associated with the transcription of each audio segment in the current window; identifying a preferred transcription of each audio segment based on the confidence score associated with the transcription of each audio segment in the current window and any confidence scores associated with transcriptions of the same audio segment in previous window iterations; and updating the transcript to comprise the preferred transcription of each audio segment. For example, the updated transcript may comprise the preferred transcription of each audio segment in the current window, and the most recent preferred transcription of each audio segment not in the current window. In some examples, the method may further comprise applying weightings to the confidence scores based on the number of transcriptions of the same audio segment (e.g. from of previous window iterations) which are identical and / or similar. Identifying a preferred transcription of each audio segment may then be based on the confidence score and weighting associated with the transcription of each audio segment in the current window, and the confidence scores and weightings associated with transcriptions of the same audio segments from at least one previous window iteration. This may be advantageous to favourably weight transcription variants which are the same as previous transcriptions of the same audio segment, and therefore may be more reliable. The confidence score and weightings may be combined into overall confidence scores. Alternatively, or in addition, applying weightings to the confidence scores may be based on a determination whether the audio segment comprises at least one word from a predetermined set of words. For example, the predetermined set of words may be a custom dictionary or vocabulary set selected based on the expected context of the stream of audio data. In particular, the expected context of the stream of audio data may be determined by the intended application of the speech to text transcription. Purely for illustration, if the speech to text transcription method is intended for use with an automated customer service line, the custom vocabulary set may include related terms that may be used, such as “refund”, “fault”, “return”, etc. By contrast, if the speech to text transcription method is intended for use with a scam detection application for monitoring telecommunication calls, the custom vocabulary set may include related terms that are likely to be used, such as “bank details”, “HMRC”, “PIN number”, etc. Providing a predetermined set of words I custom word set may be advantageous to increase the accuracy of the transcript by preferentially identifying transcription variants comprising expected words based on the context of the stream of audio data. For example, if the method determines that a transcription of an audio segment comprises at least one word from the custom word set, the method may apply a greater weighting or increase the confidence score compared to transcriptions of the same audio segment which do not comprise words from the custom word set. This may be advantageous to favour transcription variants which include expected words based on the context which may be more likely to be accurate. The skilled person will understand that, by contrast, in some examples the method may reduce the weighting or confidence score based on a determination that a transcription of an audio segment comprises at least one word from the custom word set, for example wherein the custom word set comprises common transcription errors based on the context. Common transcription errors based on the context may include words or terms which rhyme or are phonetically similar to words or terms which are commonly used based on the context. Purely for illustration, if the speech to text transcription method is intended for use with an automated customer service line, a second custom vocabulary set may include terms such as “halt” which may be a common transcription error for the word “fault” which is more likely to be used in the context of a customer service line. For example, generating a confidence score associated with the transcription of an audio segment may comprise generating a confidence score associated with a transcription of an audio segment, determining whether the audio segment comprises at least one word from a predetermined set of words (e.g., a custom vocabulary set); and applying a first weighting to the confidence score in the event that the audio segment comprises at least one word from the predetermined set of words, wherein identifying the preferred transcription of each audio segment is based on the confidence score and the weighting associated with the transcription of each audio segment. Alternatively, or in addition, generating a confidence score associated with the transcription of an audio segment may comprise generating a confidence score associated with a transcription of an audio segment, determining whether the audio segment comprises at least one word from a predetermined set of words (e.g., a custom vocabulary set), and applying a second weighting to the confidence score in the event that the audio segment does not comprise at least one word from the predetermined set of words, wherein identifying the preferred transcription of each audio segment is based on the confidence score and the weighting associated with the transcription of each audio segment. The skilled person will also understand that, alternatively to applying a weighting to the confidence score, the confidence score may itself be updated. For example, such that generating a confidence score associated with the transcription of an audio segment may comprise generating a confidence score associated with a transcription of an audio segment; determining whether the audio segment comprises at least one word from a predetermined set of words; and in the event that the audio segment comprises at least one word from the predetermined set of words, updating the confidence score. In some examples, updating the window may comprise obtaining a current window using dynamic sliding window segmentation, wherein the current window comprises a maximum number of audio segments. In some examples, updating the window comprises obtaining a current window using dynamic sliding window segmentation, wherein the maximum number of audio segments of the current window is configured to be variable, for example compared to the maximum number of audio segments of previous window iterations. This may advantageously facilitate a smaller default window size (i.e., a lower maximum number of audio segments in the current window by default) in order to reduce the amount of audio context considered in the transcription of the current window and thereby reduce the processing time and required processing power, wherein the window size is then configured to be increased as needed. Therefore, the method may only use high computational resources at times when it is needed. For example, the maximum number of audio segments in the current window may be based on the confidence score associated with the transcription of at least one audio segment in the preceding window iteration. As such, updating the window may comprise obtaining a current window comprising the maximum number of audio segments, wherein the maximum number of audio segments in the current window is based on the confidence score associated with the transcription of at least one audio segment in the preceding window iteration. For example, the confidence score associated with the transcription of at least one audio segment in the most recent window iteration may be compared to a threshold. This may be advantageous to enable the window size (i.e., the maximum number of audio segments in the current window) to be increased in the event that a confidence score associated with the transcription of at least one audio segment in the preceding window is below (or equal to) a lower threshold, in order to increase the amount of audio context considered in the transcription of the current window. By contrast, this may also be advantageous to reduce the window size in the event that a confidence score associated with the transcription of at least one audio segment in the preceding window is above (or equal to) an upper threshold, in order to reduce the amount of audio context considered in the transcription of the current window and thereby reduce the processing time and required processing power. The window size (i.e. the maximum number of audio segments in the current window) may be increased up to an upper limit. Implementing an upper limit may be advantageous to ensure that the processing cost does not become too great as window size increases, and ensures that the transcription remains in real time without significant processing delays. In some examples, the upper and lower confidence thresholds may be the same, however in other examples the upper confidence threshold may be greater than the lower confidence threshold. In some examples, the upper threshold and / or lower threshold may be pre-determined. However, in other examples, the upper threshold and / or lower threshold may be configured to be adjusted based on an indication of a property of at least one audio segment in the current and / or preceding window iteration (or an indication of a property of the stream of audio data as a whole). Example properties may include, but are not limited to, at least one of audio quality, speed of speech, presence of multiple speakers, number of speakers, background noise, transmission method of audio (e.g. phone line, or microphone, etc.), speech and / or speaker characteristics, and contextual factors. For example, the upper threshold and / or lower threshold may be configured to be adjusted based on an indication of audio quality of at least one audio segment in the current and / or preceding window iteration, or the stream of audio data as a whole. This may be advantageous, for example, as if the audio quality is low, the method may reduce the lower threshold as most audio would have a low confidence score. The skilled person will also understand that the upper threshold and / or lower threshold, or the maximum number of audio segments in the current window, may be configured to be adjusted based on a determination that a transcription of an audio segment in the current and / or preceding window iteration comprises at least one word from the custom word set. Alternatively, or in addition, the skilled person will also understand that the upper threshold and / or lower threshold, or the maximum number of audio segments in the current window, may be configured to be adjusted based on the number of transcriptions of the same audio segment (e.g. from of previous window iterations) which are identical and / or similar, for example as discussed above. Alternatively, updating the window may comprise obtaining a current window, wherein the maximum number of audio segments in the current window is configured to be based on an indication of a property of at least one audio segment in the current and / or preceding window iteration. Example properties of an audio segment may include at least one of audio quality, speed of speech, presence of multiple speakers, number of speakers, background noise, transmission method of audio (e.g., phone line, or microphone, etc.), speech and / or speaker characteristics, and contextual factors. For example, the maximum window size may be increased for low quality audio streams. This may be advantageous to improve the accuracy of transcriptions by increasing the context. Identifying the preferred transcription of each audio segment based on the confidence score associated with the transcription of each audio segment in the current window and any confidence scores associated with transcriptions of the same audio segment in previous window iterations may further comprise applying weightings to the confidence scores associated with the transcription of an audio segment based on the total number of audio segments in the window during transcription of that audio segment. For example, a confidence score associated with a transcription of an audio segment transcribed based on a larger window comprising more audio segments may have a larger weighting than a confidence score associated with a transcription of the same audio segment based on a smaller window having fewer audio segments in the window. Outputting the transcript comprising the first transcription of the first audio segment may comprise sending a signal to a display device, such as but not limited to a display screen, to display the transcript. Updating the transcript may further comprise sending a second signal to the display device to display the updated transcript. This may be advantageous to enable the user to view the most accurate and up-to-date version of the transcript. Alternatively, updating the transcript may comprise identifying at least a portion of the second transcription of the first audio segment which is different to a corresponding portion of the first transcription of the first audio segment, determining which of the first transcription of the first audio segment and the second transcription of the first audio segment is most likely to be correct, and replacing the at least one identified portion of the first transcription of the first audio segment with the corresponding portion of the first or second transcription of the first audio segment which was determined to be most likely correct. Determining which of the at least one identified portion of the first transcription and the corresponding identified portion of the second transcription is most likely to be correct may comprise comparing confidence scores for the identified portion and / or transcription of the audio segment. This skilled person will understand that the identified portion may be any of, but is not limited to, a word, phoneme, sentence, or otherwise. The skilled person will understand that whilst the method of updating the transcript is described in relation to comparing transcription versions of the first audio segment, updating the transcript may be based on comparing transcription versions of any audio segment within the stream of audio data. Optionally, the confidence scores may additionally be based on the recency of the transcription of the audio segment. In some examples, the method may further comprise applying post-processing to the transcription of the audio segment(s) or transcript. Example post-processing may include, but is not limited to, at least one of adding or correcting punctuation and / or grammar. Alternatively or in addition, example post-processing may include, but is not limited to, anonymisation or pseudo-anonymisation. For example, the method may further comprise applying a first post-processing step to the transcript, the first post-processing step being configured to identify at least one type of alphanumeric data. The method then replaces in the transcription each instance of the at least one type of alphanumeric data with an indication of the type of alphanumeric data. This may be performed using natural language processing. For example, types of alphanumeric data may include, but are not limited to, names, addresses, date of birth, bank details, including account number and sort code, any type of personally identifiable information (Pll), and / or any type of sensitive personal information (SPI). Purely for illustrative purposes, “My bank details are 12345678 and 01-01-01” may be replaced with “My bank details are [Type I data] and [Type II data]”, wherein Type I data signifies a bank account number, and Type II data signifies a sort code. The first processing step may be performed before outputting the transcript to remove personally identifiable information and address privacy concerns, for example by anonymising or pseudo-anonymising the transcription. The first processing step may also be performed before comparing the transcription versions of an audio segment for updating the transcript in order to reduce the amount of data needed to be processed, thus reducing bandwidth and / or overhead processing power, and increasing efficiency of the data. In some examples, the stream of audio data may comprise a stream of call data. For example, the stream of audio data may comprise a stream of audio call data relating to a telephone call to or from a user equipment device. However, the skilled person will understand that this example is intended to be in no way limiting and that other audio data may be used. In another aspect of the invention, there is provided a computer program product comprising instructions configured to program a programmable device to perform the method of any preceding aspect of the invention. In another aspect of the invention, there is provided a remote server configured to perform a method comprising (i) obtaining a stream of audio data, (ii) segmenting the stream of audio data into a plurality of audio segments, (iii) obtaining a window comprising at least a first audio segment of the plurality of audio segments, (iv) transcribing the window via speech to text transcription to generate a first transcription of at least the first audio segment, (v) outputting a transcript comprising the first transcription of the first audio segment, (vi) updating the window such that the window comprises at least the first audio segment and a second audio segment of the plurality of audio segments, (vii) re-transcribing the window via speech to text transcription to generate at least a second transcription of the first audio segment and a first transcription of the second audio segment, and (viii) updating the transcript based on the second transcription of the first audio segment and the first transcription of the second audio segment. Outputting the transcript and updating the transcript may comprise the server being configured to send a signal to a remote device, such as a user equipment device, wherein the signal is configured to output an indication of the transcript to the remote device. In another aspect of the invention, there is provided a system comprising a user equipment device, such as a smartphone, and a remote server. The user equipment device is configured to send a stream of audio data to the remote server. Optionally, the stream of audio data may be captured by the user equipment device, for example wherein the user equipment device comprises an audio capture device, such as a microphone, and wherein the stream of audio data is at least partially captured by the audio capture device. The remote server is configured to obtain the stream of audio data from the user equipment device, and perform the method as described in preceding aspects of the invention. In particular, the remote server is configured to (i) segment the stream of audio data into a plurality of audio segments, (ii) obtaining a window comprising at least a first audio segment of the plurality of audio segments, (iii) transcribe the window via speech to text transcription to generate a first transcription of at least the first audio segment, (iv) output a transcript comprising the first transcription of the first audio segment, (v) update the window such that the window comprises at least the first audio segment and a second audio segment of the plurality of audio segments, (vi) re-transcribe the window via speech to text transcription to generate at least a second transcription of the first audio segment and a first transcription of the second audio segment, and (vii) update the transcript based on the second transcription of the first audio segment and the first transcription of the second audio segment. In particular, outputting the transcript and outputting the updated transcript may comprise the remote server being configured to send a signal to the user equipment device, wherein the signal is configured to display the outputted transcript on a display device of the user equipment device, such as a screen. In some examples, the stream of audio data may comprise a stream of call data, for example wherein the stream of audio data comprises a stream of audio call data relating to a telephone call to or from the user equipment device. However, the skilled person will understand that this example is intended to be in no way limiting and that other audio data may be used. Drawinqs Embodiments of the disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which: Fig. 1 shows a flow diagram illustrating an example method of speech to text transcription according to the present invention. Fig. 2 shows a schematic illustrating a method of obtaining a sliding window of audio data for speech to text transcription, for example according to an embodiment of the method of Fig. 1, in particular using a sliding window of a fixed size. Fig. 3 shows a schematic illustrating another method of obtaining a sliding window of audio data for speech to text transcription, for example according to an embodiment of the method of Fig. 1, in particular using a dynamic sliding window. Fig. 4 shows a schematic illustrating a method of speech to text transcription, for example according to an embodiment of the method of Figs. 1 and 2. Fig. 5 shows a schematic illustrating an example computing device configured to perform the method of the present invention, such as the method of Fig. 1. Fig. 6 shows a schematic illustrating an alternative example computing system configured to perform the method of the present invention, such as the method of Fig. 1. Specific description Embodiments of the claims relate to a method of speech to text transcription, in particular a method of self-improving speech to text transcription. It will be appreciated from the discussion above that the embodiments shown in the Figures are merely exemplary, and include features which may be generalised, removed, or replaced as described herein and as set out in the claims. Fig. 1 shows a flow diagram illustrating an example method 100 of speech to text transcription according to the present invention. A first embodiment of the method according to Fig. 1 is discussed below with reference to the schematics shown in Fig. 2 and Fig. 4. Fig. 5 shows an example computing device 500 configured to perform the method 100 of Fig. 1. The computing device 500 comprises a processor 502 and a memory 504. The computing device 500 is also coupled to a display device 520. Firstly, the processor 502 obtains a stream of audio data 202 (102). Referring to the computing device 500 of Fig. 5A, the processor 502 may obtain the stream of audio data from memory storage 504. Alternatively, the processor 502 may obtain the stream of audio data from an audio capture device, such as a microphone 506, or an external audio capture device. The processor 502 then segments the stream of audio data 202 into a plurality of audio segments 204 (104). This may be performed on either of real-time audio or static (e.g., pre-recorded) audio. The processor 502 then obtains a window 206 comprising at least a first audio segment 204A. In this example, the first audio segment 204A corresponds to the first audio segment in the stream of audio data 202. As shown in Fig. 4, the initial window 206 (also depicted as window 206.1) is then transcribed via speech to text transcription (108). A transcript is output comprising the transcription 404A.1 of the first audio segment 204A (110) and stored by the memory storage 504. The transcript is also then output to a display device 520, wherein the display device 520 is configured to display the transcript to a user. The window 206 is configured as a sliding window such that it is then updated (see window 206.2) to include the next audio segment (112), in this case a second audio segment 204B, such that the updated window 206.2 comprises both the first audio segment 204A and the second audio segment 204B. The updated window 206.2 is then re-transcribed via speech to text transcription (114). This comprises generating a second transcription 404A.2 of the first audio segment 204A, and a first transcription 404B.1 of the second audio segment 204B. In this example, the speech to text model used in both the initial transcription 108 and the re-transcription 114 is the same. The transcript saved to the memory 504 and output to the display device 520 is then updated (116), based on the re-transcription of the window 206.2. The sliding window 206 is configured to continuously slide and update to include new audio segments from the stream of audio data 202. This may be advantageous as the sliding window is configured to provide an overlap of audio data between successive transcriptions, which allows audio segments to be transcribed with different audio contexts. This is illustrated in Fig. 2 as different sliding window iterations (206.1, 206.2, 206.3, .... 206.7) are shown as time progresses. As such, method steps 114, and 116 may also be repeated for each window iteration across the duration of the stream of audio data. In the example shown in Fig. 2, the sliding window 206 has a fixed maximum window size. As such, the window 206 continuously incorporates additional audio segments up to a maximum capacity. Once the maximum capacity is reached, the sliding window 206 continues to slide and update by discarding or excluding the oldest audio segment from the window 206, whilst including the next most recent audio segment within the window 206, as shown in window iterations 206.5 onwards. In the schematic shown in Figs. 2 and 4, the maximum window size is shown as four audio segments 204, however the skilled person will understand that this is merely for illustration and is not intended to be limiting. The transcription of each audio segment 204 has a corresponding confidence score. For repeat transcriptions, the associated confidence score is compared with other re-transcriptions of the same audio segment. In the example shown in Fig. 4, confidence score 406A.1 is associated with the first transcription 404A.1 of the first audio segment 204A, confidence score 406A.2 is associated with the second transcription 404A.2 of the first audio segment 204A, and so on. Similarly, confidence score 406B.1 is associated with the first transcription 404B.1 of the second audio segment 204B, and so on. The transcription variant with the highest confidence score replaces the older transcription which had the previous highest confidence score within a text buffer 408. The outputted transcript is then updated (116) based on the text buffer 408. As such, an updated and self-improving transcript may be displayed to the user in real time via the display device 520. 10 In addition to the confidence score alone, the preferred transcription may also be selected based on weightings, such as weightings based on the number of transcriptions of the same audio segment which are identical. Purely for illustration, a table illustrating example transcriptions of an audio segment is included below. Transcriptio n iteration of audio segment 204A Transcriptio n of audio segment 204A Associate d confidenc e score Associated weighting based on the number of identical transcription s Overall confidenc e score Preferred transcriptio n in text buffer 1st(404A.1) There is a cat inside 0.5 1 / 1 406A.1 = 0.5 There is a cat inside 2nd (404A.2) There is a rat inside 0.6 1 / 2 406A.1 = 0.25 406A.2 = 0.3 There is a rat inside 3rd (404A.3) There is a bat inside 0.4 1. / 3 406A.1 = 0.17 406A.2 = 0.2 There is a rat inside 406A.3= 0.13 4th (404A.4) There is a cat inside 0.5 2 / 4 406A.1 = 0.25 406A.2 = 0.15 406A.3 = 0.1 406A.4 = 0.25 There is a cat inside In the example shown above, the preferred transcription is identified based on comparing an overall score associated with the transcription of the audio segment in the current window, relative to overall scores associated with transcriptions of the same audio segments from previous window iterations. In this example, the overall score comprises the product of the confidence score of the transcription and the weighting based on the number of transcriptions of the same audio segment which are identical. Differing words are indicated in bold in the second column of the table. In this example, the first transcription of audio segment 204A transcribes the audio segment as “There is a cat inside”. This transcription is added to the text buffer, and output to the transcript (110). The second transcription of the audio segment 204A transcribes the audio segment as “There is a rat inside”. This variant of the transcription is assigned a higher confidence score than the first transcription, so this replaces the first transcription within the text buffer and the outputted transcript is updated (116) to “There is a rat inside”. In this example, an overall score is determined which comprises a product of the confidence score and the weighting, wherein the weighting is based on the incidence of the transcription variant. The transcription variant with the highest overall score is identified as the preferred transcription and used to update the outputted transcript. Referring again to the table above, the third transcription of audio segment 204A transcribes the audio segment as “There is a bat inside”. This has a lower confidence score and overall score than the second transcription, so the text buffer is not altered, and the transcript remains as “There is a rat inside”. The fourth transcription transcribes audio segment 204A as “There is a cat inside”. As this is the second incidence of this transcription variant, a greater weighting is applied. As such, the fourth transcription may be chosen as the preferred transcription based on the highest overall score due to the greater weighting, despite having a lower confidence score than the second transcription. As such, the fourth transcription may be used to update the text buffer and the outputted transcript. Fig. 3 shows an alternative embodiment of the method according to Fig. 1. The method may proceed in substantially the same manner as described above in relation to Fig. 2 and Fig. 4, however the sliding window 206 in Fig. 3 has a dynamic window size. In this case, the processor 502 compares the confidence score of each audio segment in the current window to a lower threshold, and optionally an upper threshold. If the confidence score for a given audio segment is below a lower threshold, the processor 502 is configured to increase the window size. This may advantageously allow the system to only utilise larger context when it is necessary for harder to transcribe audio segments. For example, as shown in Fig. 3, the default (or minimum) window size is two audio segments. The audio segment 204C within the third window iteration 206 is highlighted to indicate that the processor 502 has identified that the first transcription of the audio segment 204C has an associated confidence score below a pre-determined lower threshold. As such, the processor 502 is configured to dynamically expand the window size for the next iteration of the window 206, in this case the fourth window iteration. In this example, the window size is increased from two audio segments to a maximum of four audio segments, however the skilled person will understand that this is intended to be purely illustrative and is in no way limiting. Whilst the example shown in Fig. 3 illustrates that the expanded window 206.4 is expanded forwards and backwards relative to the previous window iteration to capture a larger context including both more recent and earlier audio segments, the skilled person will understand that this is intended to be purely illustrative and is in no way limiting. For example, an expanded window may alternatively be configured to expand only forwards, e.g. to capture a larger context of more recent audio data in the audio stream, or alternatively, an expanded window may be configured to expand only backwards to capture a larger context of the earlier audio data from the audio stream. Either way, the transcription of the fourth window iteration therefore accounts for a larger context of audio data. The processor 502 then compares the confidence scores of the audio segments in the most recently iterated window to the lower threshold and the upper threshold.. In this example, the second transcription of the audio segment 204C within the enlarged window 206 has an associated confidence score above a pre-determined upper threshold. In response, the processor 502 is configured to dynamically reduce the window size based on the indication that all confidence scores are above the upper threshold. By reducing the window this ensures that computational resources are not unnecessarily used with easy to transcribe audio segments. The skilled person will understand that the upper threshold may be the same as the lower threshold, or alternatively that the upper threshold may be higher than the lower threshold. The skilled person will also understand that resizing of the window is not limited to binary resizing, and the processor 502 may be configured to adapt the size of the window based on the confidence score, for example wherein the window size is determined by the confidence score or range of confidence scores. Optionally, for example, the size of the window may be inversely proportional to the confidence score. Alternatively, or in addition, the skilled person will also understand that the window size may be configured to be adjusted based on other features of the stream of audio data and / or audio segments, such as but not limited to the current audio quality. For example, the processor 502 may be configured to assess the quality of the stream of audio data and / or audio segments. This may comprise the processor 502 comparing the sample rate of the audio data to a sample rate threshold or range. If the audio quality is determined to be poor (e.g., below a pre-determined sample rate threshold), the processor 502 may expand at least one of the maximum and minimum window sizes to provide more context during transcription. This may be advantageous to facilitate improved transcriptions for low quality audio streams, including for audio streams having a sample rate of less than 16 kHz, for example such as audio streams sampled at 8 kHz. The skilled person will also understand that weightings may also be applied to the confidence scores, as described above with reference to Fig. 2 and Fig. 4, wherein the preferred transcription may be selected based on an overall score based on the confidence score and the weightings. Alternatively, or in addition, to weightings based on the number of transcriptions of the same audio segment which are identical, as described above, weightings may also be applied based on the size of the window, for example such that transcription variants derived from larger windows are more favourably weighted because the transcription accounts for more context and is therefore more likely to be accurate. The skilled person will also understand that the configuration of computing device 500 is merely one example and that other devices or systems may alternatively be used to perform the method of the present invention. Purely as an example, Fig. 6 shows an alternative system configured to perform the method of the present invention, such as method 100 of Fig. 1. The system of Fig. 6 comprises a user equipment device 620, such as a smartphone, and a remote server 600 comprising a processor 502 configured for cloud computing. In this example, the cloudside processor 502 may be configured to obtain the stream of audio data 202 (102) from the user equipment device 620. The obtained stream of audio data 202 may also be stored at a cloud-side memory 504. Optionally, the stream of audio data 202 may be generated by the user equipment device 620, for example wherein the stream of audio data 202 may comprise audio data from a telephone call, or other audio data captured by an audio capture device or microphone of the user equipment device 620. The processor 502 may then proceed to perform method steps 104 to 116 as described above. In particular, outputting the transcript (110) and outputting the updated transcript (116) may comprise sending a signal to the user equipment device 620, wherein the signal is configured to display the outputted transcript on a display device 520, such as a screen, of the user equipment device 620. In the context of the present disclosure other examples and variations of the apparatus and methods described herein will be apparent to a person of skill in the art. It will be appreciated in the context of the foregoing disclosure that the described embodiments are not to be construed as limiting. For example, where ranges are recited these are to be understood as disclosures of the limits of said range and any intermediate values between the two limits. As another example, with reference to the drawings in general, it will be appreciated that schematic functional block diagrams are used to indicate functionality of systems and apparatus described herein. It will be appreciated however that the functionality need not be divided in this way and should not be taken to imply any particular structure of hardware other than that described and claimed below. The function of one or more of the elements shown in the drawings may be further subdivided, and / or distributed throughout apparatus of the disclosure. In some embodiments the function of one or more elements shown in the drawings may be integrated into a single functional unit. In some examples the functionality of the controllers and processing means described herein may be provided by mixed analogue and / or digital processing and / or control functionality. It may comprise any general-purpose processor, which may be configured to perform a method according to any one of those described herein. In some examples the controller may comprise digital logic, such as field programmable gate arrays, FPGA, application specific integrated circuits, ASIC, a digital signal processor, DSP, or by any other appropriate hardware. In some examples, one or more memory elements can store data and / or program instructions used to implement the operations described herein. Embodiments of the disclosure provide computer program products such as tangible, nontransitory storage media comprising program instructions operable to program a processor to perform any one or more of the methods described and / or claimed herein and / or to provide data processing apparatus as described and / or claimed herein. Such a 5 controller may comprise an analogue control circuit which provides at least a part of this control functionality. An embodiment provides an analogue control circuit configured to perform any one or more of the methods described herein. The above embodiments are to be understood as illustrative examples. Further 10 embodiments are envisaged. It is to be understood that any feature described in relation to any one embodiment may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the embodiments, or any combination of any other of the embodiments. Furthermore, equivalents and modifications not described above may also be employed without 15 departing from the scope of the invention, which is defined in the accompanying claims. These claims are to be interpreted with due regard for equivalents.
Claims
1. A computer-implemented method of speech to text transcription, the method comprising:obtaining a stream of audio data;segmenting the stream of audio data into a plurality of audio segments;obtaining a window comprising a first audio segment of the plurality of audio segments;transcribing the window via speech to text transcription to generate a first transcription of the first audio segment;outputting a transcript comprising the first transcription of the first audio segment;updating the window such that the window comprises the first audio segment and a second audio segment of the plurality of audio segments;re-transcribing the window via speech to text transcription to generate a second transcription of the first audio segment and a first transcription of the second audio segment;updating the transcript based on the second transcription of the first audio segment and the first transcription of the second audio segment.
2. The computer-implemented method of claim 1 wherein the audio data of the first audio segment precedes the audio data of the second audio segment in the stream of audio data.
3. The computer-implemented method of any preceding claim wherein the first audio segment corresponds to the first segment of audio data in the stream of audio data.
4. The computer-implemented of any preceding claim further comprising:generating a confidence score associated with the first transcription of the first audio segment;generating a confidence score associated with the second transcription of the first audio segment;identifying a preferred transcription of the first audio segment based on theconfidence score associated with the first transcription of the first audio segment and the confidence score associated with the second transcription of the first audio segment; and wherein updating the transcript based on the second transcription of the first audio segment and a first transcription of the second audio segment comprises updating the transcript to comprise (i) the preferred transcription of the first audio segment, and (ii) first transcription of the second audio segment.
5. The computer-implemented method of any preceding claim comprising updating the window iteratively to comprise sequential audio segments of the plurality of audio segments up to a maximum window size;the method further comprising, for each window iteration:re-transcribing the window via speech to text transcription to generate a transcription of each audio segment in the current window;generating a confidence score associated with the transcription of each audio segment in the current window;identifying a preferred transcription of each audio segment based on the confidence score associated with the transcription of each audio segment in the current window and any confidence scores associated with transcriptions of the same audio segment in previous window iterations; andupdating the transcript to comprise the preferred transcription of each audio segment.
6. The computer-implemented method of claim 5 wherein once the maximum window size is reached, the method further comprises updating the window based on sliding window segmentation to exclude the oldest audio segment from the current window, and include the next sequential audio segment of the plurality of audio segments the current window;the method further comprising, for each window iteration:re-transcribing the window via speech to text transcription to generate a transcription of each audio segment in the current window;outputting a confidence score associated with the transcription of each audio segment in the current window;identifying a preferred transcription of each audio segment based on the confidence score associated with the transcription of each audio segment in the current window and any confidence scores associated with transcriptions of the same audio segment of previous window iterations; andupdating the transcript to comprise the preferred transcription of each audio segment in the current window and the most recent preferred transcription of each audio segment not in the current window.
7. The computer-implemented method of any claims 4 to 6 wherein identifying the preferred transcription of the first audio segment comprises:comparing the confidence score associated with the first transcription of the first audio segment and the confidence score associated with the second transcription of the first audio segment; andidentifying which transcription of the first audio segment has the highest confidence score.
8. The computer-implemented method of any of claims 4 to 7 wherein generating the confidence score associated with the transcription of an audio segment comprises applying a weighting to the confidence score based on the number of transcriptions of the same audio segment of previous window iterations which are identical.
9. The computer-implemented method of any of claims 4 to 7 generating the confidence score associated with the transcription of an audio segment comprises applying a weighting to the confidence score based on determining whether the audio segment comprises at least one word from a predetermined set of words.
10. The computer-implemented method of any preceding claim wherein updating the window comprises obtaining a current window using sliding window segmentation, wherein the current window comprises up to a maximum number of audio segments.
11. The computer-implemented method of any of claims 5 to 10 wherein updating the window comprises obtaining a current window using dynamic sliding windowsegmentation, wherein the current window comprises up to a maximum number of audio segments, and wherein the maximum number of audio segments of the current window is configured to be variable relative to the maximum number of audio segments of previous window iterations.
12. The computer-implemented method of claim 11 wherein the maximum number of audio segments in the current window is based on the confidence score associated with the transcription of at least one audio segment in the preceding window iteration.
13. The computer-implemented method of claim 12 wherein the maximum number of audio segments in a current window is increased in the event that the confidence score associated with the transcription of at least one audio segment in the preceding window iteration equal to or below a lower threshold.
14. The computer-implemented method of claim 12 or claim 13 wherein the maximum number of audio segments in a current window is reduced in the event that the confidence score associated with the transcription of at least one audio segment in the preceding window iteration is equal to or above an upper threshold.
15. The computer-implemented method of any of claims 10 to 14 wherein the maximum number of audio segments in a current window is configured to be adjusted based on an indication of a property of (i) at least one audio segment in the current and / or preceding window iteration, or (ii) the stream of audio data; optionally wherein the indication of a property comprises an indication of audio quality.
16. The computer-implemented method of any of claims 10 to 15 wherein generating the confidence score associated with the transcription of an audio segment further comprises applying a weighting to the confidence score based on the total number of audio segments in the window during transcription of that audio segment, for example wherein a confidence score associated with a transcription of an audio segment transcribed based on a larger window comprising more audio segments will have a larger weighting than a confidence score associated with a transcription of the same audio segment based on asmaller window having fewer audio segments in the window.
17. The computer-implemented method of any preceding claim wherein:obtaining the stream of audio data comprises obtaining a real-time stream of audio data; andsegmenting the stream of audio data into a plurality of audio segments is performed in real-time.
18. The computer-implemented method of any of claims 1 to 16 wherein obtaining the audio data comprises obtaining static audio data.
19. The computer-implemented method of any preceding claim wherein outputting the transcript comprising the first transcription of the first audio segment further comprises sending a signal to a display device to display the transcript; andwherein updating the transcript further comprises sending a signal to the display device to display the updated transcript.
20. The computer-implemented method of any preceding claim wherein updating the transcript based on the second transcription of the first audio segment and the first transcription of the second audio segment comprises:comparing the first transcription of the first audio segment and the second transcription of the first audio segment to identify at least a portion of the first transcription which is different to a corresponding portion of the second transcription;determining which of: (i) the portion of the first transcription of the first audio segment and (ii) the portion of the second transcription of the first audio segment, is most likely to be correct; andreplacing the at least one identified portion of the first transcription of the first audio segment with the corresponding portion of the first or second transcription of the first audio segment which was determined to be most likely correct.
21. The computer-implemented method of claim 20, wherein determining which of theat least one identified portion of the first transcript and the corresponding identified portion of the second transcript is most likely to be correct comprises applying confidence scores to each identified portion.5 22. A computer program product comprising instructions configured to program a programmable device to perform the method of any preceding claim.
Citation Information
Patent Citations
Methods and systems for transcription
US20200286485A1
Computer-implemented method of transcribing an audio stream and transcription mechanism
US20210407515A1
Word replacement in transcriptions
US20220059075A1
Hypothesis stitcher for speech recognition of long-form audio
US20230154468A1
ViewUS20200286485A1onEspacenetopensinnewtab