Streaming speech-to-speech model with automatic speaker turn detection
The speech-to-speech model with a turn detector automatically detects speech breakpoints, addressing the inconvenience and error of manual input in conventional models, enhancing user experience and intelligibility for atypical speech.
Patent Information
- Application Number
- JP2024567521
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-03
- Filing Date
- 2023-05-17
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2043-05-17
AI Technical Summary
Conventional speech-to-speech models require manual user input to indicate the start and end of an utterance, leading to inconvenience and potential human error, especially for users with atypical speech patterns.
A speech-to-speech model with an integrated turn detector that automatically determines speech breakpoints, enabling streaming voice conversion without manual user input, using a deep neural network to analyze high-order feature representations and generate synthetic speech at appropriate intervals.
Provides a more natural user experience by mimicking normal conversation, improving intelligibility and reducing errors in speech conversion for users with atypical speech patterns or language translation.
Smart Images

Figure 2025515870000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a speech-to-speech model with a turn detector. [Background technology]
[0002] Speech-to-speech (S2S) models can be used to convert source speaker speech into synthetic speech without changing the linguistic information of the original speech. For example, S2S models can be used to generate canonical fluent synthetic speech for users with dysarthria or atypical speech. Alternatively, S2S models can be used to translate a user's speech into synthetic speech in another language. Typically, S2S models are manually triggered by user inputs that indicate when the input speech starts and ends. After the user's speaking ends, the S2S model then processes the input speech to generate synthetic speech. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] US Patent Application Publication No. 2020 / 0226327 Summary of the Invention [Problem to be solved by the invention]
[0004] There is an opportunity to provide a streaming voice-to-voice model with improved automatic speaker turn detection. [Means for solving the problem]
[0005] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations comprising receiving as input to a speech-to-speech (S2S) model a sequence of acoustic frames corresponding to spoken utterances made by a user in streaming audio captured by a client device associated with the user. At each of the plurality of output steps, the operations also include generating, by an audio encoder of the speech-to-speech S2S model, a high-order feature representation of a corresponding audio frame in the sequence of audio frames, determining, by a turn detector of the speech-to-speech S2S model, at the corresponding output step, whether an utterance is at a breakpoint at the corresponding output step based on the high-order feature representation generated by the audio encoder, and synthesizing, when the turn detector determines that the utterance is at a breakpoint, the sequence of output audio frames output by the speech decoder of the speech-to-speech S2S model into a time-domain audio waveform of synthetic speech representing the spoken utterance uttered by the user, where each output audio frame of the sequence of output audio frames is based on a corresponding one of the high-order feature representations generated by the audio encoder until the corresponding output step at which the turn detector determines that the utterance is at a breakpoint.
[0006] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, the operations further include providing, as output from the client device when the speech is at a breakpoint, a speaker turn indication informing the user to stop speaking and a time-domain audio waveform of a synthetic speech representative of the speech produced by the user. The speech produced by the user in the streaming audio captured by the client device may be associated with atypical speech, and the time-domain audio waveform of the synthetic speech representative of the speech may include a time-domain audio waveform of a synthesized canonical fluent speech of the same speech produced by the user. Additionally, the speech produced by the user in the streaming audio captured by the client device may be in a first language, and the time-domain audio waveform of the synthetic speech representative of the speech may include a time-domain audio waveform of a synthesized translation of the same speech in a second language different from the first language.
[0007] In some examples, determining whether the utterance is at a breakpoint at the corresponding output step is further based on one or more of the high-level feature representations generated by the audio encoder at the output step preceding the corresponding output step. Additionally or alternatively, the operations may further include generating, at each of the plurality of output steps, an output audio frame of the corresponding high-level feature representation generated by the audio encoder at the corresponding output step by a speech decoder of the speech-to-speech S2S model.
[0008] In some implementations, the operations also include, in response to determining that the utterance is at a break point of the corresponding output step, receiving the sequence of high-level feature representations generated by the audio encoder as input to a speech decoder of the speech-to-speech S2S model up to the corresponding output step, and generating, by the speech decoder of the speech-to-speech S2S model, a sequence of output audio frames. The turn detector may include a deep neural network that may be disposed between the audio encoder and the speech decoder.
[0009] In some additional implementations, determining whether the utterance is at a break point in a corresponding output step comprises generating a turn output as an output from the turn detector indicating whether the utterance is at a break point, where the turn output comprises a bit or a probability distribution.
[0010] In some examples, the operations further include training the speech-to-speech S2S model by receiving a set of training utterances, where each training utterance in the set of training utterances comprises a corresponding sequence of training acoustic frames, and each training utterance is paired with a corresponding ground truth synthetic speech representation of the training utterance. Each training acoustic frame in the sequence of training acoustic frames is annotated with a label indicating whether the corresponding training acoustic frame corresponds to a breakpoint frame or a non-breakpoint frame. In these examples, the speech-to-speech S2S model is further trained by obtaining a first label of the training input audio data indicative of a target output spectrogram, obtaining a second label of the training input audio data indicative of a target turn output, and generating a training output using the speech conversion model and the training input audio data. The training output comprises a training output spectrogram corresponding to the synthetic speech representation of the training input audio data and a training turn output indicative of the breakpoints of the training input audio data. Finally, the speech-to-speech S2S model is trained by determining a first loss by comparing the training output spectrogram with a first label, determining a second loss by comparing the training turn output with a second label, and optimizing a speech conversion model based on the first loss and the second loss associated with the training input audio data.
[0011] In some implementations, the operations also include determining a voice type of speech produced by the user captured in the streaming audio by the client device and selecting a voice decoder from among a plurality of available voice decoders to generate a sequence of output audio frames, where each output audio frame in the sequence of output audio frames output by the voice decoder comprises a spectrogram frame.
[0012] Another aspect of the disclosure provides a system including data processing hardware and memory hardware in communication with the data processing hardware. The memory stores instructions that, when executed in the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving, as an input to a speech-to-speech (S2S) model, a sequence of acoustic frames corresponding to speech uttered by a user in streaming audio captured by a client device associated with the user. At each of a plurality of output steps, the operations also include generating, by an audio encoder of the speech-to-speech S2S model, a high-order feature representation of a corresponding acoustic frame in the sequence of acoustic frames; determining, by a turn detector of the speech-to-speech S2S model, at the corresponding output step based on the high-order feature representation generated by the audio encoder, whether the speech is at a break point at the corresponding output step; and synthesizing, when the turn detector determines that the speech is at a break point, the sequence of output audio frames output by the speech decoder of the speech-to-speech S2S model into a time-domain audio waveform of synthetic speech representing the speech uttered by the user. Here, each output audio frame in the sequence of output audio frames is based on a corresponding one of the high-level feature representations generated by the audio encoder up to the corresponding output step when the turn detector determines that the speech is at a break point.
[0013] This aspect may include one or more of the following optional features. In some implementations, the operations further include providing as output from the client device when the speech is at a break point a speaker turn indication informing the user to stop speaking and a time-domain audio waveform of the synthetic speech representative of the speech produced by the user. The speech produced by the user in the streaming audio captured by the client device may be associated with atypical speech. The time-domain audio waveform of the synthetic speech representative of the speech may include a time-domain audio waveform of a synthesized standard fluent speech of the same speech produced by the user. Additionally, the speech produced by the user in the streaming audio captured by the client device may be in a first language. The time-domain audio waveform of the synthetic speech representative of the speech may include a time-domain audio waveform of a synthesized translation of the same speech in a second language different from the first language.
[0014] In some examples, determining whether the utterance is at a breakpoint at the corresponding output step is further based on one or more of the high-level feature representations generated by the audio encoder at an output step preceding the corresponding output step. Additionally or alternatively, the operations may further include generating, at each of the multiple output steps, an output audio frame of the corresponding high-level feature representation generated by the audio encoder at the corresponding output step, by a speech decoder of the speech-to-speech S2S model.
[0015] In some implementations, the operations also include, in response to determining that the utterance is at a break point of the corresponding output step, receiving the sequence of high-level feature representations generated by the audio encoder as input to a speech decoder of the speech-to-speech S2S model up to the corresponding output step, and generating, by the speech decoder of the speech-to-speech S2S model, a sequence of output audio frames. The turn detector may include a deep neural network that may be disposed between the audio encoder and the speech decoder.
[0016] In some additional implementations, determining whether the utterance is at a break point in a corresponding output step comprises generating a turn output as an output from the turn detector indicating whether the utterance is at a break point, where the turn output comprises a bit or a probability distribution.
[0017] In some examples, the operations further include training the speech-to-speech S2S model by receiving a set of training utterances, whereby each training utterance in the set of training utterances comprises a corresponding sequence of training acoustic frames, and each training utterance is paired with a corresponding ground truth synthetic speech representation of the training utterance. Each training acoustic frame in the sequence of training acoustic frames is annotated with a label indicating whether the corresponding training acoustic frame corresponds to a breakpoint frame or a non-breakpoint frame. In these examples, the speech-to-speech S2S model is further trained by obtaining a first label of the training input audio data indicative of a target output spectrogram, obtaining a second label of the training input audio data indicative of a target turn output, and generating a training output using the speech conversion model and the training input audio data. The training output comprises a training output spectrogram corresponding to the synthetic speech representation of the training input audio data, and a training turn output indicative of the breakpoints of the training input audio data. Finally, the speech-to-speech S2S model is trained by determining a first loss by comparing the training output spectrogram with a first label, determining a second loss by comparing the training turn output with a second label, and optimizing the speech conversion model based on the first loss and the second loss associated with the training input audio data.
[0018] In some implementations, the operations also include determining a voice type of speech produced by the user captured in the streaming audio by the client device and selecting a voice decoder from among a plurality of available voice decoders to generate a sequence of output audio frames, where each output audio frame in the sequence of output audio frames output by the voice decoder comprises a spectrogram frame.
[0019] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief description of the drawings]
[0020] [Figure 1] FIG. 1 is a schematic diagram of an exemplary voice-to-voice system with a turn detector. [Diagram 2] FIG. 2 is a schematic diagram of an exemplary speech-to-speech model with a turn detector. [Diagram 3] FIG. 2 is a schematic diagram of an example training process for a speech-to-speech model with a turn detector. [Figure 4] 4 is a flow chart of an exemplary operational arrangement of a method for performing voice-to-voice conversion with a turn detector. [Diagram 5] FIG. 1 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0021] Like reference symbols in the various drawings indicate like elements. There is growing interest in developing more inclusive speech technologies, especially those that are enabled to assist people with speech disorders. Speech-to-speech (S2S) conversion has made great advances with the introduction of end-to-end (E2E) deep learning-based models for converting speech of speakers with dysarthria or atypical speech patterns into synthesized speech. For example, atypical speech patterns may include, but are not limited to, speech disorders resulting from physical or neurological conditions (e.g., speakers with amyotrophic lateral sclerosis (ALS) disease), heavily accented speech, and deaf speech. Another application of speech-to-speech S2S models is translation, where a user utters words in a first language while a speech-to-speech S2S model generates synthesized speech in a second language.
[0022] A conventional voice-to-voice S2S model needs to receive the entire input speech (i.e., utterance) before generating synthetic speech. Typically, a user provides inputs indicating the start and end of an utterance to the voice-to-voice S2S model. For example, a user may hold down a button in a user interface corresponding to the voice-to-voice S2S model when the user starts speaking (i.e., starts providing input speech to the voice-to-voice S2S model). The user then releases the button when the user finishes speaking. The voice-to-voice S2S model then converts the received input speech while the user is holding down the button into synthetic speech. Such manual methods for activating a voice-to-voice S2S model can be inconvenient for users because they require a large number of user inputs. Also, requiring user inputs can lead to human error (e.g., the user may release the button before completing the utterance).
[0023] The present disclosure introduces a turn detector that automatically determines appropriate break points in a user's speech to output a synthesized speech converted from the user's speech by a speech-to-speech S2S model in streaming format at the break points. Thus, instead of requiring a user input indicating when speech begins / ends, the present disclosure is enabled to receive streaming speech from the user and automatically determine appropriate moments in the speech, so-called "break points," for the user to pause to receive / listen to the synthesized speech generated by the speech-to-speech S2S model. By automatically determining break points, embodiments of the present disclosure provide a more natural user experience since the turn-by-turn nature of the system mimics normal conversation (i.e., the user speaks when it is their turn, while the system speaks when it is its turn).
[0024] As used herein, unless otherwise specified, the terms "speech-to-speech system" and "speech-to-speech model" may refer to any system / model that directly reconverts input speech to synthetic speech without performing intermediate speech recognition on the input speech. In other words, a speech-to-speech system / model is configured to directly convert an input audio waveform, sequence of acoustic frames, or spectrogram corresponding to the input speech, to an output audio waveform or spectrogram corresponding to the synthetic speech, without converting to an intermediate representation (e.g., text or phonemes). As will become apparent, speech-to-speech models and techniques for training speech-to-speech models enable users of atypical voices to speak to and be understood by both other humans and speech interfaces (e.g., digital assistants) by enabling recognition and / or playback of the user's intended speech.
[0025] Although the examples herein show a speech-to-speech model that receives an input utterance corresponding to an atypical speech and converts it into a synthetic speech corresponding to a standard fluent speech, the speech-to-speech model can be similarly adapted to perform other types of speech conversion tasks without departing from the scope of the present disclosure. For example, a speech-to-speech S2S model can be enabled to convert an input utterance in a first language into a synthetic speech corresponding to a translation of the input utterance in a different second language. A speech-to-speech S2S model can similarly receive an input spoken by a user and output a synthetic speech having the same linguistic content as the spoken input but different speech characteristics of a target speaker.
[0026] 1 illustrates a voice-to-voice system 100 that includes a voice-to-voice (S2S) model 200. The voice-to-voice (S2S) model 200 is configured to directly convert input audio data 102 (e.g., a sequence of acoustic frames or an input audio waveform) corresponding to an utterance 108 produced by a target speaker 104 into output audio data 106 (e.g., a sequence of output audio frames or an output audio waveform) corresponding to a synthetic speech representation of the same utterance 114 produced by the target speaker 104. Notably, the voice-to-voice S2S conversion model 200 is configured to directly convert the input audio data 102 into the output audio data 106 without performing speech recognition or otherwise requiring the generation of any intermediate discrete representation (e.g., text or phonemes) from the input audio data 102.
[0027] The speech-to-speech S2S conversion model 200 comprises an audio encoder 212 configured to encode the input audio data 102 into a hidden feature representation (e.g., a series of vectors), a turn detector 215 configured to determine breakpoints in the utterance 108 based on the hidden feature representation output by the encoder, and a speech decoder 220 configured to decode the hidden representation into an output audio data 106 corresponding to a synthesized standard fluent speech representation. For example, when the audio encoder 200 receives the input audio data 102 of the utterance 108, the audio encoder 200 is processing 5 frames of audio. The audio encoder 200 is enabled to convert these 5 frames of audio into 10 vectors. The vectors are not transcriptions of the frames of the audio data 102, but are mathematical representations of the frames of the audio data 102. The turn detector 215 is then enabled to determine whether a breakpoint exists among the 10 vectors. If a break point exists, the speech decoder 220 is enabled to generate output audio data 106 corresponding to a synthesized standard fluent speech representation based on the vectors received from the audio encoder 200 for vectors prior to the break point. For example, the turn detector 215 may determine that the tenth vector is a break point, such that the speech decoder 220 is enabled to receive ten vectors representing five frames of audio from the audio encoder 200. The speech decoder 220 is now enabled to generate five frames of output audio data 106 corresponding to a synthesized standard fluent speech representation of the utterance 114 that comprises the intended word or part of a word as the five frames of the input audio data 102, but without the atypical speech dysfluency.
[0028] The voice-to-voice S2S conversion system 100 may further include a synthesizer 275 for synthesizing the output audio data 106 into a time-domain waveform for audible output as the synthesized utterance 114 of the utterance 108. The time-domain audio waveform comprises an audio waveform that defines the amplitude of the audio signal over time. The synthesizer 275 may include a unit selection module or a WaveNet module for synthesizing the output audio data 106 into a time-domain waveform of standard fluent speech. In some implementations, the synthesizer 275 comprises a vocoder network, i.e., a neural vocoder that is separately trained and conditioned on a mel-frequency spectrogram for conversion to a time-domain audio waveform. In additional implementations, the synthesizer 275 comprises a streaming vocoder configured to convert / invert in real time a log-magnitude spectrogram output from the speech decoder 220 as the output audio data 106 into a time-domain audio waveform.
[0029] In some implementations, the target speaker 104 is associated with an atypical voice, and the target speaker 104 speaks with atypical speech patterns that may be difficult to understand. Atypical speech patterns may include, but are not limited to, speech disorders resulting from physical or neurological conditions (e.g., speakers with amyotrophic lateral sclerosis (ALS) disease), heavily accented speech, and deaf speech. In other implementations, the target speaker 104 speaks in a first language, while the speech-to-speech S2S translates the first speech into a second language.
[0030] Thus, the speech-to-speech conversion system 100 is trained to directly convert input audio data 102 corresponding to utterances 108 produced by a target speaker 104 associated with an atypical speech into output audio data 106 corresponding to a synthesized standard fluent speech representation of the same utterance 108. The synthesized standard fluent speech representation provided by the output audio data 106 thus improves the intelligibility of the atypical speech (e.g., heavily accented speech or amyotrophic lateral sclerosis (ALS) speech) produced by the target speaker 104.
[0031] Without departing from the scope of the present disclosure, the voice-to-voice conversion system 100 may be trained to directly convert input audio data 102 corresponding to utterances 108 associated with atypical speech in a first language into output audio data 106 corresponding to a synthesized standard fluent speech representation of the same utterance 108 in the same voice but in a different second language.
[0032] Further, the turn detector 215 may be trained to determine break points in the input audio data 102 in real time as the speech 108 produced by the target speaker 104 is captured in streaming audio by the user device 110 associated with the target speaker 104. In some implementations, the turn detector 215 comprises a deep neural network. The turn detector 215 may be disposed between the audio encoder 212 and the speech decoder 220, such that the turn detector 215 receives an output 213 (e.g., a high-level feature representation of each of the sequence of acoustic frames) from the audio encoder 212 and sends a turn output 216 to the decoder 220. The turn output indicates whether the speech is at a break point. In some implementations, the turn output is a series of bits (e.g., “1” and “0”), where “1” indicates that the speech is at a break point and “0” indicates that it is not at a break point. Here, each bit in the series of bits may indicate a corresponding acoustic frame of the input audio data 102. In other implementations, the turn output is a probability score (i.e., a number between "0" and "1") indicating the likelihood of a breakpoint in the corresponding acoustic frame. For example, if the probability score of the corresponding acoustic frame meets a breakpoint threshold, the acoustic frame indicates a breakpoint. If the turn output 216 indicates that the utterance is at a breakpoint, the speech-to-speech S2S model 200 may provide instructions 117 to the user device 110 to cause the user device 110 to output a turn indication 115 indicating that the user should stop speaking.
[0033] The speech decoder 220 may then receive the output 213 from the encoder 212 during each of the multiple time steps, as well as the turn output 216 from the turn detector 215. In some implementations, the turn detector 215 modifies the output 213 from the encoder 212 such that the output 213 from the encoder 212 indicates a break point.
[0034] A user device (interchangeably referred to as a "computing device") 110 associated with a target speaker 104 may capture the utterance 108 produced by the target speaker 104 in streaming audio and transmit the corresponding input audio data 102 to the voice-to-voice conversion system 100 for conversion into output audio data 106. The voice-to-voice conversion system 100 may then transmit the output audio data 106 corresponding to a synthetic voice representation of the same utterance 114 produced by the target speaker 104 to another computing device 116 associated with a user 118, such that the other computing device 116 audibly outputs the synthetic voice representation of the utterance 108 produced by the target speaker 104. In this example, the target speaker 104 and the user 118 converse with each other through their respective computing devices 110, 116 via telephone calls or other types of voice communication protocols, such as voice over Internet Protocol. Although the target speaker 104 and the other users 118 may be speaking in the same language, the target speaker 104 may have atypical speech due to the disease ALS, making it difficult for the other users 118 to understand the target speaker 104. Thus, while the target speaker 104 is speaking in an atypical speech (e.g., ALS speech) that is difficult to understand, the other users 118 listening to the synthesized standard fluent speech representation have an easier time to understand the utterance 108 intended by the target speaker 104. In other words, the synthesized standard fluent speech representation provides a more consistent cadence that may be easier for another user to understand than the original utterance 108 produced by the target speaker in atypical speech. In particular, the synthesized standard fluent speech representation is in the voice of the target speaker 104.
[0035] In some other examples, the voice-to-voice S2S conversion system 100 instead passes output audio data 106 corresponding to a synthetic speech representation of the utterance produced by the target speaker 104 to an output audio device for audibly outputting the synthetic speech representation in the voice of the target speaker 104 to an audience. For example, the target speaker 104 may be a psychology professor lecturing to a class of students, and the utterance produced by the target speaker 104 comprises medical terminology belonging to a particular specific domain, e.g., psychology. As will become apparent, the voice-to-voice S2S conversion system 100 is trained to learn linguistic diversity derived from the linguistic content present in the training utterances and acoustic diversity associated with a particular type of atypical speech associated with the speaker who produced the target utterance.
[0036] Alternatively, the other computing device 116 may be associated with a downstream automatic speech recognition (ASR) system in which the speech-to-speech conversion system 100 acts as a front end providing output audio data 106 corresponding to the synthetic speech representation as input to the automatic speech recognition ASR system for conversion to recognized text. The recognized text may be presented to another user 118 and / or provided to a natural language understanding (NLU) system for further processing.
[0037] In any of the above examples, if the turn detector 215 determines that the target speaker 104 reaches a pause during speech, the voice-to-voice S2S model 200 may provide instructions 117 to cause the user device 110 associated with the target speaker 104 to output a turn indication 115. For example, the user device 110 of FIG. 1 shows the user device 110 displaying a graphical turn indication 115a as a series of exclamation points to inform the target speaker 104 to pause and to enable the voice-to-voice S2S model 200 to generate output speech data 106 corresponding to the synthesized standard fluent speech representation of the input audio data 102. In an additional example, the user device 110 outputs an audible turn indication 115b (e.g., emits a tone or series of tones) to inform the target speaker 104 to pause. The user device 110 may be enabled to output other types of turn indications 115, such as by vibrating and / or flashing a light to inform the target speaker 104 to pause speaking. In some implementations, the voice-to-voice S2S model 200 provides instructions 117 to cause another device (not shown) associated with the target speaker 104 to output the turn indication 115. In these implementations, a device other than the user device 110 that captured the speech 108 may be enabled to receive the instructions 117 and output the turn indication 115. For example, a smartwatch worn by the user may output the turn indication 115 by vibrating / beeping to inform the target speaker 104 to pause. When instructions 117 are provided to cause the user device 110 (or another device) to output the turn indication 115, the synthesizer 275 may provide a time-domain audio waveform of the synthetic speech 114 for output from the device 110, 116, or from any other device.
[0038] The functionality of the voice-to-voice conversion system 100 is enabled to reside on either or both of the remote server 112, the computing devices 110, 116, or any combination of the remote server and the computing devices 110, 116. The computing devices 110 and 116 may include, but are not limited to, smartphones, tablets, desktop / laptop computers, smart speakers, smart displays, smart appliances, assistant-enabled wearable devices (e.g., smart watches, smart headphones, smart glasses, etc.), or vehicle infotainment systems.
[0039] FIG. 2 illustrates an example of the speech-to-speech model 200 of FIG. 1, comprising an encoder 212, a turn detector 215, and a decoder 220. The speech-to-speech S2S model 200 processes an input audio signal 102 corresponding to an utterance 108 produced by a target speaker 104 (FIG. 1) to generate a sequence of output audio frames 222 corresponding to a synthetic speech. The encoder 212 is enabled to receive the input audio data 102 of the utterance 108 and to generate a high-level feature representation 213 (also referred to herein as a hidden feature representation) for each frame 102 of the sequence of audio frames 102 of the audio data 102. Each high-level feature representation 213 is generated by the encoder 212 at an output step corresponding to a frame 102 of the sequence of audio frames. The turn detector 215 is enabled to receive the high-level feature representation 213 generated by the encoder 212 at each corresponding output step to generate a turn output 216 as a single bit. That is, for each corresponding output step, the turn detector 215 outputs a turn output 216 indicating whether the corresponding audio frame 102 is at a break point. Here, the turn detector 215 may predict a single bit ("0" or "1") based on the encoder output 213 at each output step by matching a deep neural model with attention to a logistic function. In some implementations, the turn output 216 comprises a probability distribution indicating the likelihood that the encoder output 212 at each output step is at a break point. The turn detector 215 may determine the turn output 216 of the corresponding high-level feature representation 213 by analyzing each high-level feature representation 213 individually. In some implementations, the encoder 212 processes the audio data 102 with an attention mechanism to obtain one or more high-level feature representations 213 (which may be transmitted as a single vector) that the turn detector 215 can use to generate the turn output 216.Thus, the turn detector 215 may determine whether the utterance 108 is at a break point based on the current high-level feature representation 213 (i.e., the current output step) and one or more high-level feature representations 213 already generated by the encoder in a previous output step (i.e., the turn detector 215 receives the history of the encoder state to predict whether the utterance 108 is currently at a break point).
[0040] The turn detector 215 may be located between the encoder 212 and the decoder 220. Alternatively, the turn detector 215 may be part of the encoder 212. In either case, the turn detector 215 is enabled to provide a turn output 216 to the decoder 220 together with the high-level feature representation 213 output by the encoder 212 at each corresponding output step, so that the decoder 220 can generate a sequence of output audio frames 222. The decoder 220 can generate a sequence of output audio frames 222 by processing the sequence of high-level feature representations 213 received at each output step, including the output step before and including the output step corresponding to the break point. Each output audio frame of the sequence of output audio frames 222 is based on a corresponding one of the high-level feature representations 213 generated by the audio encoder 212.
[0041] In some implementations, the decoder 220 does not begin generating the sequence of output audio frames 222 until it receives a turn output 216 that corresponds to a break point. In other implementations, the decoder 220 operates in a streaming mode by generating a corresponding output audio frame 222 for each high-level feature representation 213 generated by the encoder 212. Here, the synthesizer 275 may not begin synthesizing the sequence of output audio frames 212 until the turn detector 215 determines that the utterance 108 is at a break point (i.e., the synthesizer begins synthesizing output audio data 222 at an output step that corresponds to the turn output 216 that indicates a break point).
[0042] In some implementations, the voice-to-voice S2S model 200 includes multiple decoders 220 and selects an appropriate decoder 220 based on the speech type of the utterance 108. In these implementations, the voice-to-voice S2S model first determines the speech type and then selects the appropriate decoder 220. The voice-to-voice S2S model 200 is enabled to determine the speech type in various ways. In one implementation, the voice-to-voice S2S model 200 receives the utterance 108 and determines the speech type based on the characteristics of the utterance 108 (e.g., the speech may be accented, delayed, choppy, dysarthria, etc.). In other implementations, the voice-to-voice S2S model 200 receives the speech type from a user profile associated with the target speaker 104 (e.g., the user profile indicates that the user has a speech dysarthria). In yet another embodiment, the voice-to-voice S2S model 200 is enabled to receive an indication via user input indicating a voice type (e.g., user input to convert the voice to a particular language).
[0043] Based on the speech type, the speech-to-speech S2S model 200 is then enabled to select an appropriate decoder 220 to use in generating a sequence of output audio frames 222. For example, if the speech-to-speech S2S model 200 receives speech corresponding to a speech type of accented speech, the speech-to-speech S2S model 200 may select the streaming decoder 220. Here, the streaming decoder 220 may be suitable for synthesizing accented speech to non-accented speech when each word of the accented speech is directly synthesized to non-accented speech. However, when the speech type is translation, the streaming decoder 220 may not be suitable when there may be syntactic differences between languages and when the entire sentence or phrase may be required for correct translation. For example, the Turkish word "Benimle_gitti" is literally translated into English as "me_with_he_went". However, the proper translation of this phrase into English would be "he went with me." Therefore, for the speech type associated with the translation, the speech-to-speech S2S model 200 is enabled to select the standard decoder 220 to process the entire sentence. The standard decoder 220 may synthesize speech at each received break point.
[0044] FIG. 3 illustrates a training process 300 of the speech-to-speech model 200 with the turn detector 215. In some implementations, the speech-to-speech S2S model 200 is initialized / pre-trained before being further fine-tuned through training. For example, pre-training may include initiating the speech-to-speech S2S model 200 with pre-trained data 301 comprising a plurality of spoken utterances by a typical speaker associated with a standard fluent speech. The pre-trained data 301 may further include the spoken utterances paired with a corresponding synthesized ground truth standard fluent speech representation of the spoken utterance. In some implementations, the pre-trained data 301 comprises a respective break point or break points for each utterance of the plurality of utterances. The pre-trained speech-to-speech S2S model is then trained with data 302 from atypical voice users, allowing the speech-to-speech S2S model 200 to be further fine-tuned for atypical voice users. In another example, the speech-to-speech S2S model 200 is pre-configured with general parameters and then fine-tuned throughout the training process 300.
[0045] The training process 300 may include training any of the audio encoder 212, the turn detector 215, or the speech decoder 220, individually or jointly in any suitable combination. The process 300 includes providing training data 302 to the speech-to-speech S2S model 200. In some implementations, the training data 302 includes training input audio data 320 associated with a plurality of training atypical speech samples produced by one or more speakers associated with atypical speech. That is, for training each atypical speech sample, the training input audio data 320 may include a corresponding training sequence of acoustic frames. In these implementations, the training data 302 also includes training output audio data 321 associated with a plurality of ground truth standard speech samples each paired with a corresponding one of the training atypical speech samples. Here, each ground truth standard speech sample of the training output audio data 321 may include a sequence of acoustic frames corresponding to a standard fluent speech representation of the same utterance as the training input audio data 320. Additionally, the training data 302 may also include, for each of a plurality of training atypical speech samples of the input speech data 320, a sequence of break point labels 322 comprising a series of bits labeled with "1" and "0" indicating whether a corresponding acoustic frame of the input speech data 320 is at a break point. For example, the training data 302 may comprise training input audio data 320 corresponding to training atypical speech samples for an utterance produced by an atypical voice user, and training output audio data 321 corresponding to a standard fluent speech representation of the same utterance in the training input audio data 320. The first labels 321 may indicate which acoustic frames of the training input audio data 320 comprise break points of the utterance (e.g., respective periods in a transcription of the utterance). In particular, the first labels 321 may further correspond to a ground truth transcription of the training input audio data 320 indicating where the break points occur.
[0046] Further, the training data 302 may be based on user interaction with the output of the trained voice-to-voice S2S model 200. In this way, the voice-to-voice S2S model 200 is enabled to iteratively fine-tune based on real-world feedback. For example, the voice-to-voice S2S model 200 may receive streaming voice from a user. When the voice-to-voice S2S model 200 determines that the speech has reached a break point, the voice-to-voice S2S model 200 provides instructions 117 to cause the user device to output a turn indication 115 informing the user to stop speaking. However, if the user continues to speak (continue to speak) or overrides the turn indication 115 by other means, the determined break point may be considered incorrect. Here, the user's speech may be used as training data. The point at which the user stops speaking is used as a second label 322 corresponding to the target break point.
[0047] Upon receiving the training data 302, the voice-to-voice S2S model 200 is enabled to generate an output 350 (e.g., a sequence of output audio frames 222, a time-domain audio waveform of a synthetic voice representing speech uttered by a user, and / or a turn output 216). The voice-to-voice S2S model 200 is enabled to process the training data 302 in the manner described with respect to FIG. 2 above, or any other suitable method of voice-to-voice conversion.
[0048] In some implementations, the output 350 is used by one or more loss functions 305, 305a-b to generate one or more losses 310, 310a-c. That is, the first loss function 305a generates the first loss 310a by comparing the output 350 with corresponding ground truth standard speech samples of the training output audio data 321. The first loss 310a indicates a discrepancy between the predicted output 350 and the corresponding ground truth standard speech samples of the training output audio data 321. For example, the first loss function 305a is enabled to determine the first loss 310a by comparing the output 350 of the speech-to-speech S2S model 200 (i.e., a sequence of predicted output audio frames 222) with the training output audio data 321 comprising a sequence of ground truth output speech frames representing a standard fluent speech representation of the training input speech data 120.
[0049] Further, the second loss function 305b generates a second loss 310b by comparing the output 350 with the breakpoint label 322. The second loss 310b indicates a discrepancy between the breakpoint label 322 and the location of the predicted breakpoint in the output 350. For example, the second loss function 305b is enabled to determine the second loss 310b by comparing the output by the speech-to-speech S2S model 200 (i.e., the turn output 216) with the breakpoint label 322 corresponding to the target breakpoint in the training data 302. In some implementations, the turn output 216 can be a bit (i.e., “0” or “1”) generated by the turn detector 215 of the speech-to-speech S2S model 200 based on a high-level feature representation corresponding to each frame of the sequence of acoustic frames of the utterance, where “0” corresponds to a non-breakpoint and “1” corresponds to a breakpoint. Alternatively, the turn output 216 can be a probability distribution. In some implementations, the losses 310 may be combined to form a third loss 310c, or the losses 310 may be sent to the speech-to-speech S2S model 200 separately.
[0050] The loss function 305 may implement any suitable technique to determine the loss 310, such as regression loss, mean squared error, mean squared log error, mean absolute error, binary classification, binary cross entropy, hinge loss, multi-class loss, etc. The loss 310 may then be fed directly to the speech-to-speech S2S model 200, which may then be enabled to not only process the loss, but also adjust one or more parameters to account for the loss.
[0051] 4 is a flow chart of an example operational configuration of a method 400 of the speech-to-speech model 200 with turn detector 215. Data processing hardware 510 (FIG. 5) is enabled to perform the operations of the method 400 by executing instructions stored in memory hardware 520 (FIG. 5) that are communicated to the data processing hardware 510. The data processing hardware 510 and the memory hardware 520 may reside on the computing device 500 (FIG. 5), such as on the remote server 112 and / or the user computing device 110 of FIG. 1.
[0052] At operation 402, the method 400 comprises receiving, as an input to a speech-to-speech (S2S) model 200, a sequence of acoustic frames 102 corresponding to an utterance 108 uttered by the user 104 in streaming audio captured by a client device 110 associated with the user 104. At each of the plurality of output steps, operation 404 of the method 400 comprises generating, by an audio encoder 212 of the speech-to-speech S2S model 200, a high-level feature representation 213 of a corresponding acoustic frame 102 in the sequence of acoustic frames 102. At each of the plurality of output steps, operation 406 of the method 400 comprises determining, by a turn detector 215 of the speech-to-speech S2S model 200, whether the utterance 108 is at a break point at the corresponding output step based on the high-level feature representation 213 generated by the audio encoder 212 at the corresponding output step. At operation 408, the method comprises synthesizing the sequence of output audio frames 222 output by the speech decoder 220 of the speech-to-speech S2S model 200 into a time-domain audio waveform of synthetic speech representing the speech 108 uttered by the user 104 when the turn detector 215 determines that the speech is at a break, where each output audio frame 222 of the sequence of output audio frames 222 is based on a corresponding one of the high-level feature representations 213 generated by the audio encoder 212 up to the corresponding output step when the turn detector 215 determines that the speech is at a break. The sequence of output audio frames 222 corresponds to the output audio data 106.
[0053] A software application (i.e., software resource) may be referred to as computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0054] Non-transient memory may be a physical device used to store programs (e.g., instruction sequences) or data (e.g., program state information) temporarily or permanently for use by a computing device. Non-transient memory may be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0055] 5 is a schematic diagram of an exemplary computing device 500 that may be used to implement the systems and methods described herein. Computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown, their connections and relationships, and their functions are for illustrative purposes only and are not intended to limit the implementation of the invention described and / or claimed in this document.
[0056] Computing device 500 includes a processor 510, a memory 520, a storage device 530, a high-speed interface / controller 540 that connects to memory 520 and a high-speed expansion port 550, and a low-speed interface / controller 560 that connects to a low-speed bus 570 and storage device 530. Each component (510, 520, 530, 540, 550, and 560) is interconnected using various buses and may be mounted on a common motherboard or may exist otherwise as desired. Processor 510 is enabled to process instructions for execution within computing device 500, comprising instructions stored in memory 520 or storage device 530, to display graphical information of a graphical user interface (GUI) on an external input / output device, such as a display 580 coupled to high-speed interface 540. In other implementations, multiple processors and / or multiple buses may be used as needed, along with multiple memories and types of memories. Also, multiple computing devices 500 may be connected, each performing some of the required operations (eg, as a bank of servers, a group of blade servers, or a multi-processor system).
[0057] The memory 520 stores information non-transiently within the computing device 500. The memory 520 may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). The non-transient memory 520 may be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by the computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0058] The storage device 530 can provide mass storage for the computing device 500. In some implementations, the storage device 530 is a computer-readable medium. In various different implementations, the storage device 530 can be an array of devices, including floppy disk devices, hard disk devices, optical disk devices, or tape devices, flash memory or other similar solid-state memory devices, or devices in a storage area network or other configuration. In additional implementations, the computer program product is tangibly embodied in an information carrier. The computer program product comprises instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as the memory 520, the storage device 530, or a memory on the processor 510.
[0059] The high-speed controller 540 manages the bandwidth-intensive operations of the computing device 500, and the low-speed controller 560 manages the low-bandwidth intensive operations. Such an allocation of roles is merely an example. In some implementations, the high-speed controller 540 is coupled to the memory 520, the display 580 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 550 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 560 is coupled to the storage device 530 and the low-speed expansion port 590. The low-speed expansion port 590 may include various communication ports (USB, Bluetooth, Ethernet, wireless Ethernet, etc.) and can connect to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a network device, such as a switch or a router, via a network adapter or the like.
[0060] The computing device 500 can be implemented in many different forms, as shown in the figure: for example, it can be implemented as a standard server 500a, or as a plurality in a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.
[0061] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may be specialized or general purpose and may include implementations in one or more computer programs executable and / or interpretable by a programmable system having at least one programmable processor, at least one input device, and at least one output device coupled to receive data and instructions from the storage system and to transmit data and instructions to the storage system.
[0062] These computer programs (also known as programs, software, software applications, or codes) comprise machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor comprising a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0063] The processes and logic flows described herein can be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to operate on input data and generate output to perform functions. The processes and logic flows can also be implemented by special purpose logic circuitry, such as FPGAs (field programmable gate arrays) or ASICs (application specific integrated circuits). Processors suitable for executing computer programs include, for example, both general purpose and special purpose processors, as well as any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from a read-only memory, a random access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic disks, magneto-optical disks, or optical disks, for storing data, or is operatively connected to receive data from or transmit data to them, or both. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0064] To interact with a user, one or more aspects of the present disclosure can be implemented on a computer that optionally has a display device, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, for the user to provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic, verbal, or tactile input. Additionally, the computer can interact with the user by sending documents to a device used by the user and receiving documents from the device, such as by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0065] Although several embodiments have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the disclosure. Accordingly, other embodiments are within the scope of the following claims.
Claims
1. A computer-implemented method (400) that, when executed on data processing hardware (510), causes the data processing hardware (510) to perform operations, the operations comprising: receiving, as input to a speech-to-speech S2S model (200), a sequence of acoustic frames (102) corresponding to speech (108) produced by a user in streaming audio captured by a client device (110) associated with said user; In each of the plurality of output steps, generating, by an audio encoder (212) of the speech-to-speech S2S model (200), a high-level feature representation (213) of a corresponding acoustic frame (102) in the sequence of acoustic frames (102); determining, by a turn detector (215) of the speech-to-speech S2S model (200), whether the utterance (108) is at a break point at the corresponding output step based on the high-level feature representation (213) generated by the audio encoder (212); and synthesizing a sequence of output audio frames (222) already output by a speech decoder (220) of the speech-to-speech S2S model (200) into a time-domain audio waveform of a synthetic speech (114) representing the speech (108) uttered by the user when the turn detector (215) determines that the utterance (108) is at the break point, where each output audio frame (222) of the sequence of output audio frames (222) is based on a corresponding one of the high-level feature representations (213) generated by the audio encoder (212) up to the corresponding output step when the turn detector (215) determines that the utterance (108) is at the break point; A computer-implemented method (400) comprising:
2. The operation further comprises: When the utterance (108) is at the interruption point, As an output from the client device (110), a speaker turn indication (115) to inform the user to stop speaking; and a time-domain audio waveform of the synthesized speech (114) representing the speech (108) produced by the user; providing a The computer-implemented method (400) of claim 1.
3. the speech (108) produced by the user in the streaming audio captured by the client device (110) is associated with atypical speech; the time-domain audio waveform of the synthesized speech (114) representing the utterance (108) comprises a time-domain audio waveform of a synthesized standard fluent speech of the same utterance (108) produced by the user.
3. A computer-implemented method (400) according to claim 1 or 2.
4. the speech (108) uttered by the user in the streaming audio captured by the client device (110) is in a first language; the time-domain audio waveform of the synthesized speech (114) representing the utterance (108) comprises a time-domain audio waveform of a synthesized translation of the same utterance (108) in a second language different from the first language. A computer implemented method (400) according to any one of claims 1 to 3.
5. determining whether the utterance (108) is at the breakpoint in the corresponding output step further comprises determining whether the utterance (108) is at the breakpoint in the corresponding output step based on one or more of the high-level feature representations (213) generated by the audio encoder (212) in an output step preceding the corresponding output step; A computer implemented method (400) according to any one of claims 1 to 4.
6. The operations further comprise generating, at each of a plurality of said output steps, by the speech decoder (220) of the speech-to-speech S2S model (200), an output audio frame (222) of a corresponding high-level feature representation (213) generated by the audio encoder (212) at a corresponding said output step. A computer implemented method (400) according to any one of claims 1 to 5.
7. The operations further include, in response to determining that the utterance (108) is at the break point of the corresponding output step: receiving the sequence of high-level feature representations (213) generated by the audio encoder (212) as inputs to the speech decoder (220) of the speech-to-speech S2S model (200) up to the corresponding output step; and generating, by the speech decoder (220) of the speech-to-speech S2S model (200), a sequence of the output audio frames (222); Equipped with A computer implemented method (400) according to any one of the preceding claims.
8. The turn detector (215) comprises a deep neural network. A computer implemented method (400) according to any one of the preceding claims.
9. the turn detector (215) is disposed between the audio encoder (212) and the speech decoder (220); A computer implemented method (400) according to any one of the preceding claims.
10. determining whether the utterance (108) is at the break point in a corresponding output step includes generating, as an output from the turn detector (215), a turn output (216) indicative of whether the utterance (108) is at the break point; A computer implemented method (400) according to any one of the preceding claims.
11. The turn output (216) comprises a bit.
11. The computer-implemented method (400) of claim 10.
12. the turn output (216) comprises a probability distribution.
11. The computer-implemented method (400) of claim 10.
13. The operations further comprise training the speech-to-speech S2S model (200), the speech-to-speech S2S model (200) comprising: receiving a set of training utterances (302), each of the training utterances (302) in the set of training utterances (302) comprising a corresponding sequence of training acoustic frames (320), each of the training acoustic frames (320) being annotated with a label (321, 322) indicating whether the corresponding training acoustic frame (320) corresponds to a breakpoint frame or a non-breakpoint frame, each of the training utterances (320) being paired with a corresponding ground truth synthetic speech representation of the training utterance (302); obtaining a first label (321) of the training utterance (302) indicative of a target output spectrogram; obtaining a second label (322) for the training utterance (302), the second label (322) being indicative of a target turn output (216); generating training outputs (350) using the speech-to-speech S2S model (200) and the training utterances (302), the training outputs comprising training output spectrograms corresponding to synthetic speech representations of the training acoustic frames (320) and training turn outputs indicative of break points of the training acoustic frames (320); determining a first loss (310) by comparing the training output spectrograms with the first labels (321); determining a second loss (310) by comparing the training turn output (216) with the second label (322); and optimizing the speech-to-speech S2S model (200) based on the first loss (310) and the second loss (310) associated with training input audio data (320); Trained by A computer implemented method (400) according to any one of the preceding claims.
14. The operation further comprises: determining a voice type of the speech (108) produced by the user that has been captured in the streaming audio by the client device (110); and selecting, from among a plurality of available audio decoders, the audio decoder (220) for generating the sequence of output audio frames (222); Equipped with A computer implemented method (400) according to any one of the preceding claims.
15. each said output audio frame (222) of the sequence of output audio frames (222) output by said speech decoder (220) comprising a spectrogram frame; A computer implemented method (400) according to any one of the preceding claims.
16. Data processing hardware (510); and memory hardware (520) storing instructions that, when executed by said data processing hardware (510), cause said data processing hardware (510) to perform operations; A system (100) comprising: receiving, as input to a speech-to-speech S2S model, a sequence of acoustic frames (102) corresponding to speech (108) produced by a user in streaming audio captured by a client device (110) associated with said user; and In each of the plurality of output steps, generating, by an audio encoder (212) of the speech-to-speech S2S model (200), a high-level feature representation (213) of a corresponding acoustic frame (102) in the sequence of acoustic frames (102); determining, by a turn detector (215) of the speech-to-speech S2S model (200), whether the utterance (108) is at a break point at the corresponding output step based on the high-level feature representation (213) generated by the audio encoder (212); and synthesizing, when the turn detector (215) determines that the utterance (108) is at the break point, a sequence of output audio frames (222) already output by a speech decoder (220) of the speech-to-speech S2S model (200) into a time-domain audio waveform of a synthetic speech (114) representing the utterance (108) uttered by the user, where each output audio frame (222) of the sequence of output audio frames (222) is based on a corresponding one of the high-level feature representations (213) already generated by the audio encoder (212) until the corresponding output step when the turn detector (215) determines that the utterance (108) is at the break point; The system (100) comprises:
17. The operation further comprises: When the utterance (108) is at the interruption point, As an output from the client device (110), a speaker turn indication (115) to inform the user to stop speaking; and a time-domain audio waveform of the synthesized speech (114) representing the speech (108) produced by the user; providing a 17. The system (100) of claim 16.
18. the speech (108) produced by the user in the streaming audio captured by the client device (110) is associated with atypical speech; the time-domain audio waveform of the synthesized speech (114) representing the utterance (108) comprises a time-domain audio waveform of a synthesized standard fluent speech of the same utterance (108) produced by the user.
18. A system (100) according to claim 16 or 17.
19. the speech (108) uttered by the user in the streaming audio captured by the client device (110) is in a first language; the time-domain audio waveform of the synthesized speech (114) representing the utterance (108) comprises a time-domain audio waveform of a synthesized translation of the same utterance (108) in a second language different from the first language. A system (100) according to any one of claims 16 to 18.
20. determining whether the utterance (108) is at the breakpoint in the corresponding output step further comprises determining whether the utterance (108) is at the breakpoint in the corresponding output step based on one or more of the high-level feature representations (213) generated by the audio encoder (212) in an output step preceding the corresponding output step; A system (100) according to any one of claims 16 to 19.
21. The operations further comprise generating, at each of a plurality of said output steps, by the speech decoder (220) of the speech-to-speech S2S model (200), an output audio frame (222) of a corresponding high-level feature representation (213) generated by the audio encoder (212) at a corresponding said output step. A system (100) according to any one of claims 16 to 20.
22. The operations further include, in response to determining that the utterance (108) is at the break point of the corresponding output step: receiving the sequence of high-level feature representations (213) generated by the audio encoder (212) as inputs to the speech decoder (220) of the speech-to-speech S2S model (200) up to the corresponding output step; and generating, by the speech decoder (220) of the speech-to-speech S2S model (200), a sequence of the output audio frames (222); Equipped with A system (100) according to any one of claims 16 to 21.
23. The turn detector (215) comprises a deep neural network. A system (100) according to any one of claims 16 to 22.
24. the turn detector (215) is disposed between the audio encoder (212) and the speech decoder (220); A system (100) according to any one of claims 16 to 23.
25. determining whether the utterance (108) is at the break point in a corresponding output step includes generating, as an output from the turn detector (215), a turn output (216) indicative of whether the utterance (108) is at the break point; A system (100) according to any one of claims 16 to 24.
26. The turn output (216) comprises a bit.
26. The system (100) of claim 25.
27. the turn output (216) comprises a probability distribution.
26. The system (100) of claim 25.
28. The operations further comprise training the speech-to-speech S2S model (200), the speech-to-speech S2S model (200) comprising: receiving a set of training utterances (302), each of the training utterances (302) in the set of training utterances (302) comprising a corresponding sequence of training acoustic frames (320), each of the training acoustic frames (320) being annotated with a label (321, 322) indicating whether the corresponding training acoustic frame (320) corresponds to a breakpoint frame or a non-breakpoint frame, and each of the training utterances (302) being paired with a corresponding ground truth synthetic speech representation of the training utterance (302); obtaining a first label (321) of the training utterance (302) indicative of a target output spectrogram; obtaining a second label (322) for the training utterance (302), the second label (322) being indicative of a target turn output (216); generating training outputs (350) using the speech-to-speech S2S model (200) and the training utterances (302), the training outputs comprising training output spectrograms corresponding to synthetic speech representations of the training acoustic frames (320) and training turn outputs indicative of break points of the training acoustic frames (320); determining a first loss (310) by comparing the training output spectrograms with the first labels (321); determining a second loss (310) by comparing the training turn output (216) with the second label (322); and optimizing the speech-to-speech S2S model (200) based on the first loss (310) and the second loss (310) associated with training input audio data (320); Trained by A system (100) according to any one of claims 16 to 27.
29. The operation further comprises: determining a voice type of the speech (108) produced by the user that has been captured in the streaming audio by the client device (110); and selecting, from among a plurality of available audio decoders, the audio decoder (220) for generating the sequence of output audio frames (222); Equipped with A system (100) according to any one of claims 16 to 28.
30. each said output audio frame (222) of the sequence of output audio frames (222) output by said speech decoder (220) comprising a spectrogram frame; A system (100) according to any one of claims 16 to 29.
Citation Information
Patent Citations
Turn-taking timing identification apparatus, turn-taking timing identification method, program and recording medium
JP2018132678A
System and method for end-to-end speech recognition with trigger door tension
JP2022522379A
Synthesized data augmentation using voice conversion and speech recognition models
WO2022046526A1
System and method for direct speech translation system
US20200226327A1