Streaming speech-to-speech model with automatic speaker turn detection

The speech-to-speech model with a turn detector addresses the inconvenience of manual input by automatically detecting breakpoints, enhancing the natural flow and intelligibility of synthesized speech.

JP7776673B2Active Publication Date: 2025-11-26GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024567521
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-06-03
Filing Date
2023-05-17
Publication Date
2025-11-26
Estimated Expiration
2043-05-17

AI Technical Summary

Technical Problem

Conventional speech-to-speech models require manual user input to indicate the start and end of an utterance, leading to inconvenience and potential human error, which disrupts the natural flow of conversation.

Method used

A speech-to-speech model with an integrated turn detector that automatically detects breakpoints in streaming audio, allowing for seamless conversion of user speech into synthesized speech without requiring manual input.

Benefits of technology

The model provides a more natural user experience by mimicking normal conversation, improving intelligibility and reducing errors by automatically determining appropriate moments for synthesized speech output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007776673000001
    Figure 0007776673000001
  • Figure 0007776673000002
    Figure 0007776673000002
  • Figure 0007776673000003
    Figure 0007776673000003
Patent Text Reader

Abstract

A method (400) for turn detection in a speech-to-speech (S2S) model (200) receives a sequence of acoustic frames (102) corresponding to an utterance (108). At each of a plurality of output steps, an audio encoder (210) generates a high-level feature representation (213) of a corresponding acoustic frame in the sequence of acoustic frames. At a corresponding output step, a turn detector (215) of the speech-to-speech S2S model determines whether the utterance is at a break point at the corresponding output step based on the high-level feature representation generated by the audio encoder. When the turn detector determines that the utterance is at a break point, the method synthesizes a sequence of output audio frames (222) output by a speech decoder (220) of the speech-to-speech S2S model into a time-domain audio waveform of synthetic speech representing the speech uttered by the user.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a speech-to-speech model with a turn detector. [Background technology]

[0002] Speech-to-speech (S2S) models can be used to convert source speaker speech into synthetic speech without changing the linguistic information of the original speech. For example, S2S models can generate canonical, fluent synthetic speech for users with dysarthria or atypical speech. Alternatively, S2S models can translate a user's speech into synthetic speech in another language. Typically, S2S models are manually triggered by user input indicating when input speech begins and ends. After the user finishes speaking, the S2S model then processes the input speech to generate synthetic speech. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] US Patent Application Publication No. 2020 / 0226327 Summary of the Invention [Problem to be solved by the invention]

[0004] There is an opportunity to provide a streaming speech-to-speech model with improved automatic speaker turn detection. [Means for solving the problem]

[0005] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations comprising receiving as input to a speech-to-speech (S2S) model a sequence of acoustic frames corresponding to spoken utterances made by a user in streaming audio captured by a client device associated with the user. At each of the plurality of output steps, the operations also include generating, by an audio encoder of the speech-to-speech S2S model, a high-order feature representation of a corresponding audio frame in the sequence of audio frames, determining, by a turn detector of the speech-to-speech S2S model, at the corresponding output step, whether an utterance is at a breakpoint at the corresponding output step based on the high-order feature representation generated by the audio encoder, and synthesizing, when the turn detector determines that the utterance is at a breakpoint, the sequence of output audio frames output by the speech decoder of the speech-to-speech S2S model into a time-domain audio waveform of synthetic speech representing the spoken utterance uttered by the user, wherein each output audio frame in the sequence of output audio frames is based on a corresponding one of the high-order feature representations generated by the audio encoder until the corresponding output step at which the turn detector determines that the utterance is at a breakpoint.

[0006] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, the operations further include providing, as output from the client device when the utterance is at the breakpoint, a speaker turn indication informing the user to stop speaking and a time-domain audio waveform of synthesized speech representing the utterance uttered by the user. The utterance uttered by the user in the streaming audio captured by the client device may be associated with atypical speech, and the time-domain audio waveform of the synthesized speech representing the utterance may include a time-domain audio waveform of synthesized canonical fluent speech of the same utterance uttered by the user. Further, the utterance uttered by the user in the streaming audio captured by the client device may be in a first language, and the time-domain audio waveform of the synthesized speech representing the utterance may include a time-domain audio waveform of synthesized translation speech of the same utterance in a second language different from the first language.

[0007] In some examples, determining whether the utterance is at a breakpoint at a corresponding output step is further based on one or more of the high-level feature representations generated by the audio encoder at an output step preceding the corresponding output step. Additionally or alternatively, the operations may further include generating, at each of the plurality of output steps, an output audio frame of the corresponding high-level feature representation generated by the audio encoder at the corresponding output step, by a speech decoder of the speech-to-speech S2S model.

[0008] In some implementations, the operations also include, in response to determining that the utterance is at a break point of a corresponding output step, receiving as input to a speech decoder of the speech-to-speech S2S model the sequence of high-level feature representations generated by the audio encoder up to the corresponding output step, and generating, by the speech decoder of the speech-to-speech S2S model, a sequence of output audio frames. The turn detector may include a deep neural network that may be disposed between the audio encoder and the speech decoder.

[0009] In some additional embodiments, determining whether the utterance is at a break point in a corresponding output step comprises generating a turn output as an output from the turn detector indicating whether the utterance is at a break point, where the turn output comprises a bit or a probability distribution.

[0010] In some examples, the operations further include training the speech-to-speech S2S model by receiving a set of training utterances, where each training utterance in the set of training utterances comprises a corresponding sequence of training acoustic frames, and each training utterance is paired with a corresponding ground truth synthetic speech representation of the training utterance. Each training acoustic frame in the sequence of training acoustic frames is annotated with a label indicating whether the corresponding training acoustic frame corresponds to a breakpoint frame or a non-breakpoint frame. In these examples, the speech-to-speech S2S model is further trained by obtaining a first label for the training input audio data indicative of a target output spectrogram, obtaining a second label for the training input audio data indicative of a target turn output, and generating a training output using the speech conversion model and the training input audio data. The training output comprises a training output spectrogram corresponding to the synthetic speech representation of the training input audio data and a training turn output indicative of the breakpoints in the training input audio data. Finally, the speech-to-speech S2S model is trained by determining a first loss by comparing the training output spectrogram with a first label, determining a second loss by comparing the training turn output with a second label, and optimizing the speech conversion model based on the first loss and the second loss associated with the training input audio data.

[0011] In some implementations, the operations also include determining a voice type of speech uttered by the user captured in the streaming audio by the client device, and selecting a voice decoder from among a plurality of available voice decoders to generate a sequence of output audio frames, wherein each output audio frame in the sequence of output audio frames output by the voice decoder comprises a spectrogram frame.

[0012] Another aspect of the present disclosure provides a system including data processing hardware and memory hardware in communication with the data processing hardware. The memory stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving, as input to a speech-to-speech (S2S) model, a sequence of acoustic frames corresponding to speech uttered by a user in streaming audio captured by a client device associated with the user. At each of a plurality of output steps, the operations also include generating, by an audio encoder of the speech-to-speech S2S model, a high-order feature representation of the corresponding acoustic frame in the sequence of acoustic frames; determining, by a turn detector of the speech-to-speech S2S model, at the corresponding output step based on the high-order feature representation generated by the audio encoder, whether the speech is at a break point at the corresponding output step; and synthesizing, when the turn detector determines that the speech is at a break point, the sequence of output audio frames output by the speech decoder of the speech-to-speech S2S model into a time-domain audio waveform of synthetic speech representing the speech uttered by the user. Here, each output audio frame in the sequence of output audio frames is based on a corresponding one of the high-level feature representations generated by the audio encoder up to the corresponding output step when the turn detector determines that the speech is at a break point.

[0013] This aspect may include one or more of the following optional features. In some implementations, the operations further include providing, as output from the client device when the utterance is at a pause point, a speaker turn indication informing the user to stop speaking and a time-domain audio waveform of synthesized speech representing the utterance uttered by the user. The utterance uttered by the user in the streaming audio captured by the client device may be associated with atypical speech. The time-domain audio waveform of the synthesized speech representing the utterance may include a time-domain audio waveform of synthesized standard fluent speech of the same utterance uttered by the user. Further, the utterance uttered by the user in the streaming audio captured by the client device may be in a first language. The time-domain audio waveform of the synthesized speech representing the utterance may include a time-domain audio waveform of synthesized translation speech of the same utterance in a second language different from the first language.

[0014] In some examples, determining whether the utterance is at a breakpoint at a corresponding output step is further based on one or more of the high-level feature representations generated by the audio encoder at an output step preceding the corresponding output step. Additionally or alternatively, the operations may further include generating, at each of the plurality of output steps, an output audio frame of the corresponding high-level feature representation generated by the audio encoder at the corresponding output step, by a speech decoder of the speech-to-speech S2S model.

[0015] In some implementations, the operations also include, in response to determining that the utterance is at a break point of a corresponding output step, receiving as input to a speech decoder of the speech-to-speech S2S model the sequence of high-level feature representations generated by the audio encoder up to the corresponding output step, and generating, by the speech decoder of the speech-to-speech S2S model, a sequence of output audio frames. The turn detector may include a deep neural network that may be disposed between the audio encoder and the speech decoder.

[0016] In some additional embodiments, determining whether the utterance is at a break point in a corresponding output step comprises generating a turn output as an output from the turn detector indicating whether the utterance is at a break point, where the turn output comprises a bit or a probability distribution.

[0017] In some examples, the operations further include training the speech-to-speech S2S model by receiving a set of training utterances, whereby each training utterance in the set of training utterances comprises a corresponding sequence of training acoustic frames, and each training utterance is paired with a corresponding ground truth synthetic speech representation of the training utterance. Each training acoustic frame in the sequence of training acoustic frames is annotated with a label indicating whether the corresponding training acoustic frame corresponds to a breakpoint frame or a non-breakpoint frame. In these examples, the speech-to-speech S2S model is further trained by obtaining a first label for the training input audio data indicative of a target output spectrogram, obtaining a second label for the training input audio data indicative of a target turn output, and generating a training output using the speech conversion model and the training input audio data. The training output comprises a training output spectrogram corresponding to the synthetic speech representation of the training input audio data and a training turn output indicative of the breakpoints in the training input audio data. Finally, the speech-to-speech S2S model is trained by determining a first loss by comparing the training output spectrogram with a first label, determining a second loss by comparing the training output with a second label, and optimizing the speech conversion model based on the first loss and the second loss associated with the training input audio data.

[0018] In some implementations, the operations also include determining a voice type of speech uttered by the user captured in the streaming audio by the client device, and selecting a voice decoder from among a plurality of available voice decoders to generate a sequence of output audio frames, wherein each output audio frame in the sequence of output audio frames output by the voice decoder comprises a spectrogram frame.

[0019] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0020] [Figure 1] FIG. 1 is a schematic diagram of an exemplary voice-to-voice system with a turn detector. [Figure 2] FIG. 1 is a schematic diagram of an exemplary speech-to-speech model with a turn detector. [Figure 3] FIG. 1 is a schematic diagram of an exemplary training process for a speech-to-speech model with a turn detector. [Figure 4] 1 is a flowchart of an exemplary operational arrangement of a method for performing voice-to-voice conversion with a turn detector. [Figure 5] FIG. 1 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION

[0021] Like reference symbols in the various drawings indicate like elements. There is growing interest in developing more inclusive speech technologies, especially those that can assist people with speech disorders. Speech-to-speech (S2S) conversion has made significant advances with the introduction of end-to-end (E2E) deep learning-based models for recognizing and converting the speech of speakers with dysarthria or atypical speech patterns into synthesized speech. For example, atypical speech patterns can include, but are not limited to, speech disorders resulting from physical or neurological conditions (e.g., speakers with amyotrophic lateral sclerosis (ALS)), heavily accented speech, and deaf speech. Another application of speech-to-speech S2S models is translation, where a user speaks in a first language while a speech-to-speech S2S model generates synthesized speech in a second language.

[0022] A conventional voice-to-voice S2S model must receive the entire input speech (i.e., an utterance) before generating synthesized speech. Typically, a user provides input indicating the start and end of an utterance to the voice-to-voice S2S model. For example, the user may hold down a button on a user interface corresponding to the voice-to-voice S2S model when the user begins speaking (i.e., begins providing input speech to the voice-to-voice S2S model). The user then releases the button when the user finishes speaking. The voice-to-voice S2S model then converts the received input speech into synthesized speech while the user held down the button. Such manual methods for activating a voice-to-voice S2S model can be inconvenient for users because they require a lot of user input. Furthermore, requiring user input can lead to human error (e.g., the user may release the button before completing the utterance).

[0023] The present disclosure introduces a turn detector that automatically determines appropriate break points in a user's speech so that synthesized speech converted from the user's speech by a speech-to-speech S2S model can be output at those break points in a streaming manner. Thus, instead of requiring user input indicating when speech begins / ends, the present disclosure receives streaming audio from the user and automatically determines appropriate moments in the speech, so-called "break points," at which the user can pause to receive / listen to synthesized speech generated by a speech-to-speech S2S model. By automatically determining break points, embodiments of the present disclosure provide a more natural user experience because the turn-by-turn nature of the system mimics normal conversation (i.e., the user speaks when it is their turn, and then the system speaks when it is its turn).

[0024] As used herein, unless otherwise specified, the terms “speech-to-speech system” and “speech-to-speech model” can refer to any system / model that directly reconverts input speech to synthetic speech without performing intermediate speech recognition on the input speech. In other words, a speech-to-speech system / model is configured to directly convert an input audio waveform, sequence of acoustic frames, or spectrogram corresponding to the input speech into an output audio waveform or spectrogram corresponding to synthetic speech without converting it to an intermediate representation (e.g., text or phonemes). As will become apparent, speech-to-speech models and techniques for training speech-to-speech models enable users with atypical voices to speak to and be understood by both other humans and voice interfaces (e.g., digital assistants) by enabling the recognition and / or playback of user-intended speech.

[0025] Although the examples herein illustrate a speech-to-speech model that receives an input utterance corresponding to atypical speech and converts it into synthetic speech corresponding to standard fluent speech, the speech-to-speech model can be similarly adapted to perform other types of speech conversion tasks without departing from the scope of this disclosure. For example, a speech-to-speech S2S model can be enabled to convert an input utterance in a first language into synthetic speech that corresponds to a translation of the input utterance in a different second language. A speech-to-speech S2S model can similarly receive input spoken by a user and output synthetic speech having the same linguistic content as the spoken input but different speech characteristics of a target speaker.

[0026] 1 illustrates a speech-to-speech system 100 that includes a speech-to-speech (S2S) model 200. The speech-to-speech (S2S) model 200 is configured to directly convert input audio data 102 (e.g., a sequence of acoustic frames or an input audio waveform) corresponding to an utterance 108 produced by a target speaker 104 into output audio data 106 (e.g., a sequence of output audio frames or an output audio waveform) corresponding to a synthetic speech representation of the same utterance 114 produced by the target speaker 104. Notably, the speech-to-speech S2S conversion model 200 is configured to directly convert the input audio data 102 to the output audio data 106 without performing speech recognition or otherwise requiring the generation of any intermediate discrete representation (e.g., text or phonemes) from the input audio data 102.

[0027] The speech-to-speech S2S conversion model 200 includes an audio encoder 212 configured to encode input audio data 102 into a hidden feature representation (e.g., a series of vectors), a turn detector 215 configured to determine breakpoints in an utterance 108 based on the hidden feature representation output by the encoder, and a speech decoder 220 configured to decode the hidden representation into output audio data 106 corresponding to a synthesized standard fluent speech representation. For example, when the audio encoder 200 receives input audio data 102 of an utterance 108, the audio encoder 200 is processing five frames of audio. The audio encoder 200 is enabled to convert these five frames of audio into ten vectors. The vectors are not transcriptions of the frames of the audio data 102, but are mathematical representations of the frames of the audio data 102. The turn detector 215 is then enabled to determine whether a breakpoint exists among the ten vectors. If a break point is present, the speech decoder 220 is enabled to generate output audio data 106 corresponding to a synthesized standard fluent speech representation based on the vectors received from audio encoder 200 for vectors prior to the break point. For example, the turn detector 215 may determine that the tenth vector is a break point, resulting in the speech decoder 220 being enabled to receive ten vectors representing five frames of audio from audio encoder 200. The speech decoder 220 is now enabled to generate five frames of output audio data 106 corresponding to a synthesized standard fluent speech representation of the utterance 114 that comprises the intended word or part of a word as the five frames of input audio data 102 but is free of the atypical speech disfluency.

[0028] The voice-to-voice S2S conversion system 100 may further include a synthesizer 275 for synthesizing the output audio data 106 into a time-domain waveform for audible output as the synthesized utterance 114 of the utterance 108. The time-domain audio waveform comprises an audio waveform that defines the amplitude of the audio signal over time. The synthesizer 275 may include a unit selection module or a WaveNet module for synthesizing the output audio data 106 into a time-domain waveform of standard fluent speech. In some implementations, the synthesizer 275 comprises a vocoder network, i.e., a neural vocoder separately trained and conditioned on a Mel-frequency spectrogram for conversion to a time-domain audio waveform. In additional implementations, the synthesizer 275 comprises a streaming vocoder configured to convert / invert a log-magnitude spectrogram output from the speech decoder 220 as the output audio data 106 into a time-domain audio waveform in real time.

[0029] In some implementations, the target speaker 104 is associated with an atypical speech, and the target speaker 104 speaks with atypical speech patterns that may be difficult to understand. Atypical speech patterns may include, but are not limited to, speech disorders due to physical or neurological conditions (e.g., speakers with amyotrophic lateral sclerosis (ALS)), heavily accented speech, and the speech of deaf people. In other implementations, the target speaker 104 speaks in a first language, while the speech-to-speech S2S translates the first speech into a second language.

[0030] Thus, the speech-to-speech conversion system 100 is trained to directly convert input audio data 102 corresponding to utterances 108 produced by a target speaker 104 associated with atypical speech into output audio data 106 corresponding to a synthesized standard fluent speech representation of the same utterance 108. The synthesized standard fluent speech representation provided by the output audio data 106 thus improves the intelligibility of the atypical speech (e.g., heavily accented speech or amyotrophic lateral sclerosis (ALS) speech) produced by the target speaker 104.

[0031] Without departing from the scope of this disclosure, the speech-to-speech conversion system 100 may be trained to directly convert input audio data 102 corresponding to utterances 108 associated with atypical speech in a first language into output audio data 106 corresponding to a synthesized standard fluent speech representation of the same utterance 108 in the same voice but in a different second language.

[0032] Furthermore, the turn detector 215 may be trained to determine break points in the input audio data 102 in real time as the utterances 108 produced by the target speaker 104 are captured in streaming audio by the user device 110 associated with the target speaker 104. In some implementations, the turn detector 215 comprises a deep neural network. The turn detector 215 may be disposed between the audio encoder 212 and the speech decoder 220, such that the turn detector 215 receives the output 213 from the audio encoder 212 (e.g., a high-level feature representation of each of the sequence of acoustic frames) and sends a turn output 216 to the decoder 220. The turn output indicates whether the speech is at a break point. In some implementations, the turn output is a series of bits (e.g., “1” and “0”), where “1” indicates a break point in the speech and “0” indicates no break point. Here, each bit in the series of bits may indicate a corresponding acoustic frame in the input audio data 102. In other implementations, the turn output is a probability score (i.e., a number between 0 and 1) indicating the likelihood of a break point in the corresponding acoustic frame. For example, if the probability score of the corresponding acoustic frame meets a break point threshold, the acoustic frame indicates a break point. If the turn output 216 indicates that the utterance is at a break point, the speech-to-speech S2S model 200 may provide instructions 117 to the user device 110 to cause the user device 110 to output a turn indication 115 indicating that the user should stop speaking.

[0033] The speech decoder 220 may then receive the output 213 from the encoder 212 and the turn output 216 from the turn detector 215 during each of the multiple time steps. In some implementations, the turn detector 215 modifies the output 213 from the encoder 212 such that the output 213 from the encoder 212 indicates a break point.

[0034] A user device (interchangeably referred to as a “computing device”) 110 associated with a target speaker 104 may capture utterances 108 produced by the target speaker 104 in streaming audio and transmit corresponding input audio data 102 to a voice-to-voice conversion system 100 for conversion into output audio data 106. The voice-to-voice conversion system 100 may then transmit output audio data 106 corresponding to a synthesized speech representation of the same utterance 114 produced by the target speaker 104 to another computing device 116 associated with a user 118, which then audibly outputs the synthesized speech representation of the utterance 108 produced by the target speaker 104. In this example, the target speaker 104 and the user 118 converse with each other through their respective computing devices 110, 116 via telephone calls or other types of voice communication protocols, such as voice over Internet Protocol. Although the target speaker 104 and other users 118 may be conversing in the same language, the target speaker 104 may have atypical speech due to amyotrophic lateral sclerosis (ALS), making it difficult for the other users 118 to understand the target speaker 104. Thus, while the target speaker 104 speaks in an atypical speech (e.g., ALS speech) that is difficult to understand, the other users 118 listening to the synthesized standard fluent speech representation have an easier time understanding the utterance 108 intended by the target speaker 104. In other words, the synthesized standard fluent speech representation provides a more consistent cadence that may be easier for another user to understand than the original utterance 108 produced by the target speaker in atypical speech. Notably, the synthesized standard fluent speech representation is in the voice of the target speaker 104.

[0035] In some other examples, the speech-to-speech S2S conversion system 100 instead passes output audio data 106 corresponding to a synthesized speech representation of the utterances uttered by the target speaker 104 to an output audio device for audibly outputting the synthesized speech representation to an audience in the voice of the target speaker 104. For example, the target speaker 104 may be a psychology professor lecturing to a class of students, and the utterances uttered by the target speaker 104 comprise medical terminology belonging to a particular specific domain, e.g., psychology. As will become apparent, the speech-to-speech S2S conversion system 100 is trained to learn linguistic diversity derived from the linguistic content present in the training utterances and acoustic diversity associated with particular types of atypical speech associated with the speakers who uttered the target utterances.

[0036] Alternatively, the other computing device 116 may be associated with a downstream automatic speech recognition (ASR) system in which the speech-to-speech system 100 acts as a front end that provides output audio data 106 corresponding to the synthesized speech representation as input to the automatic speech recognition ASR system for conversion into recognized text. The recognized text may be presented to another user 118 and / or provided to a natural language understanding (NLU) system for further processing.

[0037] In any of the above examples, if the turn detector 215 determines that the target speaker 104 has reached a pause in an utterance during the utterance, the voice-to-voice S2S model 200 may provide instructions 117 to cause the user device 110 associated with the target speaker 104 to output a turn indication 115. For example, the user device 110 in FIG. 1 shows the user device 110 displaying a graphical turn indication 115a as a series of exclamation points to signal the target speaker 104 to pause and to enable the voice-to-voice S2S model 200 to generate output speech data 106 corresponding to a synthesized standard fluent speech representation of the input audio data 102. In a further example, the user device 110 outputs an audible turn indication 115b (e.g., emits a tone or series of tones) to signal the target speaker 104 to pause. The user device 110 may be enabled to output other types of turn indications 115, such as by vibrating and / or flashing a light, to inform the target speaker 104 to pause from speaking. In some implementations, the voice-to-voice S2S model 200 provides instructions 117 to cause another device (not shown) associated with the target speaker 104 to output the turn indication 115. In these implementations, a device other than the user device 110 that captured the utterance 108 is enabled to receive the instructions 117 and output the turn indication 115. For example, a smartwatch worn by the user may output the turn indication 115 by vibrating / beeping to inform the target speaker 104 to pause. When instructions 117 to cause the user device 110 (or another device) to output the turn indication 115 are provided, the synthesizer 275 may provide a time-domain audio waveform of the synthesized speech 114 for output from the device 110, 116, or from any other device.

[0038] The functionality of the voice-to-voice conversion system 100 is enabled to reside on either or both the remote server 112, the computing devices 110, 116, or any combination of the remote server and the computing devices 110, 116. The computing devices 110 and 116 may include, but are not limited to, smartphones, tablets, desktop / laptop computers, smart speakers, smart displays, smart appliances, assistant-enabled wearable devices (e.g., smart watches, smart headphones, smart glasses, etc.), or vehicle infotainment systems.

[0039] FIG. 2 illustrates an example of the speech-to-speech model 200 of FIG. 1, which includes an encoder 212, a turn detector 215, and a decoder 220. The speech-to-speech S2S model 200 processes an input audio signal 102 corresponding to an utterance 108 produced by a target speaker 104 (FIG. 1) to generate a sequence of output audio frames 222 corresponding to synthetic speech. The encoder 212 receives the input audio data 102 of the utterance 108 and is enabled to generate a high-level feature representation 213 (also referred to herein as a hidden feature representation) for each frame 102 of the sequence of audio frames 102 of the audio data 102. Each high-level feature representation 213 is generated by the encoder 212 at an output step corresponding to a frame 102 of the sequence of audio frames. The turn detector 215 is enabled to receive the high-level feature representation 213 generated by the encoder 212 at each corresponding output step to generate a turn output 216 as a single bit. That is, for each corresponding output step, the turn detector 215 outputs a turn output 216 indicating whether the corresponding audio frame 102 is at a break point. Here, the turn detector 215 may predict a single bit ('0' or '1') based on the encoder output 213 at each output step by associating a deep neural model with attention with a logistic function. In some implementations, the turn output 216 comprises a probability distribution indicating the likelihood that the encoder output 212 at each output step is at a break point. The turn detector 215 may determine the turn output 216 for the corresponding high-level feature representation 213 by analyzing each high-level feature representation 213 individually. In some implementations, the encoder 212 processes the audio data 102 with an attention mechanism to obtain one or more high-level feature representations 213 (which may be transmitted as a single vector), which the turn detector 215 can use to generate the turn output 216.Thus, the turn detector 215 may determine whether the utterance 108 is at a break point based on the current high-level feature representation 213 (i.e., the current output step) and one or more high-level feature representations 213 already generated by the encoder in the previous output step (i.e., the turn detector 215 receives the history of the encoder state to predict whether the utterance 108 is currently at a break point).

[0040] The turn detector 215 may be located between the encoder 212 and the decoder 220. Alternatively, the turn detector 215 may be part of the encoder 212. In either case, the turn detector 215 is enabled to provide a turn output 216 to the decoder 220 along with the high-order feature representation 213 output by the encoder 212 at each corresponding output step, so that the decoder 220 can generate a sequence of output audio frames 222. The decoder 220 can generate the sequence of output audio frames 222 by processing the sequence of high-order feature representations 213 received at each output step, including the output step before and including the output step corresponding to the break point. Each output audio frame in the sequence of output audio frames 222 is based on a corresponding one of the high-order feature representations 213 generated by the audio encoder 212.

[0041] In some implementations, the decoder 220 does not begin generating the sequence of output audio frames 222 until it receives a turn output 216 corresponding to a break point. In other implementations, the decoder 220 operates in a streaming mode by generating a corresponding output audio frame 222 for each high-level feature representation 213 generated by the encoder 212. Here, the synthesizer 275 may not begin synthesizing the sequence of output audio frames 212 until the turn detector 215 determines that the utterance 108 is at a break point (i.e., the synthesizer begins synthesizing output audio data 222 at the output step corresponding to the turn output 216 indicating the break point).

[0042] In some implementations, the voice-to-voice S2S model 200 includes multiple decoders 220 and selects an appropriate decoder 220 based on the speech type of the utterance 108. In these implementations, the voice-to-voice S2S model first determines the speech type and then selects the appropriate decoder 220. The voice-to-voice S2S model 200 is enabled to determine the speech type in various ways. In one implementation, the voice-to-voice S2S model 200 receives the utterance 108 and determines the speech type based on characteristics of the utterance 108 (e.g., the speech may be accented, delayed, choppy, dysarthria present, etc.). In other implementations, the voice-to-voice S2S model 200 receives the speech type from a user profile associated with the target speaker 104 (e.g., the user profile indicates that the user has a speech dysarthria). In yet another embodiment, the voice-to-voice S2S model 200 is enabled to receive an indication via user input indicating a voice type (e.g., user input to convert the voice to a particular language).

[0043] Based on the speech type, the speech-to-speech S2S model 200 is then enabled to select an appropriate decoder 220 to use in generating a sequence of output audio frames 222. For example, if the speech-to-speech S2S model 200 receives speech corresponding to a speech type of accented speech, the speech-to-speech S2S model 200 may select the streaming decoder 220. Here, the streaming decoder 220 may be suitable for synthesizing accented speech to unaccented speech when each word of the accented speech is directly synthesized into unaccented speech. However, if the speech type is translation, the streaming decoder 220 may not be suitable when there may be syntactic differences between languages ​​and when entire sentences or phrases may be required for correct translation. For example, the Turkish word "Benimle_gitti" literally translates to English as "me_with_he_went." However, the proper translation of this phrase into English would be "he went with me." Therefore, for the speech type associated with the translation, the speech-to-speech S2S model 200 is enabled to select the standard decoder 220, which processes the entire sentence. The standard decoder 220 can synthesize speech at each received break point.

[0044] 3 illustrates a training process 300 for the speech-to-speech model 200 including the turn detector 215. In some implementations, the speech-to-speech S2S model 200 is initialized / pre-trained before being further fine-tuned through training. For example, pre-training may include initiating the speech-to-speech S2S model 200 with pre-trained data 301 including a plurality of spoken utterances by a typical speaker associated with standard fluent speech. The pre-trained data 301 may further include the spoken utterances paired with corresponding synthesized ground truth standard fluent speech representations of the spoken utterances. In some implementations, the pre-trained data 301 includes a respective break point or break points for each utterance of the plurality of utterances. The pre-trained speech-to-speech S2S model is then trained with data 302 from atypical speech users, allowing the speech-to-speech S2S model 200 to be further fine-tuned for atypical speech users. In another example, the speech-to-speech S2S model 200 is pre-configured with general parameters and then fine-tuned throughout the training process 300.

[0045] The training process 300 may include training any of the audio encoder 212, the turn detector 215, or the speech decoder 220, individually or jointly in any suitable combination. The process 300 includes providing training data 302 to the speech-to-speech S2S model 200. In some implementations, the training data 302 includes training input audio data 320 associated with a plurality of training atypical speech samples produced by one or more speakers associated with atypical speech. That is, for each training atypical speech sample, the training input audio data 320 may include a corresponding training sequence of acoustic frames. In these implementations, the training data 302 also includes training output audio data 321 associated with a plurality of ground truth standard speech samples, each paired with a corresponding one of the training atypical speech samples. Here, each ground truth standard speech sample in the training output audio data 321 may include a sequence of acoustic frames corresponding to a standard fluent speech representation of the same utterance as the training input audio data 320. Additionally, the training data 302 may also include, for each of a plurality of training atypical speech samples of the input speech data 320, a sequence of break point labels 322 comprising a series of bits labeled with "1" and "0" indicating whether a corresponding acoustic frame of the input speech data 320 is at a break point. For example, the training data 302 may comprise training input audio data 320 corresponding to training atypical speech samples for utterances uttered by an atypical speech user and training output audio data 321 corresponding to a standard fluent speech representation of the same utterance in the training input audio data 320. The first labels 321 may indicate which acoustic frames of the training input audio data 320 comprise break points in the utterance (e.g., respective periods in a transcription of the utterance). In particular, the first labels 321 may further correspond to a ground truth transcription of the training input audio data 320 indicating where the break points occur.

[0046] Additionally, the training data 302 may be based on user interaction with the output of the trained voice-to-voice S2S model 200. In this manner, the voice-to-voice S2S model 200 is enabled to iteratively fine-tune based on real-world feedback. For example, the voice-to-voice S2S model 200 may receive streaming audio from a user. When the voice-to-voice S2S model 200 determines that the speech has reached a pause point, the voice-to-voice S2S model 200 provides instructions 117 that cause the user device to output a turn indication 115 informing the user to stop speaking. However, if the user continues to speak or overrides the turn indication 115 by other means, the determined pause point may be considered incorrect. The user's speech may then be used as training data. The point at which the user stops speaking is used as a second label 322 corresponding to the target pause point.

[0047] Upon receiving training data 302, the speech-to-speech S2S model 200 is enabled to generate output 350 (e.g., a sequence of output audio frames 222, a time-domain audio waveform of synthesized speech representing speech uttered by a user, and / or turn output 216). The speech-to-speech S2S model 200 is enabled to process the training data 302 in the manner described with respect to FIG. 2 above, or any other suitable method of speech-to-speech conversion.

[0048] In some implementations, the output 350 is used by one or more loss functions 305, 305a-305b to generate one or more losses 310, 310a-310c. That is, the first loss function 305a generates the first loss 310a by comparing the output 350 with corresponding ground truth standard speech samples in the training output audio data 321. The first loss 310a indicates a discrepancy between the predicted output 350 and the corresponding ground truth standard speech samples in the training output audio data 321. For example, the first loss function 305a may determine the first loss 310a by comparing the output 350 of the speech-to-speech S2S model 200 (i.e., a sequence of predicted output audio frames 222) with the training output audio data 321, which comprises a sequence of ground truth output speech frames representing a standard fluent speech representation of the training input speech data 120.

[0049] Additionally, the second loss function 305b generates a second loss 310b by comparing the output 350 with the breakpoint label 322. The second loss 310b indicates a discrepancy between the breakpoint label 322 and the predicted breakpoint location in the output 350. For example, the second loss function 305b may determine the second loss 310b by comparing the output by the speech-to-speech S2S model 200 (i.e., the turn output 216) with the breakpoint label 322 corresponding to the target breakpoint in the training data 302. In some implementations, the turn output 216 may be a bit (i.e., “0” or “1”) generated by the turn detector 215 of the speech-to-speech S2S model 200 based on a high-dimensional feature representation corresponding to each frame of the sequence of acoustic frames of the utterance, where “0” corresponds to a non-breakpoint and “1” corresponds to a breakpoint. Alternatively, the turn output 216 may be a probability distribution. In some implementations, the losses 310 may be combined to form a third loss 310c, or the losses 310 may be sent to the speech-to-speech S2S model 200 separately.

[0050] The loss function 305 may implement any suitable technique to determine the loss 310, such as regression loss, mean squared error, mean squared logarithmic error, mean absolute error, binary classification, binary cross entropy, hinge loss, multi-class loss, etc. The loss 310 may then be fed directly into the speech-to-speech S2S model 200, which may then be enabled to not only process the loss but also adjust one or more parameters to account for the loss.

[0051] 4 is a flowchart of an example operational configuration of a method 400 of a speech-to-speech model 200 with a turn detector 215. Data processing hardware 510 (FIG. 5) is enabled to perform the operations of method 400 by executing instructions stored in memory hardware 520 (FIG. 5) that are in communication with the data processing hardware 510. The data processing hardware 510 and memory hardware 520 may reside on the computing device 500 (FIG. 5), for example, on the remote server 112 and / or the user computing device 110 of FIG. 1.

[0052] At operation 402, the method 400 includes receiving, as input to a speech-to-speech (S2S) model 200, a sequence of acoustic frames 102 corresponding to an utterance 108 uttered by the user 104 in streaming audio captured by a client device 110 associated with the user 104. At each of a plurality of output steps, operation 404 of the method 400 includes generating, by an audio encoder 212 of the speech-to-speech S2S model 200, a high-level feature representation 213 of a corresponding acoustic frame 102 in the sequence of acoustic frames 102. At each of the plurality of output steps, operation 406 of the method 400 includes determining, by a turn detector 215 of the speech-to-speech S2S model 200, whether the utterance 108 is at a break point at the corresponding output step based on the high-level feature representation 213 generated by the audio encoder 212 at the corresponding output step. At operation 408, the method includes synthesizing the sequence of output audio frames 222 output by the speech decoder 220 of the speech-to-speech S2S model 200 when the turn detector 215 determines that the utterance is at a pause into a time-domain audio waveform of synthetic speech representing the utterance 108 uttered by the user 104, where each output audio frame 222 of the sequence of output audio frames 222 is based on a corresponding one of the high-level feature representations 213 generated by the audio encoder 212 up to the corresponding output step when the turn detector 215 determines that the utterance is at a pause. The sequence of output audio frames 222 corresponds to the output audio data 106.

[0053] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0054] Non-transitory memory may be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by a computing device. Non-transitory memory may be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.

[0055] 5 is a schematic diagram of an exemplary computing device 500 that may be used to implement the systems and methods described herein. Computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown here, their connections and relationships, and their functions are for illustrative purposes only and are not intended to limit the embodiments of the invention described and / or claimed in this document.

[0056] Computing device 500 includes processor 510, memory 520, storage device 530, high-speed interface / controller 540 connecting to memory 520 and high-speed expansion port 550, and low-speed interface / controller 560 connecting to low-speed bus 570 and storage device 530. Each component (510, 520, 530, 540, 550, and 560) is interconnected using various buses and may reside on a common motherboard or exist in other manners as desired. Processor 510 processes instructions for execution within computing device 500, including instructions stored in memory 520 or storage device 530, to display graphical information for a graphical user interface (GUI) on an external input / output device, such as display 580, coupled to high-speed interface 540. In other implementations, multiple processors and / or multiple buses may be used as needed, along with multiple memories and types of memories. Also, multiple computing devices 500 may be connected, each performing some of the required operations (eg, as a bank of servers, a group of blade servers, or a multi-processor system).

[0057] The memory 520 stores information non-transiently within the computing device 500. The memory 520 may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). The non-transient memory 520 may be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by the computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.

[0058] Storage device 530 can provide mass storage for computing device 500. In some implementations, storage device 530 is a computer-readable medium. In various different implementations, storage device 530 can be an array of devices, including floppy disk devices, hard disk devices, optical disk devices, or tape devices, flash memory or other similar solid-state memory devices, or devices in a storage area network or other configuration. In additional implementations, a computer program product is tangibly embodied on an information carrier. The computer program product comprises instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 520, storage device 530, or memory on processor 510.

[0059] High-speed controller 540 manages bandwidth-intensive operations for computing device 500, while low-speed controller 560 manages low-bandwidth-intensive operations. This allocation of roles is merely exemplary. In some implementations, high-speed controller 540 is coupled to memory 520, display 580 (e.g., via a graphics processor or accelerator), and high-speed expansion port 550, which can accept various expansion cards (not shown). In some implementations, low-speed controller 560 is coupled to storage device 530 and low-speed expansion port 590. Low-speed expansion port 590 may include various communication ports (USB, Bluetooth, Ethernet, wireless Ethernet, etc.) and can connect to one or more input / output devices, such as a keyboard, pointing device, scanner, or network devices, such as a switch or router, via a network adapter or the like.

[0060] The computing device 500, as shown, can be implemented in many different forms. For example, it may be implemented as a standard server 500a, or as multiple in a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.

[0061] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may be specialized or general-purpose and may include implementations in one or more computer programs executable and / or interpretable by a programmable system having at least one programmable processor, at least one input device, and at least one output device coupled to receive data and instructions from the storage system and to transmit data and instructions to the storage system.

[0062] These computer programs (also known as programs, software, software applications, or code) comprise machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor comprising a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0063] The processes and logic flows described herein can be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose processors, as well as any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from a read-only memory, a random-access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably connected to receive data from or transmit data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0064] To interact with a user, one or more aspects of the present disclosure can be implemented on a computer that optionally has a display device, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, for the user to provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic, verbal, or tactile input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.

[0065] Although several embodiments have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the disclosure. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A computer-implemented method (400) that, when executed on data processing hardware (510), causes the data processing hardware (510) to perform operations, the operations comprising: receiving, as input to a speech-to-speech S2S model (200), a sequence of acoustic frames (102) corresponding to speech (108) uttered by a user in streaming audio captured by a client device (110) associated with said user; In each of the plurality of output steps, generating, by an audio encoder (212) of the speech-to-speech S2S model (200), a high-level feature representation (213) of a corresponding acoustic frame (102) in the sequence of acoustic frames (102); determining, by a turn detector (215) of the speech-to-speech S2S model (200), whether the utterance (108) is at a break point at the corresponding output step based on the high-level feature representation (213) generated by the audio encoder (212); and synthesizing a sequence of output audio frames (222) already output by a speech decoder (220) of the speech-to-speech S2S model (200) onto a time-domain audio waveform of synthetic speech (114) representing the utterance (108) uttered by the user when the turn detector (215) determines that the utterance (108) is at the break point, wherein each output audio frame (222) of the sequence of output audio frames (222) is based on a corresponding one of the high-level feature representations (213) already generated by the audio encoder (212) up to the output step corresponding to when the turn detector (215) determines that the utterance (108) is at the break point; A computer-implemented method (400) comprising:

2. The operation further comprises: When the utterance (108) is at the interruption point, As an output from the client device (110), a speaker turn indication (115) informing the user to stop speaking; and a time-domain audio waveform of the synthesized speech (114) representing the utterance (108) uttered by the user; providing a The computer-implemented method (400) of claim 1.

3. the utterances (108) uttered by the user in the streaming audio captured by the client device (110) are associated with atypical speech; the time-domain audio waveform of the synthesized speech (114) representing the utterance (108) comprises a time-domain audio waveform of a synthesized standard fluent speech of the same utterance (108) uttered by the user; The computer-implemented method (400) of claim 1.

4. the speech (108) uttered by the user in the streaming audio captured by the client device (110) is in a first language; the time-domain audio waveform of the synthesized speech (114) representing the utterance (108) comprises a time-domain audio waveform of a synthesized translation of the same utterance (108) in a second language different from the first language; The computer-implemented method (400) of any one of claims 1 to 3.

5. determining whether the utterance (108) is at the breakpoint in the corresponding output step further comprises determining whether the utterance (108) is at the breakpoint in the corresponding output step based on one or more of the high-level feature representations (213) generated by the audio encoder (212) in an output step preceding the corresponding output step; The computer-implemented method (400) of any one of claims 1 to 3.

6. The operations further include, at each of a plurality of the output steps, generating, by the speech decoder (220) of the speech-to-speech S2S model (200), an output audio frame (222) of a corresponding high-level feature representation (213) generated by the audio encoder (212) at the corresponding output step. The computer-implemented method (400) of any one of claims 1 to 3.

7. The operation further comprises, in response to determining that the utterance (108) is at the interruption point of the corresponding output step: receiving the sequence of high-level feature representations (213) generated by the audio encoder (212) as inputs to the speech decoder (220) of the speech-to-speech S2S model (200) up to the corresponding output step; and generating a sequence of output audio frames (222) by the speech decoder (220) of the speech-to-speech S2S model (200); Equipped with The computer-implemented method (400) of any one of claims 1 to 3.

8. The turn detector (215) comprises a deep neural network. The computer-implemented method (400) of any one of claims 1 to 3.

9. The turn detector (215) is disposed between the audio encoder (212) and the speech decoder (220). The computer-implemented method (400) of any one of claims 1 to 3.

10. The step of determining whether the utterance (108) is at the break point in the corresponding output step includes generating a turn output (216) from the turn detector (215) indicating whether the utterance (108) is at the break point. The computer-implemented method (400) of any one of claims 1 to 3.

11. The turn output (216) comprises a bit. The computer-implemented method (400) of claim 10.

12. the turn output (216) comprises a probability distribution. The computer-implemented method (400) of claim 10.

13. The operations further comprise training the speech-to-speech S2S model (200), the speech-to-speech S2S model (200) comprising: receiving a set of training utterances (302), each of the training utterances (302) comprising a corresponding sequence of training acoustic frames (320), each of the training acoustic frames (320) being annotated with a label (321, 322) indicating whether the corresponding training acoustic frame (320) corresponds to a breakpoint frame or a non-breakpoint frame, and each of the training utterances (320) being paired with a corresponding ground truth synthetic speech representation of the training utterance (302); obtaining a first label (321) of the training utterance (302) indicative of a target output spectrogram; obtaining a second label (322) for the training utterance (302), the second label indicating a target turn output (216); generating training outputs (350) using the speech-to-speech S2S model (200) and the training utterances (302), the training outputs comprising training output spectrograms corresponding to synthetic speech representations of the training acoustic frames (320) and training turn outputs indicating break points of the training acoustic frames (320); determining a first loss (310) by comparing the training output spectrogram with the first label (321); determining a second loss (310) by comparing the training turn output (216) with the second label (322); and optimizing the speech-to-speech S2S model (200) based on the first loss (310) and the second loss (310) associated with training input audio data (320); Trained by The computer-implemented method (400) of any one of claims 1 to 3.

14. The operation further comprises: determining a voice type of the speech (108) uttered by the user captured in the streaming audio by the client device (110); and selecting, from among a plurality of available audio decoders, the audio decoder (220) to generate the sequence of output audio frames (222); Equipped with The computer-implemented method (400) of any one of claims 1 to 3.

15. each output audio frame (222) of the sequence of output audio frames (222) output by the speech decoder (220) comprises a spectrogram frame; The computer-implemented method (400) of any one of claims 1 to 3.

16. Data processing hardware (510); and memory hardware (520) that stores instructions that, when executed by the data processing hardware (510), cause the data processing hardware (510) to perform operations; A system (100) comprising: receiving, as input to a speech-to-speech S2S model, a sequence of acoustic frames (102) corresponding to speech (108) uttered by a user in streaming audio captured by a client device (110) associated with the user; and In each of the plurality of output steps, generating, by an audio encoder (212) of the speech-to-speech S2S model (200), a high-level feature representation (213) of a corresponding acoustic frame (102) in the sequence of acoustic frames (102); determining, by a turn detector (215) of the speech-to-speech S2S model (200), whether the utterance (108) is at a break point at the corresponding output step based on the high-level feature representation (213) generated by the audio encoder (212); and synthesizing, when the turn detector (215) determines that the utterance (108) is at the break point, a sequence of output audio frames (222) already output by a speech decoder (220) of the speech-to-speech S2S model (200) into a time-domain audio waveform of synthetic speech (114) representing the utterance (108) uttered by the user, wherein each output audio frame (222) of the sequence of output audio frames (222) is based on a corresponding one of the high-level feature representations (213) already generated by the audio encoder (212) up to the output step corresponding to when the turn detector (215) determines that the utterance (108) is at the break point; A system (100) comprising:

17. The operation further comprises: When the utterance (108) is at the interruption point, As an output from the client device (110), a speaker turn indication (115) informing the user to stop speaking; and a time-domain audio waveform of the synthesized speech (114) representing the utterance (108) uttered by the user; providing a 17. The system (100) of claim 16.

18. the utterances (108) uttered by the user in the streaming audio captured by the client device (110) are associated with atypical speech; the time-domain audio waveform of the synthesized speech (114) representing the utterance (108) comprises a time-domain audio waveform of a synthesized standard fluent speech of the same utterance (108) uttered by the user; 17. The system (100) of claim 16.

19. the speech (108) uttered by the user in the streaming audio captured by the client device (110) is in a first language; the time-domain audio waveform of the synthesized speech (114) representing the utterance (108) comprises a time-domain audio waveform of a synthesized translation of the same utterance (108) in a second language different from the first language; A system (100) according to any one of claims 16 to 18.

20. determining whether the utterance (108) is at the breakpoint in the corresponding output step further comprises determining whether the utterance (108) is at the breakpoint in the corresponding output step based on one or more of the high-level feature representations (213) generated by the audio encoder (212) in an output step preceding the corresponding output step; A system (100) according to any one of claims 16 to 18.

21. The operations further include, at each of a plurality of the output steps, generating, by the speech decoder (220) of the speech-to-speech S2S model (200), an output audio frame (222) of a corresponding high-level feature representation (213) generated by the audio encoder (212) at the corresponding output step. A system (100) according to any one of claims 16 to 18.

22. The operation further comprises, in response to determining that the utterance (108) is at the interruption point of the corresponding output step: receiving the sequence of high-level feature representations (213) generated by the audio encoder (212) as inputs to the speech decoder (220) of the speech-to-speech S2S model (200) up to the corresponding output step; and generating a sequence of output audio frames (222) by the speech decoder (220) of the speech-to-speech S2S model (200); Equipped with A system (100) according to any one of claims 16 to 18.

23. The turn detector (215) comprises a deep neural network. A system (100) according to any one of claims 16 to 18.

24. The turn detector (215) is disposed between the audio encoder (212) and the speech decoder (220). A system (100) according to any one of claims 16 to 18.

25. The step of determining whether the utterance (108) is at the break point in the corresponding output step includes generating a turn output (216) from the turn detector (215) indicating whether the utterance (108) is at the break point. A system (100) according to any one of claims 16 to 18.

26. The turn output (216) comprises a bit.

26. The system (100) of claim 25.

27. the turn output (216) comprises a probability distribution.

26. The system (100) of claim 25.

28. The operations further comprise training the speech-to-speech S2S model (200), the speech-to-speech S2S model (200) comprising: receiving a set of training utterances (302), each of the training utterances (302) comprising a corresponding sequence of training acoustic frames (320), each of the training acoustic frames (320) being annotated with a label (321, 322) indicating whether the corresponding training acoustic frame (320) corresponds to a breakpoint frame or a non-breakpoint frame, and each of the training utterances (302) being paired with a corresponding ground truth synthetic speech representation of the training utterance (302); obtaining a first label (321) of the training utterance (302) indicative of a target output spectrogram; obtaining a second label (322) for the training utterance (302), the second label indicating a target turn output (216); generating training outputs (350) using the speech-to-speech S2S model (200) and the training utterances (302), the training outputs comprising training output spectrograms corresponding to synthetic speech representations of the training acoustic frames (320) and training turn outputs indicating break points of the training acoustic frames (320); determining a first loss (310) by comparing the training output spectrogram with the first label (321); determining a second loss (310) by comparing the training turn output (216) with the second label (322); and optimizing the speech-to-speech S2S model (200) based on the first loss (310) and the second loss (310) associated with training input audio data (320); Trained by A system (100) according to any one of claims 16 to 18.

29. The operation further comprises: determining a voice type of the speech (108) uttered by the user captured in the streaming audio by the client device (110); and selecting, from among a plurality of available audio decoders, the audio decoder (220) to generate the sequence of output audio frames (222); Equipped with A system (100) according to any one of claims 16 to 18.

30. each output audio frame (222) of the sequence of output audio frames (222) output by the speech decoder (220) comprises a spectrogram frame; A system (100) according to any one of claims 16 to 18.

Citation Information

Patent Citations

  • Turn-taking timing identification apparatus, turn-taking timing identification method, program and recording medium

    JP2018132678A

  • System and method for end-to-end speech recognition with trigger door tension

    JP2022522379A

  • System and method for direct speech translation system

    US20200226327A1

  • Synthesized data augmentation using voice conversion and speech recognition models

    WO2022046526A1