Robust direct speech-to-speech translation
The direct speech-to-speech translation model addresses the challenges of translation quality and robustness by using an encoder, attention module, decoder, and synthesizer to preserve the speaking style and prosody of the source speaker, achieving results comparable to cascaded systems.
Patent Information
- Application Number
- JP2024502159
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-07-16
- Filing Date
- 2021-12-15
- Publication Date
- 2025-05-22
- Estimated Expiration
- 2041-12-15
AI Technical Summary
Existing direct speech-to-speech translation systems face challenges in translation quality, speech naturalness, and robustness, particularly due to issues with attention-based techniques that affect the preservation of paralinguistic and non-linguistic information such as speaker voice and prosody.
The development of a direct speech-to-speech translation model that includes an encoder, an attention module, a decoder, and a synthesizer, which is trained end-to-end to directly convert input speech from one language to output speech in another language, while preserving the speaking style and prosody of the source speaker.
This approach results in improved translation quality, naturalness, and robustness, comparable to cascaded systems, while also preserving the voice and prosody of the source speaker, thus mitigating the risk of creating spoofing audio artifacts.
Smart Images

Figure 0007681793000002 
Figure 0007681793000003 
Figure 0007681793000004
Abstract
Description
[Technical field]
[0001] This disclosure relates to robust direct speech-to-speech translation. [Background technology]
[0002] Speech-to-speech translation (S2ST) is highly beneficial for breaking down communication barriers between people who do not share a common language. Traditionally, S2ST systems consist of a cascade of three components: automatic speech recognition (ASR), text-to-text machine translation (MT), and text-to-speech (TTS) synthesis. Recently, advances in direct speech-to-text translation (ST) have outpaced the ASR and MT cascades, thereby making the two-component cascade of ST and TTS feasible as S2ST. Summary of the Invention [Means for solving the problem]
[0003] One aspect of the disclosure provides a direct speech-to-speech translation (S2ST) model that includes an encoder configured to receive an input speech representation corresponding to an utterance spoken in a first language by a source speaker and to encode the input speech representation into a hidden feature representation. The S2ST model also includes an attention module configured to generate a context vector that attends the hidden representation encoded by the encoder. The S2ST model also includes a decoder configured to receive the context vector generated by the attention module and to predict a phoneme representation corresponding to a translation of the utterance in a second, different language. The S2ST model also includes a synthesizer configured to receive the context vector and the phoneme representation and to generate a translated synthetic speech representation corresponding to the translation of the utterance spoken in the second, different language.
[0004] Implementations of the present disclosure may include one or more of any of the following features: In some implementations, the encoder includes a stack of conformer blocks. In other implementations, the encoder includes a stack of one of transformer blocks or lightweight convolutional blocks. In some examples, the synthesizer includes a duration model network configured to predict the duration of each phoneme in the sequence of phonemes represented by the phoneme representation. In these examples, the synthesizer may be configured to generate the translated synthetic speech representation by upsampling the sequence of phonemes based on the predicted duration of each phoneme. The translated synthetic speech representation may be configured to match the speaking style / prosody of the source speaker.
[0005] In some implementations, the S2ST model is trained on parallel source and target language utterance pairs, each including a voice spoken in the source utterance. In these implementations, at least one of the source language utterance or the target language utterance includes a voice synthesized by a text-to-speech model trained to generate a synthetic voice of the source utterance. In some examples, the S2ST model further includes a vocoder configured to receive the translated synthetic speech representation and synthesize the translated synthetic speech representation into an audible output of the translated synthetic speech representation. In some cases, the phoneme representation may include a probability distribution of possible phonemes in a phoneme sequence corresponding to the translated synthetic speech representation.
[0006] Another aspect of the disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations for direct speech-to-speech translation. The operations include receiving an input speech representation corresponding to an utterance spoken in a first language by a source speaker as an input to a direct speech-to-speech translation (S2ST) model. The operations also include encoding the input speech representation into a hidden feature representation by an encoder of the S2ST model. The operations also include generating, by a decoder of the S2ST model, a context vector that directs attention to the hidden feature representation encoded by the encoder. The operations also include receiving, at the decoder of the S2ST model, the context vector generated by the attention module. The operations also include predicting, by the decoder, a phoneme representation corresponding to a translation of the utterance in a second, different language. The operations also include receiving, at a synthesizer of the S2ST model, the context vector and the phoneme representation. The operations also include generating, by the synthesizer, a translated synthetic speech representation corresponding to the translation of the utterance spoken in the second, different language.
[0007] Implementations of the present disclosure may include one or more of any of the following features: In some implementations, the encoder includes a stack of conformer blocks. In other implementations, the encoder includes a stack of one of transformer blocks or lightweight convolutional blocks. In some examples, the synthesizer includes a duration model network configured to predict the duration of each phoneme in the sequence of phonemes represented by the phoneme representation. In these examples, generating the translated synthetic speech representation may include upsampling the sequence of phonemes based on the predicted duration of each phoneme.
[0008] The translated synthetic speech representation may be configured to match the speaking style / prosody of the source speaker. In some implementations, the S2ST model is trained on parallel source and target language utterance pairs, each including a voice spoken in the source utterance. In these implementations, at least one of the source language utterance or the target language utterance may include speech synthesized by a text-to-speech model trained to generate synthetic speech of the voice of the source utterance. In some examples, the operations further include receiving the translated synthetic speech representation at a vocoder of the S2ST model and synthesizing the translated synthetic speech representation by the vocoder into an audible output of the translated synthetic speech representation. In some cases, the phoneme representation may include a probability distribution of possible phonemes in a phoneme sequence corresponding to the translated synthetic speech representation.
[0009] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief description of the drawings]
[0010] [Figure 1] FIG. 1 is a schematic diagram of an exemplary speech environment including a direct speech-to-speech translation (S2ST) model. [Diagram 2] FIG. 1 is a schematic diagram of the S2ST model. [Diagram 3] FIG. 1 is a schematic diagram of the combiner of the S2ST model. [Figure 4] FIG. 2 is a schematic diagram of an exemplary Conformer block. [Diagram 5] 1 is a flowchart of an exemplary arrangement of operations for a method of performing direct speech-to-speech translation. [Figure 6] FIG. 1 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] Like reference numbers in the various drawings indicate like elements.
[0012] Speech-to-speech translation (S2ST) is highly beneficial for breaking down communication barriers between people who do not share a common language. Traditionally, S2ST systems consist of a cascade of three components: automatic speech recognition (ASR), text-to-text machine translation (MT), and text-to-speech (TTS) synthesis. Recently, advances in direct speech-to-text translation (ST) have outpaced the ASR and MT cascades, thereby making the two-component cascade of ST and TTS feasible as S2ST.
[0013] Direct S2ST involves directly translating speech in one language into speech in another language. In other words, a direct S2ST system / model is configured to convert an input audio waveform or spectrogram corresponding to speech spoken in a first language by a human speaker into an output audio waveform or spectrogram corresponding to synthesized speech in a second language different from the first language, without converting the input audio waveform into an intermediate representation (e.g., text or phonemes). As will become apparent, direct S2ST models, as well as techniques for training direct S2ST models, enable a user to speak in the user's native language and be understood by both other humans and speech interfaces (e.g., digital assistants) by making the user's speech recognizable and / or playable as synthetic audio in a different language. Recent direct S2ST models have fallen short of cascaded S2ST systems in terms of translation quality, while also having issues with robustness of the output synthetic speech in terms of bubbling and long interruptions. These issues stem from the use of attention-based techniques to synthesize speech.
[0014] Implementations herein are directed to robust direct S2ST models that are trained end-to-end, outperform existing direct S2ST systems, and are comparable to cascaded systems in terms of translation quality, speech naturalness, and speech robustness. In particular, compared to cascaded systems, direct S2ST systems / models have the potential to preserve paralinguistic and non-linguistic information such as speaker voice and prosody during translation, operate on language without written morphology, reduce computational requirements and inference latency, avoid error compounding across subsystems, and facilitate processing of content that does not need to be translated, such as names and other proper nouns. Implementations herein are further directed to voice preservation techniques in S2ST that do not rely on explicit speaker embeddings or identifiers. In particular, the trained S2ST model is trained to simply preserve the voice of the source speaker provided in the input speech, without the ability to generate speech in a voice different from the source speaker. In particular, the ability to preserve the source speaker's voice is useful in production environments by proactively mitigating exploits to create spoofing audio artifacts.
[0015] 1 illustrates a speech conversation environment 100 employing a direct speech-to-speech translation (S2ST) model 200 configured to directly translate an input utterance spoken in a first language by a source speaker into a corresponding output utterance in a different second language, and vice versa. As will become apparent, the direct S2ST model 200 is trained end-to-end. In particular, the direct S2ST model 200 differs from cascaded S2ST systems that employ an automatic speech recognizer (ASR) component, a text-to-text machine translation (MT) component, and a text-to-speech (TTS) synthesis component, or other cascaded S2ST systems that employ a cascade of a direct speech-to-text translation (ST) component followed by a TTS component.
[0016] In the illustrated example, the direct S2ST model 200 is configured to convert input audio data 102 corresponding to utterance 108 spoken in a first / source language (e.g., Spanish) by source speaker 104 into output audio data (e.g., a mel spectrogram) 106 corresponding to a translated synthetic speech representation of the translated utterance 114 spoken in a different second language (e.g., English) by source speaker 104. The direct S2ST model 200 may directly convert an input spectrogram corresponding to the input audio data 102 into an output spectrogram corresponding to the output audio data 106 without performing speech recognition and machine translation between texts or, in some cases, without requiring the generation of any intermediate discrete representation (e.g., text or phonemes) from the input data 102. As will be described in more detail with respect to FIGS. 2 and 3, the direct S2ST model 200 includes a spectrogram encoder 210, an attention module 220, a decoder 230, and a synthesizer (e.g., a spectrogram decoder) 300.
[0017] Vocoder 375 may synthesize the output audio data 106 output from the direct S2ST model 200 into a time-domain waveform for audible output as the translated utterance 114 spoken in the second language and in the voice of the source speaker. The time-domain audio waveform includes an audio waveform that defines the amplitude of the audio signal over time. Instead of vocoder 375, a unit selection module or a WaveNet module may synthesize the output audio data 106 into a time-domain waveform of the synthesized speech in the translated second language and in the voice of the source speaker 104. In some implementations, vocoder 375 includes a vocoder network, i.e., a neural vocoder, that is separately trained and tuned on a mel-frequency spectrogram for conversion to the time-domain audio waveform.
[0018] In the illustrated example, source speaker 104 is a native speaker of Spanish, the first / source language. The direct S2ST 200 is thus trained to directly convert the input audio data 102 corresponding to the utterance 108 spoken in Spanish by source speaker 104 into output audio data 106 corresponding to a translated synthetic speech representation corresponding to the translated utterance 114 in English (e.g., the second / target language). That is, the translated utterance 114 in English (e.g., "Hi, what are your plans this afternoon?") includes the synthetic audio of the translated version of the input utterance 108 (e.g., "Hola, cuales son tus planes esta tarde?") spoken in Spanish by source speaker 104. In this way, the translated synthetic representation provided in English by the output audio data 106 enables a native Spanish speaker to convey the utterance 108 spoken in Spanish to the receiving user 118 who speaks English as a native language. In some examples, source speaker 104 does not speak English and the receiving speaker 118 does not speak / understand Spanish. In some implementations, the direct S2ST model 200 is multilingual and is trained to also convert an input utterance spoken in English by speaker 118 into a translated utterance in Spanish. In these implementations, the direct S2ST model 200 may be configured to convert audio between one or more other pairs of languages in addition to, or instead of, Spanish and English.
[0019] In particular, the direct S2ST model 200 is trained to preserve the characteristics of the source speaker's voice such that the synthetic speech representation and the resulting output audio data 106 corresponding to the translated utterance 114 convey the voice of the source speaker, but in a different second language. In other words, the translated utterance 114 conveys the characteristics of the source speaker's voice (e.g., speaking style / prosody) as if the source speaker 104 actually spoke the different second language. In some examples, as described in more detail below, the direct S2ST model 200 is trained not only to preserve the characteristics of the source speaker's voice in the output audio data 106, but also to hinder the ability to generate speech in a voice different from the source speaker to mitigate misuse of the model 200 to create spoofing audio artifacts.
[0020] A computing device associated with the source speaker 104 may capture utterances 108 spoken by the source speaker 104 in a source / first language (e.g., Spanish) and send corresponding input audio data 102 to the direct S2ST model 200 for conversion to output audio data 106. The direct S2ST model 200 may then send output audio data 106 corresponding to a translated synthetic speech representation of the translated utterance 114 to another computing device 116 associated with a receiving user 118, which then audibly outputs the translated synthetic speech representation as the translated utterance 114 in a different second language (e.g., English). In this example, the source speaker 104 and the user 118 are talking to each other via their own computing devices 110, 116, respectively, through an audio / video call (e.g., video conference / chat), a telephone call, or other type of voice communication protocol, such as voice over Internet Protocol.
[0021] In particular, the direct S2ST model 200 may be trained to preserve the same speaking style / prosody in the output audio data 106 corresponding to the translated synthetic speech representation that was used in the input audio data 102 corresponding to the utterance 108 spoken by the source speaker 104. For example, in the illustrated example, the input audio data 102 for the Spanish utterance 108 conveys a style / prosody associated with asking a question, and therefore the S2ST model 200 generates the output audio data 106 corresponding to the translated synthetic speech representation having a style / prosody associated with asking a question.
[0022] In some other examples, the S2ST conversion model 200 instead sends the output audio data 106 corresponding to the translated synthetic speech representation of the utterance spoken by the source speaker 104 to an output audio device for audibly outputting the translated synthetic speech representation to an audience in the voice of the source speaker 104. For example, the source speaker 104 who speaks Spanish natively may be a lecturer giving a lecture to an English-speaking audience, in which case the utterance spoken in Spanish by the source speaker 104 is converted into the translated synthetic speech representation that is audibly output from the audio device to the English-speaking audience as the translated speech in English.
[0023] Alternatively, the other computing device 116 may be associated with a downstream automatic speech recognition (ASR) system in which the S2ST model 200 serves as a front end to provide output audio data 106 corresponding to the synthetic speech representation as input to the ASR system for conversion to recognized text. The recognized text may be presented to another user 118 and / or provided to a natural language understanding (NLU) system for further processing.
[0024] The functionality of the S2ST model 200 may reside on the remote server 112, or on either or both of the computing devices 110, 116, or on any combination of the remote server and the computing devices 110, 116. In particular, the data processing hardware of the computing devices 110, 116 may execute the S2ST model 200. In some implementations, the S2ST model 200 continuously generates output audio data 106 corresponding to a synthetic speech representation of the utterance as the source speaker 104 speaks the corresponding portion of the utterance in the first / source language. By continuously generating output audio data 106 corresponding to a synthetic speech representation of a portion of the utterance 108 spoken by the source speaker 104, the conversation between the source speaker 104 and the user 118 (or audience) may be more naturally paced. In some further implementations, the S2ST model 200 waits to determine / detect when the source speaker 104 stops speaking, using techniques such as voice activity detection, end pointing, end of query detection, etc., before converting the corresponding input audio data 102 of the utterance 108 in a first language into corresponding output audio data 106 corresponding to a translated synthetic speech representation of the same utterance 114 but in a different second language.
[0025] FIG. 2 illustrates the direct S2ST model 200 of FIG. 1, including an encoder 210, an attention module 220, a decoder 230, and a synthesizer 300. The encoder 210 is configured to encode the input audio data 102 into a hidden feature representation (e.g., a series of vectors) 215, where the input audio data 102 includes a sequence of input spectrograms corresponding to an utterance 108 spoken in a source / first language (e.g., Spanish) by a source speaker 104. The sequence of input phonemes may include an 80-channel mel spectrogram sequence. In some implementations, the encoder 210 includes a stack of conformer layers. In these implementations, the encoder subsamples the input audio data 102 including the input mel spectrogram sequence using a convolutional layer, and then processes the input mel spectrogram sequence with a stack of conformer blocks. Each conformer block may include a feedforward layer, a self-attention layer, a convolutional layer, and a second feedforward layer. In some examples, the stack of conformer blocks includes 16 layers of conformer blocks with 144 dimensions and a subsampling factor of four (4). Figure 4 is a schematic diagram of an example conformer block. The encoder 210 may use a stack of transformer blocks or lightweight convolution blocks instead of the conformer blocks.
[0026] The attention module 220 is configured to generate a context vector 225 that directs attention to the hidden feature representation 215 encoded by the encoder 210. The attention module 220 may include a multi-head attention mechanism. The decoder 230 is configured to receive as an input the context vector 225 indicating the hidden feature representation 215 as a source value of attention, and to predict as an output a phoneme representation 235 that represents a probability distribution of possible phonemes in a phoneme sequence 245 that corresponds to the audio data (e.g., a translated synthetic speech representation of a target) 106. That is, the phoneme representation 235 corresponds to a translation of the utterance 108 in a second different utterance (e.g., in a second language). A fully connected network plus softmax 240 layer may select a phoneme in the sequence of phonemes (e.g., English phonemes) 245 based on using the phoneme with the highest probability in the probability distribution of possible phonemes represented by the phoneme representation 235 at each of a plurality of output steps. In the illustrated example, the decoder 230 is autoregressive and at each output step generates a probability distribution of possible phonemes for a given output step based on each previous phoneme in the phoneme sequence 245 selected by Softmax 240 during each of the previous output steps. In some implementations, the decoder 230 includes a stack of long short-term memory (LSTM) cells aided by an attention module 220. Notably, the combination of the encoder 210, attention module 220, and decoder 230 is similar to direct speech-to-text translation (ST) components commonly found in cascaded S2ST systems.
[0027] The synthesizer 300 receives as input during each of a plurality of output steps a concatenation of the phoneme representation 235 (or phoneme sequence 245) and the context vector 225 at the corresponding output step, and generates as output during each of a plurality of output steps an output audio data 106 corresponding to a translated synthetic speech representation in the target / second language and in the voice of the source speaker 104. Alternatively, the synthesizer 300 may receive the phoneme representation 235 and the context vector 225 (e.g., without concatenation). The synthesizer 300 may also be referred to as a spectrogram decoder. In some examples, the synthesizer is an autoregressive type, where each predicted output spectrogram is based on a sequence of previously predicted spectrograms. In other examples, the synthesizer 300 is a parallel and non-autoregressive type.
[0028] 3 illustrates an example of the synthesizer 300 of FIG. 1. Here, the synthesizer 300 may include a phoneme duration modeling network (i.e., duration predictor) 310, an upsampler module 320, a recurrent neural network (RNN) 330, and a convolutional layer 340. The duration modeling network receives as input the phoneme representation 235 from the decoder 230 and the context vector 225 from the attention module 220. Furthermore, the duration modeling network 310 is tasked with predicting the duration 315 for each phoneme in the phoneme representation 235 corresponding to the output audio data 106 representing the translated synthetic speech representation in the target / second language. During training, the individual target durations 315 for each phoneme are unknown, and thus the duration model network 310 determines the target mean duration based on the ratio of the total frame duration T of the entire reference mel-frequency spectrogram sequence to the total number K of phonemes (e.g., tokens) in the reference phoneme sequence that corresponds to the reference mel-frequency spectrogram sequence. That is, the target mean duration is the mean duration for all phonemes using the reference mel-frequency spectrogram sequence and the reference phoneme sequence used during training. During training, a loss term (e.g., an L2 loss term) is determined between the predicted phoneme duration and the target mean duration. Thus, the duration model network 310 learns to predict phoneme durations in an unsupervised manner without using supervised phoneme duration labels provided from an external aligner. While an external aligner can provide reasonable alignment between phonemes and mel-spectrum frames, phoneme duration rounding is required by the length regulator to upsample the phonemes in the reference phoneme sequence according to their duration, which leads to rounding errors that can persist. In some cases, using supervised duration labels from an external aligner during training and predicted durations during inference creates a mismatch in phoneme durations between the training of S2ST model 200 and the inference of S2ST model 200.Furthermore, such rounding operations are not differentiable, and thus the error gradient cannot propagate through the duration model network.
[0029] The upsampler 320 receives the predicted duration 315, the context vector 225, and the phoneme representation as inputs and generates an output 235. Specifically, the upsampler 320 is configured to upsample an input sequence (e.g., the phoneme representation 235 or the phoneme sequence 245) based on the predicted duration 315 from the duration model network 310. The RNN 330 receives the output 335 and is configured to autoregressively predict the target mel spectrogram 335 corresponding to the audio data 106 (e.g., the translated synthetic speech representation of the target / the target in the second language). The RNN 330 provides the target mel spectrogram 335 to the convolutional layer 340 and the concatenator 350. The convolutional layer 340 provides a residual convolutional post-net configured to further improve the target mel spectrogram 335 and generate an output 345. That is, the convolutional layer 340 further improves the predicted translated synthetic speech representation in the second language. The concatenator 350 concatenates the output 345 and the target mel spectrogram 335 to generate a translated synthetic speech representation 355 corresponding to the translation of the utterance 108 spoken in a different second language. Thus, the translated synthetic speech representation 355 may correspond to the audio data 106 (FIG. 2). Notably, the translated synthetic speech representation 355 retains the speaking style / rhythm of the source speaker 104.
[0030] Implementations herein are further directed to a voice preservation technique that restricts the trained S2ST model 200 to only preserve the source speaker's voice and not generate synthetic speech in a different speaker's voice. This technique involves training on parallel utterances with the same speaker's voice in both the input utterances in the first language and the output utterances in the second language. Since fluent bilingual speakers are rare, a cross-lingual TTS model may be employed to synthesize training utterances in the target second language that include the source speaker's voice. Thus, the S2ST model 200 may be trained using utterances from the source speaker 104 in the first language and the synthesized training utterances of the source speaker 104 in the target second language. The S2ST model 200 may be further trained to preserve the source speaker's voice in the translated synthetic speech for each source speaker during the speaker's turn.
[0031] FIG. 4 shows an example of a Conformer block 400 from a stack of Conformer layers of the encoder 210. The Conformer block 400 includes a first half-feedforward layer 410, a second half-feedforward layer 440, a multi-head self-attention block 420 and a convolutional layer 430 disposed between the first and second half-feedforward layers 410, 440, and a concatenation operator 405. The first half-feedforward layer 410 processes the input audio data 102 including the input mel spectrogram sequence. The multi-head self-attention block 420 then receives the input audio data 102 concatenated with the output of the first half-feedforward layer 410. Intuitively, the role of the multi-head self-attention block 420 is to aggregate noise context separately for each input frame to be enhanced. The convolutional layer 430 subsamples the output of the multi-head self-attention block 420 concatenated with the output of the first half-feedforward layer 410. A second half-feedforward layer 440 then receives the concatenation of the convolutional layer 430 output and the multi-head self-attention block 420. A layernorm module 450 processes the output from the second half-feedforward layer 440. Mathematically, the conformer block 400 transforms the input features x using the modulation features m to generate output features y as follows:
[0032]
number
[0033] 5 is a flow chart of an exemplary arrangement of operations for a computer-implemented method 500 for performing direct speech-to-speech translation. At operation 502, the method 500 includes receiving an input speech representation 102 corresponding to an utterance 108 spoken in a first language by a source speaker 104. At operation 504, the method 500 includes an encoder 210 of the S2ST model 200 encoding the input speech representation 102 into a hidden feature representation 215. At operation 506, the method 500 includes an attention module 220 of the S2ST model 200 generating a context vector 225 that directs attention to the hidden feature representation 215 encoded by the encoder 210. At operation 508, the method 500 includes receiving the context vector 225 at a decoder 230 of the S2ST model 200. At operation 510, the method 500 includes the decoder 230 predicting a phoneme representation 235 corresponding to a translation of the utterance 108 in the second, different language. At operation 512, the method 500 includes receiving the context vector 225 and the phoneme representation 235 at a synthesizer 300 of the S2ST model 200. At operation 514, the method 500 includes generating, by the synthesizer 300, a translated phonetic representation 355 corresponding to the translation of the utterance 108 spoken in the second, different language.
[0034] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an "application," an "app," or a "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0035] Non-transient memory may be a physical device used to temporarily or permanently store programs (e.g., a sequence of instructions) or data (e.g., program state information) for use by a computing device. Non-transient memory may be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0036] 6 is a schematic diagram of an exemplary computing device 600 that may be used to implement the systems and methods described herein. The computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown, their connections and relationships, and their functions are intended to be exemplary only and are not intended to limit the implementation of the invention described and / or claimed herein.
[0037] Computing device 600 includes a processor 610, a memory 620, a storage device 630, a high-speed interface / controller 640 connecting to memory 620 and high-speed expansion port 650, and a low-speed interface / controller 660 connecting to low-speed bus 670 and storage device 630. Each of components 610, 620, 630, 640, 650, and 660 may be interconnected using various buses and mounted on a common motherboard or otherwise as desired. Processor 610 can process instructions for execution within computing device 600, including instructions stored in memory 620 or storage device 630, to display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 680 coupled to high-speed interface 640. In other implementations, multiple processors and / or multiple buses may be used in conjunction with multiple memories and types of memories as desired. Also, multiple computing devices 600 may be connected, with each device providing a portion of the required operations (eg, as a bank of servers, a group of blade servers, or a multi-processor system).
[0038] The memory 620 stores information non-transiently within the computing device 600. The memory 620 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-volatile memory 620 may be a physical device used to temporarily or permanently store programs (e.g., a sequence of instructions) or data (e.g., program state information) for use by the computing device 600. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), as well as disks or tapes.
[0039] The storage device 630 is capable of providing mass storage to the computing device 600. In some implementations, the storage device 630 is a computer-readable medium. In various different implementations, the storage device 630 may be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, a flash memory or other similar solid-state memory device, or a storage area network or other configuration of devices. In further implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods as described above. The information carrier is a computer-readable or machine-readable medium, such as the memory 620, the storage device 630, or a memory on the processor 610.
[0040] The high-speed controller 640 manages the bandwidth-intensive operations of the computing device 600, and the low-speed controller 660 manages the lower bandwidth-intensive operations. This allocation of duties is merely exemplary. In some implementations, the high-speed controller 640 is coupled to the memory 620, the display 680 (e.g., via a graphics processor or accelerator), and is also coupled to a high-speed expansion port 650 that may accept various expansion cards (not shown). In some implementations, the low-speed controller 660 is coupled to the storage device 630 and the low-speed expansion port 690. The low-speed expansion port 690 may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet) and may be coupled to one or more input / output devices such as a keyboard, pointing device, scanner, or to a networking device such as a switch or router, for example via a network adapter.
[0041] The computing device 600 may be implemented in a number of different forms, as shown in the figure. For example, the computing device 600 may be implemented as a standard server 600a, or multiple times in a group of such servers 600a, or as a laptop computer 600b, or as part of a rack server system 600c.
[0042] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable by a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0043] These computer programs (also known as programs, software, software applications, or codes) contain machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0044] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by dedicated logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general purpose and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions, and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices, such as magnetic disks, magneto-optical disks, or optical disks, for storing data, or be operatively coupled to receive data from or transfer data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0045] To enable interaction with a user, one or more aspects of the present disclosure may be implemented in a computer having a display device, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to enable interaction with the user, for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic input, speech input, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0046] Several implementations have been described. However, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims. [Explanation of symbols]
[0047] 100 Voice conversation environment 102 Input Audio Data 104 Source Speaker 106 Output Audio Data 108 utterances 110 Computing Devices 112 Remote Server 114 Translated utterances 116 Computing Devices 118 Receiving User 200 Direct Speech-to-Speech Translation (S2ST) Model 210 Encoder 215 Hidden Feature Representation 220 Attention Module 225 Context Vector 230 Decoder 235 Phoneme expression 245 Phoneme string 300 Synthesizer 310 Duration Model Network 315 Duration 320 Upsampler 330 Recurrent Neural Networks (RNN) 335 Target Mel Spectrogram 340 Convolutional Layer 345 Output 350 coupler 355 translated synthetic speech expressions 375 Vocoder 600 computing devices 610 Processor 620 Memory 630 Storage Devices 640 High Speed Interface / Controller 650 High Speed Expansion Port 660 Low Speed Interface / Controller 670 Slow Bus 680 Display 690 Low Speed Expansion Port
Claims
1. A program for causing a computer to execute the functions of a direct speech-to-speech translation (S2ST) model (200), the direct S2ST model (200) comprising: An encoder (210), receiving an input speech representation (102) corresponding to an utterance (108) spoken in a first language by a source speaker (104); and encoding said input speech representation (102) into a hidden feature representation (215); An encoder (210) configured to: an attention module (220) configured to generate a context vector (225) that directs attention to the hidden feature representation (215) encoded by the encoder (210); A decoder (230), receiving the context vector (225) generated by the attention module (220); and Predicting a phonemic representation (235) corresponding to a translation of said utterance (108) in a second, different language. A decoder (230) configured to: A combiner (300), receiving said context vector (225) and said phoneme representation (235); and generating a translated synthetic speech representation (355) corresponding to the translation of the utterance (108) spoken in the different second language. A combiner (300) configured to perform A program that includes:
2. The program of claim 1 , wherein the encoder (210) comprises a stack of conformer blocks (400).
3. 3. The program of claim 1, wherein the encoder comprises a stack of one of transformer blocks or lightweight convolution blocks.
4. 4. The program of claim 1, wherein the synthesizer (300) comprises a duration model network (310) configured to predict a duration (315) of each phoneme in a sequence of phonemes represented by the phoneme representation (235).
5. 5. The program of claim 4, wherein the synthesizer (300) is configured to generate the translated synthetic speech representation (355) by upsampling the sequence of phonemes based on the predicted duration (315) of each phoneme.
6. 6. The program of claim 1, wherein the translated synthetic speech representation (355) is adapted to the speaking style / prosody of the source speaker (104).
7. the direct S2ST model (200) is trained on parallel source language and target language utterance pairs; 7. The program of claim 1, wherein each pair comprises a voice spoken in the source language utterance.
8. 8. The program of claim 7, wherein at least one of the source language utterance (108) or the target language utterance comprises speech synthesized by a text-to-speech model trained to generate a synthetic speech of the voice of the source language utterance (108).
9. Vocoder (375) receiving the translated synthesized speech representation (355); synthesizing the translated synthetic speech representation (355) into an audible output of the translated synthetic speech representation (355); The program according to any one of claims 1 to 8, configured to:
10. 10. The program of claim 1, wherein the phoneme representation (235) comprises a probability distribution of possible phonemes in a phoneme sequence corresponding to the translated synthetic speech representation (355).
11. A method for generating a speech-to-speech translation (S2ST) model, comprising: receiving, as an input to a direct speech-to-speech translation (S2ST) model (200), an input speech representation (102) corresponding to an utterance (108) spoken in a first language by a source speaker (104); encoding the input speech representation (102) into a hidden feature representation (215) by an encoder (210) of the direct S2ST model (200); generating a context vector (225) by an attention module (220) of the direct S2ST model (200) that directs attention to the hidden feature representation (215) encoded by the encoder (210); receiving the context vector (225) generated by the attention module (220) at a decoder (230) of the direct S2ST model (200); predicting, by said decoder (230), a phoneme representation (235) corresponding to a translation of said utterance in a second, different language; receiving said context vector (225) and said phoneme representation (235) at a synthesizer (300) of said direct S2ST model (200); generating, by said synthesizer (300), a translated synthetic speech representation (355) corresponding to said translation of said utterance spoken in said different second language; A computer-implemented method (500).
12. The computer-implemented method of claim 11 , wherein the encoder comprises a stack of conformer blocks.
13. 13. The computer-implemented method of claim 11 or 12, wherein the encoder comprises a stack of one of transformer blocks or lightweight convolution blocks.
14. 14. The computer-implemented method (500) of any one of claims 11 to 13, wherein the synthesizer (300) includes a duration model network (310) configured to predict the duration (315) of each phoneme in a sequence of phonemes represented by the phoneme representation (235).
15. 15. The computer-implemented method (500) of claim 14, wherein generating the translated synthetic speech representation (355) comprises upsampling the sequence of phonemes based on the predicted duration (315) of each phoneme.
16. 16. The computer-implemented method (500) of any one of claims 11 to 15, wherein the translated synthetic speech representation (355) is adapted to the speaking style / prosody of the source speaker (104).
17. the direct S2ST model (200) is trained on parallel source language and target language utterance pairs; 17. The computer-implemented method (500) of any one of claims 11 to 16, wherein each pair comprises a voice spoken in the source language utterance (108).
18. 20. The computer-implemented method (500) of claim 17, wherein at least one of the source language utterance (108) or the target language utterance comprises speech synthesized by a text-to-speech model trained to generate a synthetic speech of the voice of the source language utterance (108). receiving the translated synthetic speech representation (355) at a vocoder (375); synthesizing the translated synthetic speech representation (355) into an audible output of the translated synthetic speech representation (355) by the vocoder (375); 19. The computer-implemented method (500) of any one of claims 11 to 18, further comprising:
20. 20. The computer-implemented method (500) of any one of claims 11 to 19, wherein the phoneme representation (235) comprises a probability distribution of possible phonemes in a phoneme sequence corresponding to the translated synthetic speech representation (355).
Citation Information
Patent Citations
Speech translation method and system using a multilingual text-to-speech synthesis model
JP2021511534A
System and method for direct speech translation system
US20200226327A1
Direct Speech-to-Speech Translation via Machine Learning
US20210209315A1