METHOD AND DEVICE FOR LANGUAGE SYNTHESIS FOR MULTIPLE LANGUAGES AND MULTIPLE SPEAKERS

The speech synthesis model addresses the limitations of conventional methods by removing speaker and speech information during training and adding it during inference, ensuring accurate and natural-sounding speech synthesis for multiple speakers and languages, particularly in vehicle environments.

DE102025116409A1Pending Publication Date: 2026-03-05HYUNDAI MOTOR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025116409
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-28
Filing Date
2025-04-29
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Conventional speech synthesis methods struggle to generate natural-sounding speech for untrained speakers and multiple languages, requiring complex fine-tuning and leading to unstable duration predictions when dealing with multilingual text.

Method used

A speech synthesis model that removes speaker and speech information during training and adds it during inference, using a neural network architecture to generate audio signals tailored to user preferences for speaker and language, avoiding complex fine-tuning.

Benefits of technology

Enables accurate pronunciation and natural-sounding speech synthesis for multiple speakers and languages, improving quality and reliability in vehicle environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A speech synthesis device includes a memory configured to store speech information configured by a user and speaker samples corresponding to speaker information selected by the user. The speech synthesis device also includes a processor configured to generate an audio signal corresponding to the input text by applying a speech synthesis model to the input text, speech information, and audio samples in response to a speech synthesis request from the user. The speech synthesis model is trained to generate an audio signal incorporating features of the training text and features of a training audio signal. Speech information from the training text and speaker information from the training audio signal are then removed from the generated audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL AREA

[0001] The present disclosure relates to a method and a device for speech synthesis for multiple languages ​​and multiple speakers. BACKGROUND

[0002] The information in this section merely provides background information relating to the present disclosure and does not necessarily represent the state of the art.

[0003] Speech synthesis is a technology that produces sound similar to human speech and is commonly known as a TTS (Text-to-Speech) system. Speech synthesis technology delivers information to the user through speech signals rather than text or images, which is particularly useful when the user cannot see the screen of an operating machine, for example, when driving a car or if the user is blind.

[0004] Conventional speech synthesis methods involve generating a spectrogram based on input text and then generating a sound wave based on that spectrogram. A spectrogram is a tool for visualizing and understanding a sound image or waveform obtained by converting an audio signal in the time domain into frequency components along the time axis. Based on the spectrogram, the properties of a waveform and its spectrum can be visualized. Furthermore, the speech synthesis method can generate sound waves that reflect the characteristics of the speaker's speech. The speech synthesis method can generate a speech signal corresponding to the input text based on attributes such as the speaker's voice, prosody, pitch, and speaking rate.

[0005] Recently, a speech synthesis method that synthesizes speech from text based on an artificial neural network has been gaining increasing attention. One popular speech synthesis method based on an artificial neural network is a flow-based method. The flow-based method estimates the probability of text by applying an invertible transformation.

[0006] For a conventional speech synthesis model, it is difficult to synthesize the speech of an untrained (or unseen) speaker. Specifically, the training data used to train a conventional speech synthesis model consists of [text, speaker, speech] pairs. Since most speakers only speak one language, it is difficult for a speech synthesis model to generate natural-sounding speech from a speaker in another language. For example, a speech synthesis model trained on speech data from a man who speaks English will have limited ability to synthesize speech data representing a man who speaks Korean. In other words, conventional speech synthesis methods have many limitations when it comes to synthesizing natural-sounding speech that reflects the speaker's speech style, emotional expression, and so on.

[0007] Furthermore, language embeddings can be added to the text encoder included in the flow-based speech synthesis model, in addition to the input text, to address the challenges of multilingual synthesis. Based on the above procedure, the speech synthesis model can learn multiple languages; however, complex fine-tuning is required in subsequent stages of the model to produce high-quality speech. Additionally, speaker and language embeddings can be added to the duration predictor, which forecasts the duration of the input text in the flow-based speech synthesis model. Using the above procedure, the speech synthesis model can learn the duration features of multiple languages. However, if a sentence is expressed in multiple languages, the predicted duration may become unstable. SUMMARY

[0008] Embodiments of the present disclosure provide a speech synthesis method and a speech synthesis device for generating speech with multiple speakers / in multiple languages ​​with more accurate pronunciation in a vehicle environment. The speech synthesis method and the speech synthesis device train a speech synthesis model to remove speaker and speech information from the text during the training phase and to add speaker and speech information to the text during the inference phase.

[0009] At least one embodiment of the present disclosure provides a speech synthesis device. The speech synthesis device includes a memory configured to store speech information configured by a user and audio samples of a speaker corresponding to the speaker information selected by the user. The speech synthesis device also includes a processor configured to generate an audio signal corresponding to an input text by applying a speech synthesis model to the input text, the speech information, and the audio samples in response to a speech synthesis request from the user. The speech synthesis model is trained to generate an audio signal that includes features of the training text and features of a training audio signal. Speech information from the training text and speaker information from the training audio signal are removed from the generated audio signal.

[0010] A further embodiment of the present disclosure provides a speech synthesis method performed by a speech synthesis device. The speech synthesis method comprises receiving a speech synthesis request for input text, wherein the speech synthesis request includes speech information and speaker information provided by a user. The speech synthesis method also comprises generating an audio signal corresponding to the input text by applying a speech synthesis model to the input text, the speech information, and the speaker's audio samples corresponding to the speaker information. The speech synthesis model is trained to generate an audio signal that includes features of a training text and features of a training audio signal. Speech information from the training text and speaker information from the training audio signal are removed from the generated audio signal.

[0011] As described above, embodiments of the present disclosure provide a speech synthesis method and a speech synthesis device for generating speech with multiple speakers / in multiple languages ​​with more accurate pronunciation in a vehicle environment. The speech synthesis method and the speech synthesis device train a speech synthesis model to remove speaker and speech information from a text during the training phase and to add speaker and speech information to the text during the inference phase. Thus, complex fine-tuning after the training phase can be avoided, and the quality of the synthesized speech can be improved.

[0012] According to embodiments of the present disclosure, even when a sentence contains multiple languages, a vehicle occupant can receive synthetic voice guidance tailored to his preferred speaker and language characteristics. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 represents the construction of a vehicle according to an embodiment of the present disclosure. Fig. 2 represents a speech synthesis system according to an embodiment of the present disclosure. Fig. 3 represents a training of a speech synthesis model according to an embodiment of the present disclosure. Fig. 4 represents a loss of time according to one embodiment of the present disclosure. Fig. Figure 5 describes the functioning of a speech synthesis model according to an embodiment of the present disclosure. Fig. Figure 6 shows a flowchart representing a speech synthesis method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0013] Some embodiments of the present disclosure are described in detail below with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals denote the same elements, even if the elements are shown in different drawings. Furthermore, for the sake of clarity and brevity, detailed descriptions of related, known components and functions are omitted in the following description of some embodiments, insofar as these could make the subject matter of the present disclosure unclear.

[0014] Furthermore, various expressions such as "first," "second," "A," "B," "a," "b," etc., are used only to distinguish one component from another, not to imply or suggest the materials, order, or sequence of the components. When, in this patent specification, a part "comprises" or "includes" a component, this means that the part further comprises, and does not exclude, other components unless expressly stated otherwise. Terms such as "unit," "module," and the like refer to one or more units for processing at least one function or process and may be implemented by hardware, software, or a combination thereof.

[0015] Each constituent element of a device or method according to embodiments of the present disclosure can be implemented by hardware, software, or a combination of hardware and software. Likewise, the function of each constituent element can be implemented by software, and a microprocessor can execute the function of the software corresponding to each constituent element.

[0016] If a component, device, element or the like is described in the present disclosure as being designed for a specific purpose or for performing an operation, a function or the like, the component, device or element should here be considered as being "designed to" fulfill that purpose or to perform that operation or function.

[0017] The detailed descriptions provided below, together with the accompanying drawings, serve only to explain exemplary embodiments of the present disclosure, which should not be regarded as the only embodiments of the present disclosure.

[0018] The present disclosure relates to speech synthesis for multiple languages ​​and multiple speakers. Embodiments of the present disclosure provide a speech synthesis method and a speech synthesis device for generating speech with multiple speakers / in multiple languages ​​with more accurate pronunciation in a vehicle environment. The speech synthesis method and the speech synthesis device train a speech synthesis model to exclude speaker and speech information from the text during the training phase and to add speaker and speech information to the text during the inference phase.

[0019] Fig. Figure 1 represents the construction of a vehicle according to an embodiment of the present disclosure.

[0020] With reference to Fig. 1 indicates a vehicle 10 comprising a microphone 110 through which a user's voice is input, an input module 120 that receives vehicle information, a loudspeaker 130 that emits a sound required to provide a service requested by the user, a display 140 that shows an image that may be required to provide a service requested by the user, a communication module 150 that performs communication with an external device, and a control unit or controller 160 that controls the aforementioned components and other components of the vehicle.

[0021] The microphone 110 can be located at a position within the vehicle 10 where the user's voice is input. The user inputting their voice into the microphone 110 provided in the vehicle 10 can be the driver. The microphone 110 can be located in a position such as the steering wheel, the center panel or center console, the headliner, or the rearview mirror to receive the driver's voice.

[0022] In addition to the user's voice, various tones or acoustic signals generated in the vicinity of microphone 110 can be input into microphone 110. Microphone 110 outputs an audio signal corresponding to the input audio signal. The output audio signal can be processed by the controller 160 or transmitted to an external server device via the communication module 150.

[0023] In addition to the microphone 110, the vehicle 10 may include an input module 120 that receives user commands. The input module 120 may be in the form of a button or a jog shuttle located in the area of ​​the instrument cluster or speedometer, the AVN (audio, video, navigation) area of ​​the center console, the gearshift area, or the steering wheel.

[0024] To receive control commands relating to the passenger seat, the input module 120 can include an interface device provided at the door of each seat, as well as an interface device provided at the armrest of the front seat or the armrest of the rear seat.

[0025] The input module 120 can include a touchpad that is integrated with the display 140 to create a touchscreen.

[0026] The input module 120 can include a camera. The camera can capture at least one internal or external image of the vehicle 10. The camera can be mounted inside, outside, or both inside and outside the vehicle 10. The images acquired by the camera are processed by the controller 160 or an external device; based on the acquired images, the gaze or direction of gaze, mouth shape, face, behavior, or condition of the occupant in the video can be analyzed.

[0027] The loudspeaker 130 outputs an electrical signal in the form of a sound wave. The loudspeaker 130 can be positioned so that it is directed towards the interior of the vehicle 10 near each door, the roof, the windshield, or the rear window. The loudspeaker 130 can refer to various types of loudspeakers, such as array loudspeakers.

[0028] Display 140 may include an AVN display, a speedometer display, or a head-up display (HUD) provided on the center instrument panel of the vehicle 10. Alternatively, display 140 may include a rear-seat display provided on the back of the front seat headrest for rear-seat passengers. If the vehicle 10 is a multi-person vehicle, display 140 may also include a display mounted on the headliner.

[0029] The display 140 must be provided in positions where it can be seen by the occupants of vehicle 10, and there are no further restrictions regarding the number or position of the displays 140.

[0030] The Communication Module 150 can exchange signals with other devices using at least one of several wireless communication methods, such as Bluetooth, 4G, 5G, or Wi-Fi. Alternatively or additionally, the Communication Module 150 can exchange information with other devices via a cable connected to a USB (Universal Serial Bus) port, an AUX port, etc.

[0031] Since the communication module 150 is equipped with two or more communication interfaces that support different communication methods, it can also exchange information signals with two or more other devices.

[0032] For example, the communication module 150 can communicate with a mobile device located inside the vehicle 10 via Bluetooth to receive information (video, voice, contact information, schedule, etc.) received from or stored on the mobile device; transmit the user's voice by communicating with the server 1 via 4G or 5G communication; and receive signals necessary to provide a service requested by the user. Furthermore, the communication module 150 can exchange necessary signals with the server 1 via a mobile device connected to the vehicle 10.

[0033] In addition, the vehicle may include 10 a navigation device for providing route guidance, an air conditioning device for controlling or regulating the interior temperature, a window control device for controlling the opening / closing of windows, a seat heating device for warming the seats, a seat positioning device for adjusting the position, height or angle of the seats and a lighting device for adjusting the interior lighting.

[0034] The devices described above provide comfort features for the vehicle 10, and some of the devices may be omitted depending on the vehicle model and options. Furthermore, it should be noted that other devices may be included in addition to those described above. Known arrangements are used for driving the vehicle 10, and their description has been omitted from this disclosure.

[0035] The controller 160 can switch the microphone 110 on and off. The controller 160 can process or store the voice input into the microphone 110, or transmit the voice input to another device via the communication module 150.

[0036] The controller 160 can control images to be displayed on the display 140 and can control acoustic signals to be output to the speaker 130.

[0037] The control unit 160 can perform various control operations with respect to the vehicle 10. For example, the control unit 160 can, according to a command input from a user via the microphone 110 or the input module 120, control at least one of the navigation device, the climate control device, the window control device, the seat heating device, the seat positioning device, or the lighting device.

[0038] The controller 160 can include at least one memory that stores a program for performing the above-mentioned process as well as the process described in more detail below. The controller 160 can also include at least one processor that executes the stored program.

[0039] In the following description, intralingual synthesis refers to the synthesis of speech from a text in a language spoken by a speaker represented in a speaker embedding. For example, intralingual synthesis corresponds to the synthesis of speech from a Korean text by a Korean speaker.

[0040] Cross-linguistic synthesis refers to the synthesis of speech from a text in a language that a speaker represented in the speaker embedding does not speak. For example, cross-linguistic synthesis corresponds to the synthesis of speech from a Korean text by an English speaker.

[0041] Code-mix synthesis refers to the synthesis of speech from text in multiple languages. For example, code-mix synthesis corresponds to the synthesis of speech from "Korean+English" text by a Korean speaker.

[0042] In one example, intralingual synthesis can be used primarily during the training phase of a speech synthesis model. In addition to intralingual and cross-linguistic synthesis, code-mixing synthesis can be performed during the inference phase of the speech synthesis model.

[0043] According to one embodiment of the present disclosure, the controller 160 can operate as a speech synthesis device. For example, a user can request audio output as if the text displayed on the display 140 were being spoken by a preferred speaker in a preferred language. The user's preferred language and speaker can be preconfigured. The controller 160 can synthesize speech corresponding to the text by converting the text into an audio signal according to a requested speaker in a requested language. In other words, the controller 160 can perform intralingual or cross-linguistic synthesis. The user can hear natural-sounding speech as if a selected speaker had spoken the selected text naturally.In another example, if the user requests speech recognition by saying, "What does 'Encantado de conocerlo' mean in English?", the Controller 160 can perform code-mixing synthesis to generate "Encantado de conocerlo means 'Nice to meet you' in English." The Controller 160 can receive pre-stored audio samples and apply a speech synthesis model to the multilingual text and the audio samples. To process multilingual text, the Controller 160 can configure the languages ​​included in the multilingual text. The speech synthesis model can generate an audio signal as if a multilingual text were being spoken naturally in several languages. The audio signal can be output via the Speaker 130. The user can hear natural-sounding speech, as if a selected speaker had spoken the multilingual text naturally.

[0044] In another example, the 160 controller can synthesize and convert the user's voice or text according to a different speaker and language.

[0045] According to another embodiment, the control unit 160 and the communication module 150, in conjunction with an electronic device located outside the vehicle 10, can provide a speech synthesis function.

[0046] Fig. 2 represents a speech synthesis according to an embodiment of the present disclosure.

[0047] With reference to Fig. 2 comprises a speech synthesis system according to one embodiment, a vehicle 210, and an electronic device 220. A speech synthesis method according to one embodiment can be implemented by the vehicle 210 and the electronic device 220. The speech synthesis model can be implemented in the electronic device 220. The speech synthesis method can be performed by the electronic device 220.

[0048] The electronic device 220 can perform speech synthesis. The electronic device 220 can be implemented by at least one of the server device 221 or the mobile device 223.

[0049] The vehicle 210 can transmit a speech synthesis request to the electronic device 220. The electronic device 220 can respond to the vehicle 210 with an audio signal that is the result of the speech synthesis. The speech synthesis request includes text to be synthesized into speech, speech information for the text, and speaker information.

[0050] For example, the vehicle 210 can transmit a speech synthesis request to the electronic device 220, consisting of a sentence of [text, speaker] or [text, speaker, speech] pairs. The electronic device 220 can generate an audio signal indicating that the requested text is being spoken by a selected speaker. Since an actual speaker is not uttering the text, the audio signal is generated data rather than recorded data. However, the audio signal can contain natural pronunciation and voice, as if it were a recording of an actual speaker speaking fluently in the requested language. The electronic device 220 can transmit the generated audio signal to the vehicle 210 as a result of the speech synthesis.The vehicle 210 can reproduce the received audio signal, thereby outputting the audio signal according to the text, speaker, and language requested by the user. As described above, in the case of code-mix synthesis (i.e., when text contains words or characters from multiple languages), language information can be generated or configured for each word or character.

[0051] In one embodiment, the electronic device 220 may include a processor and a memory for speech synthesis.

[0052] Fig. 3 represents a training of a speech synthesis model according to an embodiment of the present disclosure.

[0053] With reference to Fig. Figure 3 shows a model architecture 30 for a training phase of a speech synthesis model. In the training phase, the model architecture 30 can include a speech embedding module 300, a character embedding module 310, an encoder 320, a duration predictor 330, a speaker encoder 340, a projection module 350, a matching data estimator 360, a decoder 370, a posterior encoder 380, and an audio generator 390. In other embodiments, a section of the constituent elements included in the model architecture 30 can be omitted and / or the order of the constituent elements can be changed. The model architecture 30 can further include a discriminator (not shown). In addition, the model architecture 30 can include a trainer (not shown) for training the speech synthesis model or can be implemented by connecting it to an external trainer.In one embodiment, the posterior encoder 380, the discriminator and the trainer are used only for training the speech synthesis model.

[0054] In embodiments of the present disclosure, conditioning an embedding can mean adding, multiplying, or subtracting an embedding to or from an input or internal element. To ensure dimensionality matching, the dimensionality of the embedding can be adjusted by a layer of the neural network (e.g., a fully connected layer or a convolutional layer). For example, in the duration predictor 330, a speech embedding can be added to a text feature vector.

[0055] To train the speech synthesis model, a training dataset can be prepared in advance. The training dataset can include text data, corresponding audio data, and speech data. The audio data can be a recording of text data actually spoken by multiple speakers. The training data can consist of pairs of [training text, speaker's training audio signal, speech]. However, since most speakers only speak a small number of languages, training data consisting of [text, speaker's audio signal, speech] can be sparse. In other words, datasets for multiple speakers and multiple languages ​​can be sparse.

[0056] The training text can consist of a sequence of characters in a natural language. For example, the string can include alphabetic characters, numbers, punctuation marks, or other special characters. The audio data corresponding to the text data is a recording of text data actually spoken by multiple speakers.

[0057] The training audio signal represents the speech data of speakers. The speaker refers to the person who spoke the audio data corresponding to the text data. The training audio signal can include vocal and / or linguistic characteristics of a speaker. Linguistic characteristics of a speaker can include at least one of several elements, such as speech rate, pause intervals, pitch, tone of voice, prosody, intonation, pronunciation, or emotion. Audio signals from multiple speakers can be prepared. Because the training audio signal represents the linguistic characteristics of a specific speaker, it may differ from the audio data corresponding to the training text.

[0058] Furthermore, linear or Mel spectrograms can be prepared from audio data for use as ground truth during the training phase. Linear spectrograms can be generated by applying the short-time Fourier transform (STFT), discrete Fourier transform (DFT), or fast Fourier transform (FFT) to the audio data. A Mel spectrogram can be obtained by fitting the frequency interval of the linear spectrogram to the Mel scale. Alternatively, a Mel spectrogram can be obtained by applying a Mel filter bank to the linear spectrogram. Linear or Mel spectrograms can be used to calculate the reconfiguration loss described later.

[0059] Language information refers to the language of the text data. This language information can be represented as a number. For example, the language information could include Korean, English, German, Japanese, and Chinese. Korean could be designated as 1, English as 2, and German as 3.

[0060] Fig. 3 represents a process in which training text, a training audio signal, a training spectrogram and speech information are processed within the training data set for training.

[0061] The Speaker Encoder 340 can convert or map the speaker's training audio signal into a speaker embed. A speaker embed represents the speaker's speech characteristics and can be expressed in vector form. The speaker embed can also include speaker identification information.

[0062] The Speaker Encoder 340 can represent discontinuous data values ​​contained in the speaker information as a vector consisting of consecutive numbers. For example, the Speaker Encoder 340 can generate a speaker embedding vector based on a combination of at least one or two or more different artificial neural network models, including a Pre-Net, a CBHG module, a Deep Neural Network (DNN), a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), a Long Short Memory Network (LSTM), and a Bidirectional Recurrent Deep Neural Network (BRDNN).

[0063] As in Fig. As shown in Figure 4, the speaker encoder 340 can perform speaker embedding (e s ) using the voice of a speaker who has spoken the training text of the encoder 320. In addition, the speaker encoder 340 can generate a speaker embedding (e s_hat) using the voice of a different speaker who speaks a different language than the one used for the training text.

[0064] The Speech Embedding Module 300 can convert the speech information corresponding to the training text into a speech embed. For example, the Speech Embedding Module 300 can map the speech information into a speech embed using one-hot coding. The speech embed can be in vector form. Since one-hot coding is a widely known technology in the field of speech synthesis, a detailed description is omitted.

[0065] The Character Embedding Module 310 can convert or map training text into a character embedding. For example, the training text might be composed of sentence or character units. The Character Embedding Module 310 can split the training text into character units and convert each split text into a character embedding. Alternatively, the Character Embedding Module 310 can split the training text into alphabetic or phonemic units and then convert these into character embeddings. For instance, the Character Embedding Module 310 can perform character embedding using a model of an artificial neural network. Character embeddings can be represented as learnable vectors.

[0066] The Encoder 320 can extract text feature vectors from character embeddings. Text feature vectors extracted by the Encoder 320 can include features of character embeddings, i.e., the training text.

[0067] In one embodiment, the encoder 320 can perform encoding in phoneme units. For this purpose, the encoder 320 can separate the character embeddings into phoneme units of the training text. In another embodiment, the encoder 320 can perform encoding for the entire set of character embeddings.

[0068] The Encoder 320 can incorporate an artificial neural network. For example, the Encoder 320 can be a transformation-based Encoder 320. The transformation-based Encoder 320 can comprise multiple transformer blocks, and each transformer block can include at least one Encoder 320, at least one decoder, and an attention module. For example, the transformation-based Encoder 320 can comprise 10 transformer blocks. The transformer block can extract context vectors from character embeddings using the Encoder 320, recognize important character embeddings using the attention module, and generate text feature vectors from a context vector and the outputs of the attention module using the Decoder 370.

[0069] Projection module 350 can output the distribution of text feature vectors for dimension matching before element-wise summation. The text feature vector distribution can be a priority distribution that includes the means and standard deviations of the text feature vectors. The distribution can include the mean and standard deviation of each text feature vector corresponding to each phoneme. Projection module 350 can be a linear projection layer.

[0070] The Posterior Encoder 380 can encode training spectrograms and output latent variables. Encoding can involve extracting features from existing data and converting or transforming those features into data with reduced size or dimensionality compared to the original data. In other words, the output from encoding can be the result obtained by compressing the input data. The latent variable can be a latent vector. Latent variables include the speaker's voice and / or features of the language. Linguistic features can also be included.

[0071] The training spectrograms can be linearly scaled spectrograms or Mel spectrograms that have been converted or transformed from the corresponding audio data of the text data. In another embodiment, an audio file format such as wav or mp3 is input into the posterior encoder 380, and the posterior encoder 380 can encode the audio signal to extract a latent vector.

[0072] The Posterior Encoder 380 can incorporate a deep neural network. For example, the Posterior Encoder 380 can be a VAE (Variational Auto-Encoder) encoder. The Posterior Encoder 380 can include non-causal WaveNet residual blocks, which are used in the WaveGlow and Glow-TTS models. For example, the Posterior Encoder 380 can contain 12 WaveNet residual blocks. The non-causal WaveNet residual block can include an extended convolution layer with controlled activation units and skip connections. A linear projection layer above the block can generate the mean and variance of the posterior normal distribution.

[0073] The Decoder 370 can output converted latent variables based on the latent variables, speech embeddings, and speaker embeddings. The Decoder 370 can generate a latent variable with a different distribution than the previous distribution of the latent variable. This different distribution can be a normal distribution.

[0074] Decoder 370 can remove speaker and speech information from latent variables. For example, Decoder 370 can receive speech embeddings corresponding to the training text and speaker embeddings corresponding to the training audio signal. Decoder 370 can remove speaker- and speech-related features within the latent variables by normalizing the speaker- and speech-related information of the latent variables. As an example, speaker and speech information removal can be performed based on the feature ratio normalization (FRN) of Equation 1. SN(x,es)=x−m(es)exp(v(es))LN(x,el)=x−m(el)exp(v(el))FRN=ρ⋅SN+(1−ρ)⋅LN

[0075] In equation 1, SN(x,g) represents the result of speaker normalization (SN). x is a normalization target and can be a latent variable. s stands for speaker embedding, m(e s) represents the mean of the speaker embedding and v(e s ) represents the variance of speaker embedding. The mean and variance of speaker embedding can be calculated for the entire dataset used for training. According to speaker normalization, the latent variable includes the speaker characteristics e s Speaker normalization can be applied to subdivisions of the latent variable.

[0076] LN(x,e l ) represents the result of language normalization (LN). x is a normalization target and can be a latent variable. e l stands for language embedding, m(e l ) represents the mean value of the language embedding and v(e) l) represents the variance of the language embedding. The mean and variance of the language embedding can be calculated for the entire dataset used for training. According to language normalization, the latent variable includes the language features e l Language normalization can be applied to subdivisions of the latent variable.

[0077] FRN is defined as a linearly weighted sum of SN and LN. The feature ratio ρ can be calculated based on the mean and variance of the speaker embedding and the mean and variance of the speech embedding. For example, the feature ratio ρ can be estimated based on the output of a neural network layer that uses the mean and variance of the speaker embedding and the mean and variance of the speech embedding as inputs. This neural network layer can be trained. According to Equation 1, speaker and speech embedding features can be excluded from the latent variable.

[0078] In this way, the Decoder 370 can normalize speaker- and speech-related features in the latent variable based on the speaker and speech embeddings and can generate a transformed latent variable by sampling the latent variable from a simpler or more complex distribution than the distribution of the preprocessed latent variable. Here, preprocessing refers to the normalization of the speaker and speech embeddings. The Decoder 370 can remove the speech information from the training text and the speaker information from the training audio signal from within the latent variable by normalizing the latent variable based on the speech and speech embeddings. The transformed latent variable includes the features of the training audio signal but may not include the speech information from the training text and the speaker information from the training audio signal.

[0079] Decoder 370 can have a normalization flow function. Decoder 370 can obtain the transformed latent variable by applying the function f to the preprocessed latent variable. Since the distribution transformation of Decoder 370 is reversible, an inverse function for Decoder f can be defined. The transformed latent variable can have the same, a different, or a more complex distribution compared to the original latent variable. Here, a complex distribution, as opposed to a simple normal distribution, means a distribution with multiple local minima and maxima.

[0080] The Decoder 370 can feature a deep neural network. For example, the Decoder 370 can be a flow-based decoder. The Decoder 370 can include a variety of affine coupling layers. For example, the Decoder 370 can include four affine coupling layers. One section of the variety of affine coupling layers can be used for speaker and speech embedding.

[0081] The affine coupling layer for excluding speaker and speech information can be called a normalized affine speaker-language coupling layer (SLNAC). The normalized affine speaker-language coupling layer can obtain a speaker-language normalization result from the latent variable according to Equation 1. In one embodiment, the affine coupling layer can generate a latent output variable by applying speaker-language normalization to a portion of the dimensions of the latent input variable, applying the affine transformation to the normalization result based on scale and bias parameters, and combining the transformation result with a speaker-language normalization result for the remaining dimensions of the latent input variable.The affine coupling layer is easily invertible and has a triangular Jacobian matrix; the determinant can be calculated on the basis of the Jacobian expression, from which the model density q can be easily calculated.

[0082] As described above, the Decoder 380 can generate a converted latent variable that is normalized using the speaker and speech embeddings.

[0083] The Matching Data Estimator 360 can output matching data based on the distribution of text feature vectors and the converted latent variables.

[0084] In one example, the Match Data Estimator 360 can estimate a matrix for sorting the duration of each phoneme in the training text based on the means, standard deviations, and transformed latent variables of the text feature vectors as match data. The dimensionality of the match data can depend on the length of the latent variable and the length of the character embedding. For example, rows can represent phonemes, while columns can represent time intervals. In the match data, the duration of each phoneme is expressed as a path; elements along the path have the value 1, while other elements have the value 0. In other words, match data refers to matching information between phonemes in the training text and their respective latent variables.

[0085] To estimate matrix A, which represents the matching data between the phonemes included in the training text, Monotonic Alignment Search (MAS), a matching search technique that maximizes the probability of data parameterized by a flow normalization function, can be used. The Matching Data Estimator 360 can estimate matching data by applying the MAS technique to the distribution of text feature vectors and transformed latent variables. Since the MAS technique is a well-known method, a detailed description is omitted.

[0086] The matching data can be used to train the duration predictor 330. The matching data can relate to the similarity between text feature vectors and the transformed latent variable.

[0087] The duration predictor 330 can receive text feature vectors, speaker embeddings, and speech embeddings and, based on the received input, predict the duration of each phoneme in the training text. In other words, the duration predictor 330 can predict phoneme duration data.

[0088] The duration predictor 330 can use speaker embeddings as conditioning information. Duration predictor 330 can condition speaker embeddings during the computation process. For example, speaker embeddings can be added to or multiplied by text feature vectors.

[0089] The duration predictor 330 can use language embeddings as conditioning information. Duration predictor 330 can condition language embeddings during the computation process. For example, language embeddings can be added to or multiplied by text feature vectors.

[0090] As in Fig. As shown in Figure 4, in one embodiment the duration predictor 330 uses a speaker embedding (e s ) based on the voice of a speaker who has spoken the training text, or a speaker embedding (e s_hat ) based on the voice of a different speaker speaking in a language other than the one used for the training text. Using speaker embedding (e s_hat Based on the voice of another speaker speaking a language different from the language used for the input sentences, the duration predictor 330 can use the language information inherent in the text and the language embedding, but ignore the language information inherent in the speaker embedding. Therefore, in the inference phase of the speech synthesis model, even if a sentence includes multiple languages, the duration predictor 330 can reliably generate the duration of a phoneme.

[0091] The Audio Generator 390 can generate an audio signal in the time domain based on latent variables. In other words, the Audio Generator 390 can generate a speech waveform based on the previous distribution of latent variables.

[0092] The Audio Generator 390 can incorporate a deep neural network. The Audio Generator 390 can be a vocoder. For example, the Audio Generator 390 can be a HiFi GAN generator. The Audio Generator 390 can include a stack of transposed convolutions, with each convolution potentially followed by an MRF (Multi-Receptive Field Fusion) module. The output of MRF is a sum of the outputs from "residual blocks" with varying receptive field sizes. The Audio Generator 390 can include a linear layer responsible for converting speaker embeddings, adding the speaker embeddings to the latent variable z, and generating an audio signal from the combination of the latent variable and the speaker embeddings.

[0093] End-to-end training can be applied to architecture 30 of the speech synthesis model described above. Architecture 30 of the speech synthesis model can be trained using a computer-implemented training device. A discriminator can be used to train audio generator 390. This discriminator can be the HiFi discriminator.

[0094] In one embodiment, at least one of the loss functions of the speech synthesis model can be reconstruction loss, Kullback-Leibler divergence loss, time loss, adversarial loss, and feature matching loss.

[0095] The reconstruction loss can be calculated based on the difference between the spectrogram of the generated audio signal and the training spectrogram. As described above, a transformer can also be used to convert the generated audio signal into a spectrogram. Furthermore, the training spectrogram can be generated from the audio data corresponding to the training text.

[0096] The KL divergence loss can be calculated based on the difference between the latent variable and the text feature vectors. The KL divergence loss can be calculated based on the difference between the posterior probability of the latent variable and the conditional prior probability of the text feature vector. In other words, the KL divergence loss can relate to the similarity between the distribution of the latent variable and the distribution of the text feature vector.

[0097] The time loss can be calculated based on the difference between the phoneme duration data predicted by the time duration predictor 430 and the phoneme duration generated by the matching data estimator 360. As described above, the matching data generated by the matching data estimator 360 can be used to determine the duration d MAS of phonemes. A duration loss can be calculated based on the mean square error (MSE). The duration loss is intended to enable the duration predictor 330 to predict the duration of each phoneme depending on the speaker and the language. As in Fig. As shown in Figure 4, the duration predictor 330 can predict the duration d. intra using speaker embedding e s generate the voice of the speaker who spoke the training text. Alternatively, the duration predictor 330 can predict the duration d. cross using speaker embedding es_hat generate a voice based on that of a different speaker in a language other than the one used for the input sentence's training text. As in Fig. As shown in 4, the time loss is defined as the minimum length Ld intra or Ld cross depending on the speaker embedding used (e s or s_hat Alternatively, the time loss can be either Ld intra as well as Ld cross include. Using speaker embedding e s based on the voice of a speaker who spoke the input sentence and the speaker embedding e s_hatBased on the voice of a different speaker in a language other than that used for the input sentence, the trainer can use the speech information inherent in the text and the speech embedding, but exclude the speech information inherent in the speaker embedding. As a result, the duration predictor 330 can reliably generate the duration of a phoneme in the inference phase of the speech synthesis model, even if a sentence includes multiple languages.

[0098] The adversarial loss can be calculated based on the discriminator's determination of whether an audio signal generated by the 390 audio generator is genuine. To reduce the adversarial loss, the discriminator must correctly identify the generated audio signal as genuine data. The adversarial loss causes the discriminator to output the value 1 in response to genuine data input and the value 0 in response to falsified data input. The feature matching loss can be calculated based on the difference between the features extracted from the generated audio signal by the discriminator and the features extracted from the actual audio signal.

[0099] By training based on adversarial loss and feature matching loss, the audio generator can produce 390 audio signals that are nearly identical to the actual data.

[0100] The loss function of the model architecture 30 can further include speaker consistency loss (SCL). SCL is calculated based on the difference between the output of the speaker encoder 340 and the ground truth. In other embodiments, the speaker encoder 340 can be pretrained.

[0101] In another embodiment, the reconstruction loss alone can be used as a loss function. Through end-to-end learning, the model architecture 30 can be updated based on the difference between audio signals generated by the model architecture 30 from training text and marked audio samples that correspond to the training text.

[0102] The trainer can update the model architecture 30 in the direction in which the above loss function decreases. Through iterative training based on the overall loss function, each component of the model architecture 30 is refined, enabling the speech synthesis model to generate natural speech signals from the speaker.

[0103] The training process described above makes the speech synthesis model robust against speakers and language diversity. In particular, it reduces its dependence on specific speakers and languages. In other words, the speech synthesis model is trained on text rather than on specific speakers and languages. During the inference phase of the speech synthesis model, it then uses speaker and language information. When learning phoneme durations, the speech synthesis model can exclude the language information inherent in speaker embeddings.

[0104] Even if the speech synthesis model receives a code-mixed text, it can generate a natural-sounding speaker's voice from the text using speaker and speech information. Even if the training dataset contains a substantial amount of [Korean text, Korean voice] data and only a limited amount of [English text, Korean voice] data, the speech synthesis model will still learn, for example, the context of the Korean / English text and the speech features without relying on speaker and speech information. During the inference phase, the speech synthesis model can then synthesize a natural-sounding speech by adding speaker and speech embeddings to the [code-mixed text].

[0105] Fig. Figure 5 describes the functioning of a speech synthesis model according to an embodiment of the present disclosure.

[0106] With reference to Fig. Figure 5 shows the arrangements of the speech synthesis model 40. The speech synthesis model 40 can generate an audio signal as if the input text were being spoken in a specific language by a specific speaker. In particular, the speech synthesis device stores speech information set by the user and pre-recorded audio samples of a selected speaker. The content of the audio samples can differ from that of the input text. The speech synthesis device can synthesize an audio signal by applying the speech synthesis model to speech information, the speaker's audio signals, and the target text.

[0107] In the inference phase, the speech synthesis model 40 can include a speech embedding module 410, a character embedding module 420, an encoder 430, a duration predictor 440, a speaker encoder 450 and a projection module 460, a matching unit 470, an inverted decoder 480 and an audio generator 490.

[0108] The speech synthesis model 40 can be achieved through the procedure of Fig. 3 are trained. The speech embedding module 410, the character embedding module 420, the encoder 430, the duration predictor 440, the speaker encoder 450, the projection module 460, the inverted decoder 480 and the audio generator 490 of Fig. 5 correspond to the speech embedding module 300, the character embedding module 310, the encoder 320, the duration predictor 330, the speaker encoder 340, the projection module 350, the decoder 370 and the audio generator 390 of Fig. 3. The inverted decoder 480 represents the inverse function of the decoder.

[0109] The speech embedding module 410 can convert the speech information of an input text into speech embeddings. In one embodiment, the speech embedding module 410 can be omitted, and speech embeddings corresponding to different languages ​​can be pre-stored. In other words, a speech embedding corresponding to the speech information of the input text can be pre-stored, and the inverted decoder 480 can receive the speech embedding. For example, in the case of cross-language synthesis, the language of the input text may differ from the language of the speaker's audio signal. The speech embedding module 410 can generate a speech embedding for each word. For example, in the case of code-mix synthesis, an input text may contain words or characters corresponding to multiple languages. The speech embedding module 410 can generate a speech embedding for each word or character.

[0110] The character embedding module 420 can convert a given input text into character embeddings. The input text can be assigned to a variable space for character embeddings.

[0111] The Encoder 430 can output text feature vectors for the input text by encoding character embeddings. Text feature vectors comprise features of each phoneme in the input text.

[0112] The Speaker Encoder 450 can receive an audio signal recording the voice of a selected speaker and output a speaker embedding by encoding the audio signal. The speaker embedding can include the speaker's voice and / or speech characteristics.

[0113] The duration predictor 440 can predict the duration of each phoneme in the input text based on text feature vectors, a speaker embedding, and a speech embedding, and output phoneme duration data including the duration of the phonemes. The phoneme duration data can include a predicted duration for each phoneme based on speech features and the speaker's voice and / or speech features. When generating the duration of each phoneme in the input text, the duration predictor can use speech information inherent in the input text and speech embedding, but excludes speech information inherent in the speaker embedding.

[0114] The phoneme duration data can be entered into the synchronization unit 470.

[0115] The projection module 460 can generate the distribution of text feature vectors. This distribution can include means and standard deviations. In this process, the text feature vector can be transformed to match the dimensionality of the matching data from the matching unit 470. The dimensionality of the data representing the distribution can correspond to one of the dimensions of the matching data.

[0116] The Matching Unit 470 can generate latent variables based on the distribution of text feature vectors and phoneme duration data. Latent variables can be generated from text feature vectors based on phoneme duration data. For example, the Matching Unit 470 can process the mean and standard deviation of text feature vectors corresponding to each phoneme with the matching data and output a latent variable as a result of the operation. The latent variable can include features of each phoneme in the input text and features related to the duration of each phoneme.

[0117] The inverted Decoder 480 can generate a transformed latent variable based on the latent variable, speech embedding, and speaker embedding. Since the inverted Decoder 480 is trained for speaker normalization and speech normalization to exclude speaker and speech information during the training phase, speaker embedding based on an audio signal from a speaker and speech embedding based on input text must be integrated into the latent variable during the inference phase.

[0118] A section of affine coupling layers within the inverted decoder 480 can be used to denormalize speaker and speech embeddings. The corresponding affine coupling layer can be referred to as a denormalized affine speaker-speech coupling layer. As an example, speaker and speech embeddings can be integrated into the latent variable based on the feature-ratio denormalization (FRDN) of Equation 2. FRDN=exp(v(es))⋅exp(v(el))ρ⋅exp(v(el))+(1−ρ)⋅exp(v(es))⋅x +ρ⋅exp(v(el))⋅m(es)+(1−ρ)⋅exp(v(es))⋅m(es)ρ⋅exp(v(el))+(1−ρ)⋅ex p(v(es))SDN(x,es)=x⋅exp(v(es))+m(es)LDN(x,el)=x⋅exp(v(el))+m(el)

[0119] In equation 2, x represents a denormalization target, which can be a latent variable. s stands for speaker embedding, m(e s ) represents the mean of the speaker embedding and v(e s ) represents the variance of speaker embedding. el stands for language embedding, m(e j ) represents the mean value of the language embedding and v(e) j ) represents the variance of the speech embedding. FRDN is the reciprocal of FRN from Equation 1. FRDN can be applied to a section of the dimensions of the latent variable. As described above, the feature ratio ρ can be calculated based on the mean and variance of the speaker embedding and the mean and variance of the speech embedding. As shown in Equation 2, if ρ equals 1, FRDN is given by SDN(x,e). s ) is replaced, i.e., the speaker's denormalization is calculated. SDN is the reciprocal of SN. If ρ equals 0, FRDN is replaced by LDN(x, e). l) is replaced, i.e., the denormalization of the language is calculated. LDN is the inverse of LN. Therefore, FRND can be viewed as a nonlinear weighted sum of SDN and LDN. According to Equation 2, features of the speaker and language embeddings can be integrated into the latent variable.

[0120] The inverted decoder 480 can convert the latent variable that has been preprocessed using speech embedding and speaker embedding. Here, the preprocessing involves denormalizing the speaker embedding and speech embedding. The inverted decoder 480 can obtain a converted latent variable by applying the inverse function f. -1The inverted decoder 480 applies a normalization flow function used in the training phase to the latent variable. The transformed latent variable may have a simpler or more complex distribution than the preprocessed latent variable. The inverted decoder 480 can transform the distribution of the latent variable based on speaker embedding and speech embedding. The transformed latent variable includes features of the input text, features of the speech information, features of the speaker's audio signals, and duration features.

[0121] The Audio Generator 490 can generate an audio signal, representing a sound wave, from the converted latent variable. The generated audio signal can be identical to or similar to the audio recording of the user-selected speaker uttering the input text. Even if the selected speaker is not familiar with the language of the input text, a result can be generated as if the selected speaker had uttered the input text in that language.

[0122] Fig. Figure 6 shows a flowchart representing a speech synthesis method according to an embodiment of the present disclosure.

[0123] With reference to Fig. 6. In operation S610, the speech synthesis device receives a speech synthesis request from the user.

[0124] Here, the speech synthesis request comprises input text to be synthesized into speech. In one embodiment, the speech synthesis request can include speaker and speech information requested by the user. In another embodiment, the speaker and speech information are preconfigured by the user, and the speech synthesis device can store the configured information in advance. The speech synthesis device can also store speaker audio samples in advance. However, there may be cases where audio samples in the requested language are not available for the requested speaker. In other words, the speech information of the requested text may differ from the language of the requested speaker's audio samples. For example, in the case of code-mix synthesis, the input text may include words or characters corresponding to multiple languages.The speech synthesis request can include speech information for each word or character.

[0125] In operation S620, the speech synthesis device applies a speech synthesis model to the input text, speech information, and audio samples to produce an output audio corresponding to the text.

[0126] Here, the speech synthesis model is pre-trained to generate an audio signal that includes features of the training text and features of the training audio signal, while removing speech information from the training text and speaker information from the training audio signal. During the training phase, the speech synthesis model can remove speech information from the training text and speaker information from the training audio signal by normalizing the latent training variables, including features of the training text and features of the training audio signal, based on the speech information from the training text and speaker information from the training audio signal. Furthermore, the speech synthesis model uses the training text, the speech embedding of the training text, and the speaker embedding of the training audio signal when generating the duration of each phoneme in the training text.At this point, the speech synthesis model is pre-trained to use the speech information inherent in the training text and the speech embedding of the training text, and to exclude the speech information inherent in the speaker embedding of the training audio signal.

[0127] In the inference phase, the speech synthesis model can include a speech embedding module, a character embedding module, an encoder, a speaker encoder, a duration predictor, a projection module, a matching unit, an inverted decoder, and an audio generator.

[0128] The speech embedding module can convert the requested speech information into a speech embed. In other embodiments, a speech embed can be pre-stored, and the speech embedding module may not be included in the speech synthesis model.

[0129] The character embedding module can convert input text into character embeddings.

[0130] The encoder can encode character embeddings into text feature vectors.

[0131] The speaker encoder can encode audio samples for outputting a speaker embedding.

[0132] The duration predictor can predict phoneme duration data, which includes the duration of each phoneme in the input text based on the text feature vectors, speech embedding, and speaker embedding. When generating the duration of each phoneme in the input text, the duration predictor can use speech information inherent in the input text and speech embedding, but excludes speech information inherent in the speaker embedding.

[0133] The projection module can generate a distribution of the text feature vectors. This distribution can include the mean and the standard deviation.

[0134] The matching unit can generate a latent variable based on the distribution of text feature vectors and phoneme duration data. As described above, in one embodiment, since the matching unit is trained for speaker normalization and speech normalization to exclude speaker and speech information during the training phase, the latent variable does not include the features based on the speaker and speech embeddings.

[0135] The inverted decoder can output a transformed latent variable based on the latent variable, the speech embedding, and the speaker embedding. For example, the inverted decoder can denormalize the latent variable based on the speaker and speech embeddings and output a transformed latent variable based on the denormalized latent variable.

[0136] The audio generator can produce an audio signal from the converted latent variable.

[0137] Although the steps or operations in the respective flowcharts are described as being carried out sequentially, they merely illustrate the technical idea of ​​some embodiments of the present disclosure. Therefore, a person skilled in the art in the field to which the present disclosure belongs could perform the steps or operations by modifying the sequences described in the respective drawings or by carrying out two or more of the steps in parallel. Thus, the steps or operations in the respective flowcharts are not limited to the chronological sequences shown.

[0138] It should be understood that the above description represents exemplary embodiments that can be implemented in various other ways. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or a combination thereof. It should also be understood that the functional components described in this disclosure are designated as "... unit" to emphasize the possibility of their independent implementation.

[0139] Various methods or functions described in some embodiments can be implemented as instructions stored on a non-volatile recording medium that can be read and executed by one or more processors. The non-volatile recording medium can, for example, include various types of recording devices in which data is stored in a form readable by a computer system. For example, the non-volatile recording medium can include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical drives, magnetic hard disks, and solid-state drives (SSDs), to name just a few.

[0140] Although embodiments of the present disclosure have been described for illustrative purposes, it is understood by a person skilled in the art in the field to which the present disclosure belongs that various modifications, additions, and substitutions are possible without departing from the idea and scope of the present disclosure. Therefore, embodiments of the present disclosure have been described for the sake of brevity and clarity. The scope of the technical idea of ​​the embodiments of the present disclosure is not limited by the descriptions. Consequently, a person skilled in the art in the field to which the present disclosure belongs should understand that the scope of the present disclosure is not limited by the embodiments expressly described above, but rather by the claims and their equivalents.

Claims

[1] Speech synthesis device comprising: a memory configured to store user-configured speech information and speaker audio samples corresponding to the speaker information selected by the user; and a processor configured to generate an audio signal corresponding to the input text by applying a speech synthesis model to the input text, speech information, and audio samples in response to a speech synthesis request from the user, wherein the speech synthesis model is trained to generate an audio signal including features of training text and features of a training audio signal, and wherein speech information of the training text and speaker information of the training audio signal are removed from the generated audio signal. [2] Speech synthesis device according to claim 1, wherein, if the input text includes characters corresponding to multiple languages, the speech information is configured for each character. [3] Speech synthesis device according to claim 1, wherein the speech synthesis model is trained to remove the speech information of the training text and the speaker information of the training audio signal by normalizing a latent training variable including the features of the training text and the features of the training audio signal on the basis of the speech information of the training text and the speaker information of the training audio signal. [4] Speech synthesis device according to claim 1, wherein the speech synthesis model is configured to use the training text, a speech embedding of the training text and a speaker embedding of the training audio signal when generating the duration of each phoneme of the training text, and wherein the speech synthesis model is trained to use speech information inherent in the training text and the speech embedding of the training text and to exclude speech information inherent in the speaker embedding of the training audio signal. [5] Speech synthesis device according to claim 1, wherein the speech synthesis model comprises: a speech embedding module that is set up to convert the speech information into a speech embed; a character embedding module that is set up to convert the input text into character embeddings; an encoder that is set up to encode the character embeddings into text feature vectors; a speaker encoder that is set up to encode the audio samples for outputting a speaker embedding; a duration predictor that is set up to predict phoneme duration data including a duration of each phoneme of the input text based on the text feature vectors, speech embedding, and speaker embedding; a projection module that is set up to generate a distribution of the text feature vectors; a matching unit that is set up to generate a latent variable based on the distribution of text feature vectors and phoneme duration data; an inverted decoder configured to output a transformed latent variable based on the latent variable, speaker embedding, and speech embedding; and an audio generator that is set up to generate the audio signal from the converted latent variable. [6] Speech synthesis device according to claim 5, wherein the inverted decoder is configured to denormalize the latent variable based on the speaker embedding and the speech embedding and to output the converted latent variable based on the denormalized latent variable. [7] Speech synthesis device according to claim 5, wherein the duration predictor is configured, when generating the duration of each phoneme of the input text, to: to use language embedding and speaker embedding; and to use the language information inherent in the input text and the speech embedding, and to exclude the language information inherent in the speaker embedding. [8] Speech synthesis method performed by a speech synthesis device comprising the method: Receiving a speech synthesis request for input text, wherein the speech synthesis request includes user-configured speech information and speaker information; and Generating an audio signal corresponding to the input text by applying a speech synthesis model to the input text, the speech information, and the speaker's audio samples that correspond to the speaker information. wherein the speech synthesis model is trained to generate an audio signal including features of a training text and features of a training audio signal, and wherein speech information of the training text and speaker information of the training audio signal are removed from the generated audio signal. [9] Method according to claim 8, wherein, if the input text includes characters corresponding to multiple languages, the language information is configured for each character. [10] Method according to claim 8, wherein the speech synthesis model is trained to remove the speech information of the training text and the speaker information of the training audio signal by normalizing a latent training variable including the features of the training text and the features of the training audio signal on the basis of the speech information of the training text and the speaker information of the training audio signal. [11] Method according to claim 8, wherein generating the duration of each phoneme of the training text comprises using the training text, the speech embedding of the training text and the speaker embedding of the training audio signal, and training the speech synthesis model to use speech information inherent in the training text and the speech embedding of the training text and to exclude speech information inherent in the speaker embedding of the training audio signal. [12] Method according to claim 8, wherein generating the audio signal corresponding to the input text comprises applying the speech synthesis model to the input text: Converting speech information into a speech embedding using a speech embedding module; Converting the input text into character embeddings using a character embedding module; Encoding the character embeddings into text feature vectors using an encoder; Encoding the audio samples using an audio encoder to output a speaker embedding; Predictions of phoneme duration data, including the duration of each phoneme in the input text, based on text feature vectors, speech embedding, and speaker embedding, using a duration predictor; Generating a distribution of text feature vectors using a projection module; Generating a latent variable based on the distribution of text feature vectors and phoneme duration data using a matching unit; Output of a transformed latent variable based on the latent variable, speaker embedding, and speech embedding by an inverted decoder; and Generating the audio signal using an audio generator from the converted latent variable. [13] Method according to claim 12, wherein outputting the transformed latent variable comprises: Denormalizing the latent variable based on speaker embedding and language embedding; and Outputting the transformed latent variable based on the denormalized latent variable. [14] Method according to claim 12, wherein generating a duration of each phoneme of the input text comprises: Using speech embedding and speaker embedding; and Using speech information inherent in the input text and speech embedding, and excluding speech information inherent in the speaker embedding.