Method for training text-to-speech (TTS) model, TTS apparatus, and method for providing TTS service using TTS apparatus

The TTS model learning method addresses the limitations of existing TTS models by training a main block and speaker encoder to generate diverse voices, including those for infants and characters, achieving efficient and accurate text-to-speech conversion.

WO2025135231A1PCT designated stage expired Publication Date: 2025-06-2642 MARU INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2023/021158
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2023-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing Text-To-Speech (TTS) models, such as VITS, are limited in generating voices for infants and characters, and require additional learning to synthesize diverse voices, which is inefficient.

Method used

A TTS model learning method that trains a main block to convert text into speech and a speaker encoder to extract and reflect speech features, allowing the model to generate voices for various speakers, including infants and characters, without additional main model learning.

Benefits of technology

Enables accurate text-to-speech conversion and easily expands the range of speech outputs, including voices for infants and characters, while efficiently separating the learning processes for text-to-speech conversion and speech feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2023021158_26062025_PF_FP_ABST
    Figure KR2023021158_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a method for training a Text-To-Speech (TTS) model comprising a main block that converts text into speech and a speaker encoder that extracts an utterance feature and reflects same in the main block, the method comprising: training, by using first speech data of a first target, the main block so that the text is converted into speech in at least one of languages of a plurality of countries; extracting an utterance feature for a second target to freeze the main block; and training the speaker encoder by using second speech data of the second target.
Need to check novelty before this filing date? Find Prior Art

Description

A method for learning a TTS (Text-to-Speech) model, a TTS device, and a method for providing TTS services using a TTS device.

[0001] The present invention relates to a learning method of a TTS model, a TTS device, and a method for providing a TTS service using a TTS device.

[0002] With the advancement of artificial intelligence technology, TTS (Text-To-Speech) technology, which converts text into speech, is being utilized in various service fields.

[0003] This TTS technology learns the relationship between human voice and text, and when text is input by a user, it provides a voice that speaks that text in a specific voice.

[0004] In particular, a technology has recently been developed that learns the languages ​​of multiple countries, and when a command to select a specific country is input along with text into the TTS model, a voice is provided with the text spoken in the language of that country.

[0005] For example, YourTTS offers multilingual text-to-speech capabilities based on Zero-Shot Multi-Speaker TTS, which was developed based on the existing VITS (Variational Inference with adversarial learning for end-to-end Text-to-Speech) model, an end-to-end TTS model. However, this VITS model has the disadvantage of being able to synthesize only adult speech.

[0006] Meanwhile, the recent emergence of Large Language Models (LLMs) has led to the emergence of various services and platforms capable of natural human conversation. For example, conversational communication agent services based on generative AI models are emerging. These generative AI model-based conversational communication agent services leverage text-to-speech (TTS) technology to converse with users and even provide the ability to converse in the voice of a person (or character) of the user's or service provider's choice.

[0007] The present invention relates to a TTS model learning method and system capable of uttering text with a type of voice not present in learning data without additional learning of a main model that extracts text features and synthesizes text features and speech features to generate a new voice waveform for a TTS model.

[0008] In particular, the present invention relates to a TTS model learning method and system that improves the learning method of a TTS model, thereby facilitating data expansion for various voices, such as those of infants and characters, in addition to voices corresponding to adults.

[0009] In addition, the present invention relates to a TTS model learning method and system for efficiently learning a main block and a speaker encoder based on VITS (Variational Inference with adversarial learning for end-to-end Text-to-Speech).

[0010] In addition, the present invention relates to a method and system for providing a TTS service that trains a speaker encoder of a TTS model using the voice of a subject.

[0011] Furthermore, the present invention provides a generative AI model-based conversational communication agent service that can speak in a voice desired by a user or service provider, using a TTS model learning method and system.

[0012] In addition, the present invention provides a digital human that speaks in a voice desired by a user or service provider by using a TTS model learning method and system when speaking with a user using a digital human in a conversational communication agent service based on a generative AI model as discussed above.

[0013] In order to solve the problem discussed above, a training method of a TTS model according to the present invention trains a main model that extracts text features and synthesizes the text features with the speech features to generate a new speech waveform, and a speaker encoder that extracts speech features from a speaker's voice. More specifically, a training method of a TTS model according to the present invention is a training method of a Text-To-Speech (TTS) model including a main block that converts text into speech, and a speaker encoder that extracts speech features and reflects them to the main block, the method including: training the main block using first speech data of a first subject so that the text is converted into speech in at least one of a plurality of languages; and freezing the main block and training the speaker encoder using second speech data of the second subject so that the speech features are extracted for a second subject, different from the first subject, and reflected in the main block.

[0014] In particular, the TTS service providing method according to the present invention can additionally train a speaker encoder to speak a voice for a specific speaker using a main model for which learning has been completed.

[0015] In addition, the TTS service providing method according to the present invention can extract voices corresponding to a predetermined time interval from voice data to generate a reference voice, and train a speaker encoder using the generated reference voice.

[0016] In addition, the TTS service providing method according to the present invention can provide a waveform of a voice in which the target text is spoken by inputting a target text, a sample voice, and a country code into a TTS model.

[0017] In addition, a method for providing a TTS service according to the present invention includes a TTS service providing method of a TTS (Text-To-Speech) model including a main block and a speaker encoder, the method comprising: a step of training the TTS model; and a step of inputting a target text for conversion and a sample voice of a target person into the trained TTS model, and converting the target text into the voice of the target person and outputting the converted text by the trained TTS model, wherein the step of training the TTS model may train the main block to convert the text into voice using first voice data of a man and a woman, freeze the main block to extract speech features for the target person and reflect them in the main block, and train the speaker encoder using second voice data of the target person.

[0018] In addition, a TTS device according to the present invention includes a main block for converting text into speech; a speaker encoder for extracting speech features and reflecting them to the main block; and a circuit for driving control logic of the main block and the speaker encoder, wherein the main block is trained to convert the text into speech in at least one of a plurality of languages ​​using first speech data of a first subject, and the speaker encoder can be trained to extract the speech features for a second subject, different from the first subject, and reflect them to the main block by freezing the main block and using second speech data of the second subject. In this case, the second subject can be trained to be able to speak the voice of the second speech data (a specific speaker) to be added.

[0019] In addition, a program stored in a computer-readable recording medium according to the present invention is a program stored in a computer-readable recording medium, which is executed by one or more processes in an electronic device, and which is a program, wherein the program comprises a main block for converting text into speech, and a speaker encoder for extracting speech features and reflecting them to the main block, wherein the method for training a TTS (Text-To-Speech) model includes the steps of: training the main block to convert the text into speech in at least one of a plurality of languages ​​using first speech data of a first subject; and freezing the main block and training the speaker encoder to extract speech features for a second subject, different from the first subject, and reflecting them to the main block. In this case, the second subject can be trained to speak the voice of the second speech data (a specific speaker) to be added.

[0020] Meanwhile, the present invention is derived from research conducted as part of a national project to develop and commercialize a SaaS-type SiteAgent based on a super-large language model.

[0021] [Project ID: 0407231005, Project No.: A0407-23-1005, Ministry of Science and ICT, Project Management (Specialist) Agency: National IT Industry Promotion Agency, Research Project Name: [SaaS Advancement and Intelligence] Support for Promising SaaS Development and Promotion (Ultra-Large AI-Used SaaS), Research Project Name: (New) Development and Commercialization of SaaS-Type SiteAgent Based on Ultra-Large Language Model, Project Implementing Agency Name (Host): Forty2Maru, Research Period: 2023.06.01-2023.12.31]

[0022] According to various embodiments of the present invention, a TTS model learning method and system provides accurate text-to-speech conversion using a TTS model by learning a main model that extracts text features and synthesizes the text features and speech features to generate a new speech waveform, and a speaker encoder that extracts speech features from the speaker's voice, respectively, thereby easily expanding the speaker of the speech output from the TTS model.

[0023] In addition, according to various embodiments of the present invention, the TTS model learning method and system learns a main model using a speaker ID that can distinguish different speakers, and learns a speaker encoder that extracts speech features using the main model for which learning has been completed, thereby separating a learning process for text-to-speech conversion that requires a large amount of learning data and a learning process for extracting speech features that requires a relatively small amount of learning data, thereby performing efficient learning of the TTS model.

[0024] In addition, according to various embodiments of the present invention, the TTS model learning method and system can efficiently and flexibly learn the characteristics of speech from the speaker's speech collected at various lengths by extracting speech corresponding to a predetermined time interval from speech data to generate a reference speech, and training a speaker encoder using the generated reference speech.

[0025] Meanwhile, according to various embodiments of the present invention, a method and system for providing a TTS service can provide a waveform of a voice in which a target text is spoken so that the characteristics of the speech appearing in the sample voice are more clearly displayed by inputting a target text and a sample voice into a TTS model.

[0026] In addition, according to various embodiments of the present invention, the TTS service providing method and system can enable adaptive learning of the subject's speech by training the speaker encoder of the TTS model using the subject's voice.

[0027] Figures 1a and 1b illustrate one embodiment of a TTS device according to the present invention.

[0028] Figure 2 illustrates a learning system of a TTS model according to the present invention.

[0029] Figure 3 illustrates a TTS service providing system according to the present invention.

[0030] Figure 4 is a flowchart showing a learning method of a TTS model according to the present invention.

[0031] Figures 5 and 6 illustrate an embodiment of training the main block of a TTS model.

[0032] Figures 7 and 8 illustrate an embodiment of training a speaker encoder of a TTS model.

[0033] Figures 9a, 9b and 10 illustrate an embodiment of generating a reference voice.

[0034] Figure 11 is a flowchart showing a method for providing TTS service according to the present invention.

[0035] Figures 12 and 13 illustrate an embodiment of generating a voice waveform using a TTS model.

[0036] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Regardless of the drawing numbers, identical or similar components will be given the same reference numbers, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably only for the convenience of writing the specification, and do not in themselves have distinct meanings or roles. In addition, when describing the embodiments disclosed in this specification, if it is determined that a specific description of a related known technology may obscure the gist of the embodiments disclosed in this specification, a detailed description thereof will be omitted. In addition, the attached drawings are only intended to facilitate easy understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, and substitutes included in the spirit and technical scope of the present invention.

[0037] Terms that include ordinal numbers, such as first, second, etc., may be used to describe various components, but the components are not limited by these terms. These terms are used solely to distinguish one component from another.

[0038] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.

[0039] Singular expressions include plural expressions unless the context clearly indicates otherwise.

[0040] In this application, terms such as “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part or combination thereof described in the specification, but should be understood not to exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.

[0041] With the recent rapid development of artificial intelligence (AI) technology, language models capable of natural conversations with people (e.g., ChatGPT) have emerged.

[0042] In this way, language models, unlike existing manually built chatbots that only provide limited responses, are demonstrating technological capabilities that enable natural communication similar to humans and provide fast and accurate information, thereby demonstrating innovation in the artificial intelligence market.

[0043] Meanwhile, companies have a need to provide and promote corporate information to customers through various channels such as corporate websites and corporate messengers, and are making great efforts to promptly respond to customers' requests and inquiries about the company.

[0044] To provide these services, methods and systems for providing interactive communication agent services tailored to businesses are being developed, and methods and systems for providing interactive communication agent services capable of providing accurate and prompt answers to users' inquiries about businesses are emerging. These services are provided in the form of Software as a Service (SaaS) for businesses that have purchased or subscribed to the services according to the present invention. For example, services are provided in the form of providing answers to users (or customers) who have entered user queries for businesses that have purchased the services.

[0045] These services are implemented to provide customers with quick and accurate answers to their (or user) inquiries by utilizing an answer model learned from the company's data. By selectively utilizing an appropriate answer model among multiple answer models based on the characteristics of the customer's inquiry, such as the intent of the inquiry, more efficient and accurate answers can be provided to the customer. Furthermore, the method and system for providing a conversational communication agent service based on such a generative AI model can provide customers with highly reliable, customized answers by identifying the correct answer to the customer's inquiry from the company's data and inputting it into a large model to generate an answer.

[0046] Meanwhile, a generative AI model-based conversational communication agent service can receive user queries from users (or customers) and provide answers to those queries through various channels linked to the company. An example of such a channel could be a chatbot.

[0047] Here, a chatbot can refer to a software application capable of natural conversation with humans. A chatbot can be understood as a computer program that simulates human conversation using artificial intelligence (AI) and natural language processing (NLP), understanding user queries and automatically responding.

[0048] For example, a large language model may include at least one of Generative Pre-trained Transformer (GPT), Bidirectional Encoder Representations from Transformers (BERT), and Language Model for Dialogue Applications (LaMDA).

[0049] Meanwhile, a conversational communication agent service based on a generative AI model can also receive user queries through a digital human trained using data related to a specific company.

[0050] A digital human can be understood as a virtual human created using digital technology (e.g., computer graphics, CG).

[0051] For example, digital humans can interact with users in a variety of ways, including through voice-based conversations.

[0052] Furthermore, the chatbot discussed above can also converse with users based on voice. In this case, the digital human and chatbot can provide a service that answers user questions using a generative AI (Artificial Intelligence) model. Meanwhile, when a digital human or chatbot converses with a user based on voice, the speaking voice of the digital human or chatbot can be set in various ways, and in particular, it can be implemented to speak in a voice desired by the user or service provider. Below, the learning method of a TTS model capable of generating various voices, a TTS device, and a method of providing a TTS service using the TTS device will be examined in more detail with the attached drawings.

[0053] Figures 1a and 1b illustrate an embodiment of a TTS device according to the present invention. Figure 2 illustrates a learning system for a TTS model according to the present invention. Figure 3 illustrates a TTS service providing system according to the present invention.

[0054] Referring to FIGS. 1A and 1B, a TTS device (10) (or TTS model) according to the present invention may include a main block (11) that converts text into speech, and a speaker encoder (12) that extracts speech features and reflects them in the main block (11).

[0055] The TTS device (10) may be a learning model implemented to analyze text to extract text features between words when text and a user's voice are input, analyze the user's voice to extract speech features for the user, and generate a waveform (e.g., speech data) of speech in which the previously input text is spoken by the user's voice using the previously extracted text features and the previously extracted speech features.

[0056] For example, the TTS device (10) can be trained to convert text into speech using training data including training text and training speech data corresponding to the training text, based on an end-to-end TTS model such as YourTTS and VITS.

[0057] In this regard, the TTS device (10) may include a main block (11), a speaker encoder (12), and a circuit for driving control logic of the main block and the speaker encoder.

[0058] The main block (11) analyzes the text to extract text features based on the relationship between words included in the text, and synthesizes the text features and speech features to generate a waveform in which the previously input text is spoken in the user's voice.

[0059] For example, the main block (11) may include a text feature extraction unit (e.g., Char Embedding, Language Embedding, and Transform-Based Encoder) that analyzes text to extract text features, a text-speech alignment unit (e.g., Linear Projection and Monotonic Alignment Search) that aligns (or matches) text features and speech features to extract speech features corresponding to the text features, a posterior encoder (e.g., Posterior Encoder and Flow-Based Decoder) that samples previously extracted speech features so that the speech features extracted in response to the text features correspond to the user's speech features, and a generator (e.g., HiFi-GAN Generator) that converts the sampled speech features into a voice waveform.

[0060] Accordingly, the main block (11) can extract text features from the training text using a text feature extraction unit in a learning process based on training data, and define the distribution of the user's speech features in the training speech data using a rear encoder.

[0061] Accordingly, the main block (11) can align the speech features defined as the distribution of speech features and the previously extracted text features using the text-to-speech alignment unit in a learning process based on learning data.

[0062] Next, the main block (11) can convert speech features previously defined as a distribution of speech features into a voice waveform using a generator in a learning process based on learning data, and can compare the previously converted speech waveform with learning voice data according to the learning data using a discriminator (e.g., Hifi-GAN Discriminator).

[0063] Accordingly, the main block (11) can be trained to achieve a predetermined level of similarity between the previously converted voice waveform and the voice data for learning based on the comparison result by the discriminator in the learning process based on the learning data.

[0064] Meanwhile, the speaker encoder (12) can analyze voice data and extract speech characteristics based on the voice data. The aforementioned rear encoder can learn the text-specific speaking style of TTS by focusing on the person's speaking speed and overall speech feel. In contrast, the speaker encoder focuses on the person's intonation and vocal characteristics, and thus the voice itself is learned using the speaker encoder's speech characteristics.

[0065] In this regard, as shown in Fig. 1a, the speaker encoder (12) may be replaced with a speaker embedding layer during the learning process for the main block (11). The speaker embedding layer may output feature vectors for different speakers, and when speaker information (e.g., Speaker ID) for a specific speaker is input, a feature vector for the speaker corresponding to the input speaker information may be output.

[0066] Here, speaker information may include identification data (e.g., an identification number) for distinguishing different speakers (or users).

[0067] Accordingly, in the learning process for the main block (11), the speaker embedding layer inputs speaker information for a specific speaker corresponding to the learning voice data according to the learning data, and outputs a feature vector corresponding to the input speaker information, so that the feature vector indicating which speaker the learning voice data is a voice spoken by can be transmitted to the main block (11).

[0068] In this way, in the learning process for the main block (11), the speaker embedding layer of the VITS model is not utilized, but rather, a numerical condition is added through Speaker ID, etc., so that the main block is learned while providing information that a specific voice is the voice of a specific speaker. Accordingly, the speaker embedding layer can be defined as a speaker information providing layer with a different function from the speaker embedding layer that performs the function of classifying speakers in the VITS model, but is named as a speaker embedding layer for the convenience of explanation.

[0069] Meanwhile, as shown in FIG. 1b, the speaker embedding layer can be replaced with the speaker encoder (12) when learning for the main block (11) is completed. That is, the TTS model learning system (100) according to the present invention connects the speaker embedding layer to the main block (11) to perform primary learning (e.g., Stage 1) for the main block (11), and when learning for the main block (11) is completed, the main block (11) is frozen, and the speaker embedding layer connected to the main block (11) can be replaced with the speaker encoder (12) to perform secondary learning (e.g., Stage 2) for the speaker encoder (12).

[0070] Here, freezing the main block (11) can be understood as not performing additional learning on the main block (11) for which learning has been completed, and this can be understood as only utilizing the main block (11) for which learning has been completed.

[0071] Therefore, the speaker encoder (12) can extract speech features from learning voice data according to learning data during the learning process for the speaker encoder (12) and transfer the extracted speech features to the main block.

[0072] Accordingly, the main block (11) can generate a voice waveform by synthesizing the learning text according to the learning data and the speech features previously extracted by the speaker encoder (12) during the learning process for the speaker encoder (12), and can compare the previously generated voice waveform with the learning voice data according to the learning data using a discriminator.

[0073] Accordingly, the speaker encoder (12) can be trained so that, in the learning process for the speaker encoder (12), the similarity between the previously synthesized voice waveform and the learning voice data reaches a predetermined level based on the comparison result by the discriminator of the main block (11).

[0074] Referring to FIG. 2, the TTS model learning system (100) according to the present invention can learn a TTS model (or TTS device) including a main block that converts text into speech and a speaker encoder that extracts speech features and reflects them in the main block.

[0075] To this end, the TTS model learning system (100) may include an input unit (110), a storage unit (120), and a control unit (130).

[0076] The input unit (110) can input learning data (101) to train the TTS model (121). To this end, the input unit (110) can be connected to a separately provided server or device via a wireless or wired network, and can receive learning data (101) from the separately provided server or device.

[0077] At this time, the training data (101) may include at least one of training text, training voice data, and speaker information for the training voice data. The training voice data may include a waveform (or Mel-Spectrogram) for the voice of a speaker who utters the training text according to the speaker information.

[0078] Alternatively, the learning data (101) may include first learning data designed to train the main block included in the TTS model (121) and second learning data designed to train the speaker encoder included in the TTS model (121).

[0079] For example, the first training data may include training text and training speech data, which is label data for the training text. The training speech data may include waveforms (or Mel-Spectrograms) of speech produced by male and female speakers speaking the training text in multiple languages ​​(e.g., four countries). The first training data may further include speaker information about the speakers of the training speech data.

[0080] Additionally, the second learning data may include learning texts and learning speech data, which is label data for the learning texts. In this case, the second learning data may include a plurality of different learning texts and a plurality of learning speech data labeled for each of the plurality of learning texts. Here, the plurality of learning speech data may include waveforms (or Mel-Spectrograms) for speech uttered by the same speaker.

[0081] The storage unit (120) can store various data and commands required for the operation of the TTS model learning system (100) according to the present invention. For example, the storage unit (120) can store a TTS model (121) and learning data (101) required for learning the TTS model.

[0082] The control unit (130) can control the overall operation of the TTS model learning system (100) according to the present invention. For example, the control unit (130) can train the TTS model (121) using learning data (101).

[0083] At this time, the control unit (130) can train the TTS model (121) by connecting a speaker embedding layer or a speaker encoder to the main block according to predetermined conditions during the learning process for the TTS model (121).

[0084] In one embodiment, the control unit (130) may perform learning for the TTS model (121) by connecting a speaker embedding layer to the main block in the first learning process, and may perform learning for the TTS model (121) by disconnecting the speaker embedding layer from the main block and connecting a speaker encoder in the second learning process.

[0085] Meanwhile, referring to FIG. 3, the TTS service providing system (200) according to the present invention can provide a TTS service using a TTS model (231) including a main block and a speaker encoder.

[0086] Here, the TTS service may be a service that synthesizes a sample voice (201) of a subject and a target text (202) and provides a waveform (206) for a voice in which the target text (202) is spoken in the voice of the subject.

[0087] At this time, the TTS model (231) can input language information (203) (e.g., Language ID) together with a sample voice (201) and a target text (202), and provide a waveform (206) for a voice spoken in a language according to the language information (203) of the target text (202).

[0088] Additionally, the TTS model (231) may be learned by the TTS model learning system (100) described above, and at this time, the TTS model (231) may be provided in a form in which a speaker encoder is connected to the main block.

[0089] To this end, the TTS service providing system (200) according to the present invention may include an input unit (210), an output unit (220), a storage unit (230), and a control unit (240).

[0090] The input unit (210) can input target text (202), a sample voice (201) of the target person, and language information (203). To this end, the input unit (210) can be connected to a separate voice input device via a wireless or wired network, and receives voice data detected by the voice input device as a sample voice (201), and the target text (202) can be input through various input devices such as a keyboard, a mouse, a touch pad, an optical character reader (OCR), and a magnetic ink character reader.

[0091] In one embodiment, the input unit (210) may be connected to a separately provided server or device via a wireless or wired network, and in this case, at least one of a pre-prepared sample voice (201), target text (202), and language information (203) may be input from the separately provided server or device.

[0092] In another embodiment, the input unit (210) may input at least one of a sample voice (201), a target text (202), and language information (203) from a separate application or program.

[0093] In another embodiment, the input unit (210) may receive each of a sample voice (201), a target text (202), and language information (203) through a combination of at least two of various input devices, servers, devices, applications, and programs. For example, the input unit (210) may receive text recognized by an optical character recognition program as the target text (202), voice data detected by a voice input device as the sample voice (201), and a language item selected by a predetermined input device as the language information (203). As another example, the input unit (210) may receive voice data prepared in advance from a server as the sample voice (201), and receive text input by various input devices as the target text (202).

[0094] The output unit (220) can output a waveform (206) for a voice synthesized through a previously learned TTS model (231). To this end, the output unit (220) can be connected to a voice output device such as a speaker, headset, or earphone via a wireless or wired network, and output the waveform (206) through the voice output device.

[0095] The storage unit (230) can store data and commands necessary for the operation of the TTS service provision system (200) according to the present invention. For example, the storage unit (230) can store a TTS model (231) learned by the TTS model learning system (100).

[0096] Additionally, the storage unit (230) can store sample voice (201) and target text (202) input through the input unit (210), and can also store a waveform (206) for voice synthesized through the TTS model (231).

[0097] The control unit (240) can control the overall operation of the TTS service providing system (200) according to the present invention. For example, the control unit (240) can input the target text (202) and the target person's sample voice (201) into the previously learned TTS model (231), and obtain the waveform (206) of the voice in which the target text (202) is spoken by the target person's voice.

[0098] Based on the configuration of the TTS model learning system (100) and the TTS service providing system (200) discussed above, the TTS model learning method and the TTS service providing method will be described in more detail below.

[0099] FIG. 4 is a flowchart illustrating a method for training a TTS model according to the present invention. FIGS. 5 and 6 illustrate an embodiment of training the main block of a TTS model. FIGS. 7 and 8 illustrate an embodiment of training a speaker encoder of a TTS model. FIGS. 9a, 9b, and 10 illustrate an embodiment of generating a reference voice. FIG. 11 is a flowchart illustrating a method for providing a TTS service according to the present invention. FIGS. 12 and 13 illustrate an embodiment of generating a voice waveform using a TTS model.

[0100] Referring to FIG. 4, the TTS model learning system (100) according to the present invention can train the main block to convert text into voice of at least one of the languages ​​of a plurality of countries using the first voice data of the first target (S110).

[0101] Specifically, the TTS model learning system (100) can train a main block to which a speaker embedding layer is connected using learning data including learning texts containing languages ​​of multiple countries and first speech data (e.g., learning speech data) of a first target labeled in the learning texts.

[0102] At this time, the first target may be a plurality of people who can speak Korean or a foreign language, and the first voice data may be set to Korean and foreign language speech by the first target.

[0103] As a more specific example, the languages ​​of the above multiple countries may be Korean, English, French, and German, and the Language ID may be assigned as Korean 0, English 1, French 2, and German 3, respectively.

[0104] In this case, the text can be received from the Char Embedding layer of the text feature extraction unit, and the Language ID can be assigned from the Language Embedding layer (see Fig. 1).

[0105] Referring to FIG. 5, for example, the TTS model learning system (100) inputs speaker information for learning voice data (32) into the speaker embedding layer (13) connected to the main block (20), and inputs learning text (31) into the main block (20), thereby training the main block (20) so that the voice waveform (41) synthesized by the main block (20) becomes identical (or similar) to the learning voice data.

[0106] As another example, the TTS model learning system (100) can train the main block to extract text features between words of text (e.g., training text) through a transformer encoder (e.g., a text feature extraction unit), extract speech features of first speech data (e.g., training speech data) through a posterior encoder, and generate a waveform corresponding to the text using the text features and speech features.

[0107] The above first voice data may be a sufficient amount of public data, and as a specific example, may be voice data of N languages ​​of adult men and women, totaling approximately 200 hours or less.

[0108] At this time, the main block may be connected to a speaker embedding layer that provides speaker information related to the first voice data.

[0109] Furthermore, the TTS model learning system (100) inputs the voice waveform synthesized in the main block together with the learning voice data used to synthesize the voice waveform into the discriminator of the main block, compares the voice waveform with the learning voice data, and trains the main block based on the comparison result in the discriminator.

[0110] Referring to FIG. 6, for example, the TTS model learning system (100) can input speaker information for learning voice data (32) into the speaker embedding layer (13) connected to the main block (20) and input learning text (31) into the main block (20) to train the main block (20).

[0111] In this case, the text feature extraction unit (21) can extract text features from the previously input training text (32), the speaker embedding layer (13) can specify a feature vector corresponding to the previously input speaker information, and can transmit the specified feature vector to the main block (20), and the rear encoder (23) can input training voice data (32) together with the feature vector specified in the speaker embedding layer (13), and specify speech features in the training voice data (32) according to the voice corresponding to the previously input feature vector.

[0112] A linear spectrogram is required to align speech and text, and a speaker ID can be additionally input through a speaker embedding layer (13). The linear spectrogram input here can be, for example, the actual speech of a specific Korean speaker pronouncing the text. The speaker ID is defined as speaker information, and can include, for example, a Korean male speaker, a female speaker, and an American male speaker.

[0113] The text-to-speech alignment unit (22) can learn the alignment relationship between the text features and the utterance features by inputting the text features extracted from the text feature extraction unit (21) together with the utterance features previously specified from the rear encoder (23), and the generator (24) can convert the utterance features previously specified from the rear encoder (23) into a voice waveform (41) based on the distribution of the utterance features corresponding to the speaker's voice corresponding to the feature vector by inputting the feature vector specified according to the speaker information from the speaker embedding layer (13) together with the utterance features previously specified from the rear encoder (23).

[0114] Accordingly, the discriminator (25) can compare the voice waveform (41) generated by the generator (24) with the learning voice data (32), and train the generator (24) based on the comparison result.

[0115] Referring again to FIG. 4, the TTS model learning system (100) according to the present invention can freeze the main block to extract speech features for a second target, which is different from the first target, and reflect them in the main block, and train the speaker encoder using the second voice data of the second target (S120).

[0116] In this case, freezing the main block can be defined as meaning that the main block processes the input as a model that has completed learning without performing learning and outputs the result.

[0117] Here, the second target can be set as the target for which voice is to be implemented through the TTS model, and may include the voices of various targets, such as the voice of a specific character or the voice of a specific user. Furthermore, the second voice data may be relatively smaller (or shorter) than the first voice data, targeting the target for which voice is to be implemented.

[0118] As a more specific example, the second voice data may be voices of M individuals whose voices are to be embodied, spanning approximately 10 hours (or any time between approximately 5 and 20 hours). In this case, the second voice data may include voices spoken in various languages, such as Korean, English, French, and German.

[0119] In this case, when learning of the main block is completed, the TTS model learning system (100) can freeze the main block, replace the speaker embedding layer connected to the main block with a speaker encoder, and perform learning for the speaker encoder.

[0120] For example, when the learning of the main block is completed with the flag corresponding to Stage 1 raised (e.g., Stage 1 == True) in relation to the learning status for the TTS model, the TTS model learning system (100) can lower the flag corresponding to Stage 1 (e.g., Stage 1 == False) and raise the flag corresponding to Stage 2 (e.g., Stage 2 == True) to replace the speaker embedding layer connected to the main block with a speaker encoder.

[0121] That is, a state in which the flag corresponding to Stage 1 is raised may be a state in which a speaker embedding layer is connected to the main block, and information (e.g., a feature vector) output from the speaker embedding layer is input to each component included in the main block.

[0122] Additionally, a state in which the flag corresponding to Stage 2 is raised may be a state in which a speaker encoder is connected to the main block, and information (e.g., speech characteristics) output from the speaker encoder is input to each component included in the main block.

[0123] As another example, in the TTS model learning system (100), if the value corresponding to the Stage indicating the learning status for the TTS model is a value corresponding to the first learning for learning the main block (e.g., 1), a speaker embedding layer is connected to the main block, so that information (e.g., a feature vector) output from the speaker embedding layer is input to each component included in the main block, and if the value corresponding to the Stage is not a value corresponding to the first learning or is a value corresponding to the second learning for learning the speaker encoder (e.g., 0), a speaker encoder is connected to the main block, so that information (e.g., a speech feature) output from the speaker encoder is input to each component included in the main block.

[0124] Furthermore, the TTS model learning system (100) can input second voice data (e.g., voice data for learning) of a second target into the speaker encoder and input text for learning into the main block to learn the main block to which the speaker encoder is connected, thereby obtaining a voice waveform in which the text for learning is spoken in the voice of the second target.

[0125] Referring to FIG. 7, for example, the TTS model learning system (100) can generate a reference voice (37) containing the speech of a specific speaker (e.g., a second target) from learning voice data (36) (e.g., second voice data), and input the reference voice (37) into the speaker encoder (14) to perform learning for the speaker encoder (14).

[0126] In this case, the speaker encoder (14) extracts speech features (or voice features) based on the previously input reference voice (37), and the speech features extracted from the speaker encoder (14) are input to the rear encoder of the main block (50) where learning has been completed in advance, so that the speaker encoder (14) can be trained so that a voice waveform (42) corresponding to the text is generated.

[0127] That is, the TTS model learning system (100) can input the learning text (35) into the main block (50) together with the speech features extracted from the speaker encoder (14), and generate a voice waveform (42) in which the learning text is spoken in the voice of a specific speaker.

[0128] At this time, a Linear spectrogram and Language ID for aligning voice and text can be additionally input through the main block (50). For example, the Linear spectrogram input above can be a voice of a specific speaker who wants to extract speech features through a speaker encoder (14) pronouncing the training text (35). In addition, the training voice data (36) used to generate a reference voice (37) can be a voice of a specific speaker corresponding to the Linear spectrogram pronouncing a text different from the training text (35).

[0129] In other words, the learning voice data (36) may include multiple voices of a specific speaker pronouncing different texts, and among the multiple voices, the voice pronouncing the learning text (35) input to the main block (50) is input to the main block (50) as a linear spectrogram, and a reference voice (37) in which at least two or more voices are combined among multiple voices other than the voice corresponding to the linear spectrogram among the multiple voices may be input to the main block (50).

[0130] Therefore, the TTS model learning system (100) can train the speaker encoder (14) using the voice waveform (42) generated in the main block (50) and the voice data for learning (36).

[0131] Furthermore, the TTS model learning system (100) inputs the voice waveform synthesized in the main block together with the learning voice data used to synthesize the voice waveform into the discriminator of the main block, compares the voice waveform with the learning voice data, and trains the speaker encoder (14) based on the comparison result in the discriminator.

[0132] Referring to FIG. 8, for example, the TTS model learning system (100) can generate a reference voice (37) using learning voice data (36), and input the previously generated reference voice (37) and learning text (35) into the speaker encoder (14) connected to the main block (20) where learning has been completed, thereby training the speaker encoder (14).

[0133] In this case, the text feature extraction unit (51) can extract text features from the previously input training text (35), and the text-speech alignment unit (52) can align the text features and the speech features by specifying the speech features corresponding to the text features.

[0134] At this time, the text-speech alignment unit (52) aligns the text features and the speech features using the speech features learned in the learning process for the main block (50). At this time, the speech features learned in the learning process for the main block (50) may be different from the speech features extracted from the speaker encoder (14).

[0135] Next, the speaker encoder (14) can extract speech features for a specific speaker from the previously input reference voice (37). Accordingly, the rear encoder (53) inputs the speech features aligned in the text-to-speech alignment unit (52) and the speech features extracted from the speaker encoder (14), and can sample the speech features aligned in the text-to-speech alignment unit (52) using the speech features extracted from the speaker encoder (14).

[0136] That is, the rear encoder (53) can substitute the speech features aligned in the text-to-speech alignment unit (52) with the speech features extracted from the speaker encoder (14). At this time, the speech features extracted from the speaker encoder (14) may be a distribution of speech features extracted from the voice of a specific speaker, and therefore, the rear encoder (53) can sample the speech features aligned in the text-to-speech alignment unit (52) as speech features for the voice of a specific speaker based on the distribution of the speech features.

[0137] Through this, the generator (54) can input the speech features extracted from the speaker encoder (14) together with the speech features sampled from the rear encoder (53), and convert the speech features sampled from the rear encoder (53) into a voice waveform (42) based on the distribution of the speech features.

[0138] Accordingly, the discriminator (55) can compare the voice waveform (42) generated by the generator (54) with the learning voice data (36), and train the speaker encoder (14) based on the comparison result.

[0139] Furthermore, the TTS model learning system (100) can extract first data from second voice data (e.g., voice data for learning), cut out voices of a specific time range from the extracted first data, and generate a reference voice using the cut out voices.

[0140] Referring to FIG. 9a, for example, the TTS model learning system (100) can extract first data (61) excluding learning voice data (63), which is correct answer data for learning text, from a voice data set (60) for a specific speaker.

[0141] Here, the speech data set (60) may include a plurality of training speech data labeled for each of a plurality of training texts. That is, the speech data set (60) may include a plurality of speech data generated by a specific speaker speaking different texts.

[0142] Accordingly, the first data (61) can be randomly extracted from a voice data set (e.g., second voice data) excluding the correct answer data corresponding to the training text. In addition, the first data (61) can be data of the voice spoken by a specific speaker corresponding to the training voice data without being preprocessed.

[0143] Through this, the TTS model learning system (100) can use the voice of a specific speaker as a source of learning data for extracting speech features, rather than preprocessing the voice data into an average value or the like, and can learn a new voice of a specific speaker by using a voice other than the correct answer data corresponding to the learning text.

[0144] Next, the TTS model learning system (100) can extract a voice corresponding to a predetermined specific time range from the previously extracted first data at an arbitrary (or random) point in time, and use the voice corresponding to the extracted specific time range as a reference voice (65).

[0145] As another example with reference to FIG. 9b, the TTS model learning system (100) can extract first learning voice data (or first data (61)) and second learning voice data (or second data (62)) excluding learning voice data (63), which is correct answer data for learning text, from a voice data set (60) for a specific speaker.

[0146] Here, the first data (61) and the second data (62) can be randomly extracted from the voice data set (60), excluding the correct answer data corresponding to the training text. That is, the first data (61) and the second data (62) can be different training voice data.

[0147] Next, the TTS model learning system (100) can extract a voice corresponding to a predetermined specific time range from the previously extracted first data (61) and second data (62) at any (or random) point in time, and concatenate the voice corresponding to the specific time range extracted from the first data (61) and the voice corresponding to the specific time range extracted from the second data (62) to generate a reference voice.

[0148] Referring to FIG. 10, the TTS model learning system (100) can extract (or split) a first voice (66) corresponding to a predetermined specific time range from the first data (60a) at an arbitrary (or random) point in time, and can extract a second voice (67) corresponding to a predetermined specific time range from the second data (60b) at an arbitrary (or random) point in time.

[0149] At this time, the predetermined specific time range may be determined as a specific time interval, or, an arbitrary time interval may be specified and determined among a range for a predetermined time interval (e.g., 3 seconds to 10 seconds).

[0150] In one embodiment, the TTS model learning system (100) can specify two time ranges based on time intervals corresponding to the length of the voice data for learning, and extract voices corresponding to each of the two previously specified time ranges from the first data (60a) and the second data (60b).

[0151] Through this, the TTS model learning system (100) can generate a reference voice (65a) by connecting the first voice (66) and the second voice (67) extracted from the first data (60a) and the second data (60b), respectively.

[0152] Through the process described above, the second voice data, together with the learning text, Language ID, Ref. wav (see FIG. 10, randomly processed voice), and wav (Linear Spectrogram), can be input into the speaker encoder to perform learning to extract the speaker's speech characteristics.

[0153] In addition, through the above configurations, the TTS model learning system (100) according to the present invention learns a main model that extracts text features and synthesizes text features and speech features to create a new speech waveform for the TTS model, and a speaker encoder that extracts speech features from the speaker's voice, thereby providing accurate text-to-speech conversion using the TTS model and easily expanding the speaker of the speech output from the TTS model.

[0154] In addition, the TTS model learning system (100) according to the present invention learns a main model using a speaker embedding layer capable of distinguishing different speakers, and learns a speaker encoder that extracts speech features using the main model for which learning has been completed, thereby separating a learning process for text-to-speech conversion that requires a large amount of learning data and a learning process for extracting speech features that requires a relatively small amount of learning data, thereby enabling efficient learning of the TTS model.

[0155] In addition, the TTS model learning system (100) according to the present invention extracts voices corresponding to a predetermined time interval from voice data to generate a reference voice, and trains a speaker encoder using the generated reference voice, thereby efficiently and flexibly learning the characteristics of speech from the voice of a speaker collected at various lengths.

[0156] Meanwhile, referring to FIG. 11, the TTS service providing system (200) according to the present invention trains a TTS model (S210), and when a target text for conversion and a sample voice of a target person are input into the trained TTS model, the trained TTS model can convert the target text into the voice of the target person and output it (S220).

[0157] Referring to FIG. 12, for example, the TTS service providing system (200) inputs a target text (81) and a sample voice (82) of the target person into a learned TTS model (70), and through the TTS model (70), speech characteristics according to the sample voice (82) are applied to obtain a voice waveform (90) in which the target text (81) is spoken.

[0158] Referring to FIG. 13, for example, the TTS service providing system (200) can input target text (81) and sample voice (82) into the TTS model (70).

[0159] Accordingly, the text feature extraction unit (71) of the TTS model (70) can extract text features from the previously input target text (81), and the text-speech alignment unit (72) can align the text features and the speech features by specifying the speech features corresponding to the text features.

[0160] Additionally, the speaker encoder (73) can extract speech features of the subject from the previously input sample voice (82).

[0161] Accordingly, the rear encoder (74) receives the speech features aligned in the text-to-speech alignment unit (72) and the speech features extracted from the speaker encoder (73), and can sample the speech features aligned in the text-to-speech alignment unit (72) using the speech features extracted from the speaker encoder (73).

[0162] Through this, the generator (75) can input the speech features extracted from the speaker encoder (73) together with the speech features sampled from the rear encoder (74), and convert the speech features sampled from the rear encoder (74) into a voice waveform (90) based on the distribution of the speech features.

[0163] Furthermore, the TTS service provision system (200) may perform additional training on the speaker encoder of the TTS model using training text and a training voice in which the training text is spoken by the subject. In this case, the training voice may be correct answer data for the training text.

[0164] For example, the TTS service provision system (200) can provide a predetermined learning text so that the subject can confirm it, and can receive a learning voice as the correct answer data for the learning text.

[0165] Accordingly, the TTS service providing system (200) can generate learning data using learning text and learning voice, and perform the learning process of the speaker encoder described above using the learning data.

[0166] Through this, the TTS service provision system (200) can implement a TTS model whose learning has been completed using learning data collected from the subject, and can provide a voice waveform that more accurately imitates the subject's voice using this TTS model.

[0167] That is, the TTS service provision system (200) performs additional learning on the speaker encoder of the TTS model using learning data collected from the subject for the TTS model for which the main block has been pre-trained (or, pre-trained), and provides a voice waveform simulated as the subject's voice for the target text using the TTS model for which learning has been completed.

[0168] Through the above configurations, the TTS service providing system (200) according to the present invention can provide a waveform of a voice in which the target text is spoken so that the characteristics of the speech appearing in the sample voice are more clearly displayed by inputting the target text and sample voice into the TTS model.

[0169] In addition, the TTS service providing system (200) according to the present invention can enable adaptive learning of the subject's speech by training the speaker encoder of the TTS model using the subject's voice.

[0170] Furthermore, the present invention discussed above can be implemented as a program executed by one or more processes in an electronic device and stored in a computer-readable recording medium.

[0171] Accordingly, the present invention can be implemented as computer-readable code or instructions on a program-recorded medium. That is, the various control methods according to the present invention can be provided in the form of integrated or individual programs.

[0172] Meanwhile, computer-readable media include all types of recording devices that store data that can be read by a computer system. Examples of computer-readable media include hard disk drives (HDDs), solid-state disk drives (SSDs), silicon disk drives (SDDs), ROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices.

[0173] Furthermore, the computer-readable medium may include a storage device and may be a server or cloud storage device accessible via communication. In this case, the computer may download the program according to the present invention from the server or cloud storage device via wired or wireless communication.

[0174] Furthermore, in the present invention, the computer described above is an electronic device equipped with a processor, i.e., a CPU (Central Processing Unit), and there is no particular limitation on its type.

[0175] Meanwhile, the detailed description above should not be construed as limiting in any respect and should be considered illustrative. The scope of the present invention should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the present invention are intended to be included within the scope of the present invention.

Claims

1. A learning method for a TTS (Text-To-Speech) model including a main block that converts text into speech and a speaker encoder that extracts speech features and reflects them in the main block, A step of training the main block to convert the text into speech in at least one of the languages ​​of a plurality of countries using the first speech data of the first subject; and A method for learning a TTS model, comprising the steps of freezing the main block to extract the speech features for a second target, different from the first target, and reflecting the extracted speech features in the main block, and training the speaker encoder using second voice data of the second target.

2. In paragraph 1, The above first target is set to be Korean and foreign language speech of multiple characters, A learning method for a TTS model, characterized in that the second target is set as a target to implement voice through the TTS model.

3. In paragraph 2, A learning method for a TTS model, characterized in that the second voice data is a smaller amount of data than the first voice data, targeting the target to implement the voice.

4. In paragraph 1, The step of training the above speaker encoder is: A learning method for a TTS model, characterized by generating a reference voice containing the speech of the second subject from the second voice data and performing learning using the reference voice.

5. In paragraph 4, The above reference voice is, A learning method for a TTS model, characterized in that the method comprises cutting out a specific time range of voices from each of the first data and the second data extracted from the second voice data, and connecting the cut out voices to each other.

6. In paragraph 5, A learning method for a TTS model, characterized in that the first data and the second data are randomly extracted from the second voice data, excluding the correct answer data corresponding to the text.

7. In paragraph 1, The steps for training the above main block are: Extract text features between words in the text using a transformer encoder, Extracting voice features of the first voice data through a posterior encoder, A learning method for a TTS model, characterized by generating a waveform corresponding to the text using the text features and the voice features.

8. In paragraph 6, The step of training the above speaker encoder is: A method for learning a TTS model, characterized in that the speaker encoder is trained to generate a waveform corresponding to the text by inputting the speech features extracted from the speaker encoder into the rear encoder of the main block where the learning is completed.

9. In paragraph 6, In the step of learning the above main block, a speaker embedding layer that provides speaker information related to the first voice data is connected to the main block, A method for learning a TTS model, characterized in that in the step of learning the speaker encoder, the speaker encoder is connected to the main block by replacing the speaker embedding layer.

10. A method for providing a TTS service of a TTS (Text-To-Speech) model including a main block and a speaker encoder, a step of training the above TTS model; and When the target text for conversion and the sample voice of the subject are input to the learned TTS model, the learned TTS model includes a step of converting the target text into the voice of the subject and outputting it. The steps for training the above TTS model are: Using the first voice data of a man and a woman, the main block is trained to convert the text into voice, A method for providing a TTS service, characterized in that the main block is frozen so as to extract speech features for the subject and reflect them in the main block, and the speaker encoder is trained using the second voice data of the subject.

11. Main block that converts text to speech; A speaker encoder that extracts speech features and reflects them in the main block; and Includes a circuit for driving the control logic of the above main block and speaker encoder, The above main block is, Using the first speech data of the first subject, the text is trained to be converted into speech in at least one of the languages ​​of a plurality of countries, The above speaker encoder, A TTS (Text-To-Speech) device characterized in that the main block is frozen so as to extract the speech characteristics for a second target, different from the first target, and reflect them in the main block, and the second voice data of the second target is used for learning.

12. A program that is executed by one or more processes in an electronic device and stored in a computer-readable recording medium, The above program is, A method for learning a TTS (Text-To-Speech) model, comprising a main block that converts text into speech and a speaker encoder that extracts speech features and reflects them in the main block, A step of training the main block to convert the text into speech in at least one of the languages ​​of a plurality of countries using the first speech data of the first subject; and A program stored on a computer-readable recording medium, characterized in that it includes commands for performing a step of freezing the main block so as to extract the speech characteristics for a second target, different from the first target, and reflecting the extracted speech characteristics in the main block, and training the speaker encoder using the second voice data of the second target.

Citation Information

Patent Citations

  • Method and apparatus for providing multi-country multi-channel e-commerce fulfillment integrated solution service using big data and artificial intelligence

    KR1020230088207A

  • Display panel

    KR1020230109211A

  • Road rainwater drainage system

    KR1020250008308A

  • Multilingual neural text-to-speech synthesis

    US20220246136A1

  • KR20200143659A